PToL Year 4 Roadmap - Q3 Status

PToL Year 4 Roadmap - Q3 Status

iPToL Year 4 Roadmap

Revision 2.3 – Q3

The June-September period has been largely occupied by outreach efforts, with iPToL presence at the following conferences:

  • Evolution (oral presentation: NM)

  • iEvoBio (2 oral presentations: NM, JE)

  • Botany (workshop: NM, SM; poster: MN)

  • Plant biology (workshop: NM, poster: MN)

  • UseR (poster: NM)

In certain area, the progress of the iPToL has been affected by changes in the discovery environment development plan. This affected in particular data management, visualization and integration into the discovery environment. As a result some original objectives have been deprecated or postponed to a later date to be determined according to the overall development. In some cases, the working groups have taken over additional projects.

Data Assembly

General Data Assembly

After some initial difficulties, the different components of the Data Assembly are now either on-track for delivery or already completed. The data upload and management capabilities have been greatly improved through a refactoring of the DE backend. The work on the ingestion pipelines is also moving forward. Collaborative tools are planned as part of the future DE development.
Deliverables:

  • An industrial strength pipeline for data assembly for very large data sets;

  • Collaboration and Analysis Tools (contribute one's own data, download data, analyze data, keep data/results private) through the DE.

Strategy:

  • Engage iPlant faculty (Nirav Merchant, Sudha Ram, Eric Lyons) for domain expertise in data infrastructure, meta-data management and scientific workflows.

Tasks:

  • Robust data upload capability (IRODS; Rion Dooley, Nirav Merchant), late 11Q1?

    • COMPLETED

  • Meta-data management (Rion Dooley, Nirav Merchant, Sudha Ram), early 11Q2;

    • POSTPONED/IN PROGRESS

  • Robust data storage and retrieval, collaboration tools in DE (iPlant-wide requirement; core software), 11Q2;

    • DATA STORAGE: COMPLETED

    • COLLABORATION TOOLS: POSTPONED/IN PROGRESS

  • Advanced collaboration tools (iPlant-wide requirement; core software), 11Q3;

    • POSTPONED/IN PROGRESS

  • Input data validation (core software in collaboration with data integration ETAs);

    • DEPRECATED: The original database/metadata driven model for the discovery environment has been replaced by a file-based storage. This requires a redesign of the role metadata will have in the new context.

  • Multiple sequence alignment generation:

    • PHLAWD (John Cazes, Stephen Smith);

      • ONGOING

    • Gordon Burleigh's pipeline (John Cazes, Eric Lyons, Stephen Smith);

      • STALLED – BEING RESOLVED

    • Muscle, other alignment strategies (John Cazes, Eric Lyons).

      • COMPLETED

  • Sequence database(s) bad GIs for PHLAWD, etc (Sheldon McKay and delegates), 11Q2.

    • NCBI'S GENBANK ADDED

My-Plant/My-Crop

The backend has been successfully redesigned, which will allow the further development of My-Crop
Deliverables:

  • My-Plant: Robust, widely used scientific collaboration network for the plant sciences based on a phylogeny metaphor;

  • My-Crop: In support of integrated breeding platform, a scientific interaction site as well as data landing pad.

Strategy:

  • Build a phylogenetically-structured social networking website for information sharing and collaboration;

  • Generalize and extend to other display/networking paradigms.

Milestones:
1. Official launch 10Q3.
Status:

  • 239 users, 57 clades (Feb 3, 2010).

Tasks:

  • Refactor back-end for generic 'clade' structure to facilitate other display/organization paradigms (Matt Hanlon), late 11Q1:

    • Node based;

    • Drupal core taxonomy.

      • COMPLETED

  • Implement Drupal module to ingest and provide basic search functionality for relevant literature citations from the user community and public databases (Steve Mock and delegates), 11Q2;

  • Integration with Facebook and other similar sites. Initially with passive linkouts, possible later with Facebook page or app (Steve Mock and delegates), 11Q3;

  • Integration (as consumer) of TNRS and iPlant tree viewer (Steve Mock; Matt Hanlon);

  • My-Crop (different paradigm for display – also data repository):

    • scoping, early 11Q2;

    • early implementation, late 11Q2;

    • full implementation, early 11Q3.

Trait Evolution

Initial work on the integration of tree stretching (evolutionary) models identified major problems with the optimization routines used by the underlying R package. Because the problems are severe and affect other packages and functions (optim) that are used across the spectrum of biological sciences, the group decided to further investigate the problem, with specific application to phylogenetic questions. A statistics graduate student has been integrated into the group to work on a project that should elucidate the conditions under which the performance of the optimization routines yield unreliable results and formulate suggestions on which routines are more appropriate.
Moreover, members of the group have integrated 4 additional tools into the discovery environment. A new tool written by members of the working group will be integrated and linked from the publication describing it.
Deliverables:

  • An infrastructure for trait analysis and ancestral characters estimation.