Importing data

The visual interface for imports with uploads was added in the recent G-nom rework. It is functional but expected to undergo visual rennovations in the near future.

All imported data references a species NCBI ID, as outlined in the data hierarchy. The first step of every import is selecting the NCBI taxonomy ID in the dialogue seen below.

Import Taxon selection
Screenshot 1: The Taxon selection interface with the NCBI ID 7460 (Apis mellifera) entered

Following the data hierarchy, your imported data is either a new assembly or references an existing assembly (e.g. mappings or annotations). The next dialogue allows you to decide between importing a new assembly or adding data for an exisiting one.

Import Assembly selection
Screenshot 2: Switching between new and existing assemblies

Importing assemblies

Future version of G-nom will automatically allow you to automatically import reference assemblies from IDs. You can track progress here.
flowchart TD id1(Upload assembly) --> id2(Calculate assembly statistics) id1(Upload assembly) --> id3(bgzip file for storage) id3(bgzip file for storage) --> id4(Generate FASTA index) id4(Generate FASTA index) --> id5(JBrowse import) id5(JBrowse import) --> id6(Mark assembly completed) id2(Calculate assembly statistics) --> id6(Mark assembly completed)
  1. Specify the NCBI taxonomy ID of the species corresponding to your assembly as outline above

  2. Under Import or select a new assembly, tick the Import a new assembly tickbox

  3. Enter a name for your assembly, this will later be displayed in the assembly page header and on the assembly card

  4. If your assembly is a reference assembly sourced from a database like NCBI or GENCODE, specify the stable ID of the assembly in the reference assembly ID field

  5. At the bottom of the page click the start import button

Importing annotations

flowchart TD id1(Upload assembly) --> id2(gff3sort) id2(gff3sort) --> id3(bgzip file for storage) id3(bgzip file for storage) --> id4(Generate tabix adapters) id4(generate tabix adapters) --> id5(Mark annotation complete)
  1. Specify the NCBI taxonomy ID of the species corresponding to your annotation

  2. Select a pre-existing assembly

  3. Under Import Analyses > Annotation, select a GFF file to upload and specify a custom name for the annotation. This name will be used as track label in the genome browser

Importing Mappings

Mapping to larger genome can lead to large mapping files exceeding 20GB. Depending on your webserver configuration, you may encounter problems uploading. In this case, dispatch the import job manually in the server filesystem.
flowchart TD id1(Upload mapping) --> id2(Compress if uncompressed) id2(Compress if uncompressed) --> id3(Index) id3(bgzip file for storage) --> id5(Mark mapping complete)
  1. Specify the NCBI taxonomy ID of the species corresponding to your mapping

  2. Select a pre-existing assembly

  3. Under Import Analyses > Mapping, select a SAM or BAM file to upload and specify a custom name for the mapping. This name will be used as track label in the genome browser.

Importing BUSCO analyses

The BUSCO analysis can be generated automatically using the G-nom core pipeline.
  1. Specify the NCBI taxonomy ID of the species corresponding to your analysis

  2. Select a pre-existing assembly

  3. In your BUSCO output locate the file matching short_summary.specific.<lineage_dataset>.<output_folder>.json

  4. Under Import Analyses > BUSCO, select the aforementioned file and specify a custom name for the analysis.

Importing fCat analyses

  1. Specify the NCBI taxonomy ID of the species corresponding to your analysis

  2. Select a pre-existing assembly

  3. In your fCat output locate the file fcat_report_summary.txt

  4. Under Import Analyses > fCat, select the aforementioned file and specify a custom name for the analysis.

Importing Repeatmasker analyses

G-nom currently does not support metadata for repeat libraries used to perform the analysis. This feature will be addad in a future release.
  1. Specify the NCBI taxonomy ID of the species corresponding to your analysis

  2. Select a pre-existing assembly

  3. In your RepeatMasker Output locate the repeatmasker summary file (.tbl.out) and the corresponding FNA file (.fna)

  4. Under Import Analyses > RepeatMasker, select the aforementioned files and specify a custom name for the analysis.

Importing taXaminer analyses

flowchart TD id1(Upload analysis) --> id2(unzip) id2(unzip) --> id3(move to storage) id3(move to storage) --> id4(index protein FASTA file) id4(index protein FASTA file) --> id5(Parse DIAMOND hits into database) id5(Parse DIAMOND hits into database) --> id6(Mark import complete)

Importing taXaminer analyses requires a .zip file containing a subset of taXaminer’s output. Namely, the file should contain:

  1. contribution_of_variables.csv

  2. gene_table_taxon_assignment.csv

  3. proteins.faa

  4. summary.txt

  5. taxonomic_hits.txt

  6. Specify the NCBI taxonomy ID of the species corresponding to your analysis

  7. Select a pre-existing assembly

  8. Under Import Analyses > taXaminer, select the aforementioned .zip file and specify a custom name for the analysis.