This analysis is part of a forthcoming manuscript entitled Classifying Bacterial Genomes Below the Species Rank.
Beatriz Vieira Mourato, Sara-Lena Welk, Fabian Klötzl, and Bernhard Haubold
This tutorial has six main dependencies:
- Neighbors for finding target and neighbor genomes
- datasets and dataformat for downloading and analyzing genomes
- Fur for finding unique genome regions
- Prim for designing PCR primers
- Biobox for general sequence manipulation
- the Unix tools
curl,bzip2, andzipfor downloading and decompressing files
If you are on a Debian system like WSL/Ubuntu, you can install these
dependencies into ~/bin/ by executing from inside the o111h8 repo
bash scripts/setup.shOnce that’s done, make sure ~/bin/ is in your path by running
source ~/.profileWe have tested this setup on our “minimal box” Docker container, mix.
Execute
make datato generate the directory data and download the Neighbors database,
neidb, inside it. Since neidb is 1.3 GB large, this takes a couple
of minutes, depending on the speed of your internet connection. The
result consists of four files inside data,
neidb, the Neighbors database calculated from the Genbank assemblies, Refseq assemblies, and the taxonomy database downloaded from the NCBI on 17th June 2026eco.json, the genome summaries of the 7884 E. coli genomes assembled to level “complete”eco.nwk, the tree of the 7361 complete E. coli genomes that passed the quality filtersero.txt, the serotypes of the 7361 complete E. coli genomes calculated with ectypermlst.txt, the sequence types of the 7361 complete E. coli genomes retrieved using the program mlst
Make the tutorial and change into it.
make tutorialThis now contains the scripts and data files for following the tutorial described in the doc. We can also run all scripts to test them.
cd tutorial
bash testTut.sh