- Author: Kaitlin Weber
- Commissioned by: Rijksinstituut voor Volksgezondheid en Milieu (RIVM)
This python script combines two in-house databases, into one multi-FASTA file. In the multi-FASTA file, by giving an email (linked to NCBI account) as input and the scientific name from the files, the taxonomy ID is placed in the header. The taxonomy ID is the reason this database can be made compatible with Kraken 2.
This script was designed for creating a database for the purpose of using it with ExId16S.
- Clone the repository:
git clone https://github.com/Kaitlinweber/combine_16S_database
- Enter the directory with the pipeline and install the conda environment:
cd combine_16S_database
conda env install -f envs/database_env.yaml
-db, --databasePath to directory, which can have multiple subdirectories with 16S rDNA information. The files need to contain ORGANSIM/DEFENITION (for extraction scientific name), ORIGIN (for extraction sequence) (format derived from Genbank format, but in text file format)-e, --emailEmail which is linked to a NCBI account, this is for obtaining the taxonomy ID, which is required for building a Kraken 2 database.-o, --outputPathway to output directory, if directory does not exists, directory will be created
python combine_16S_database.py -db [path/to/input/dir/16Sdata] -o [email from NCBI] -o [path/to/output/dir]