Page 10 - Read Online
P. 10

Page 4 of 21               O’Connell et al. Microbiome Res Rep 2023;2:21  https://dx.doi.org/10.20517/mrr.2023.17

               MP taxonomy, the initial dataset of MP genomes was restricted to those available from the RefSeq database,
               thereby  reducing  the  data  set  to  752  genomes  (https://www.ncbi.nlm.nih.gov/labs/virus/vssi/#/
               virus?SeqType_s=Genome, accessed 14 December 2021). The genomes were then manually organised into
               their respective clusters, which were assembled from publications associated with the characterisation of MP
               and the Actinobacteriophage Database (https://phagesdb.org/hosts/genera/1/?sequenced=True, accessed 14
               December 2021). At this point, it was decided to introduce exclusion criteria to remove small clusters of
               genomes as the scope of this study is quite broad, and it was hypothesised that novel groupings would be
               more apparent in larger clusters due to the larger sample size. The criteria for exclusion from the study
               were: (1) MP that could not be assigned to a cluster based on the literature/database information; (2)
               clusters represented by ≤ 3 MP genomes (considered underrepresented); and (3) MP described as singletons
               (which lack sufficient nucleotide identity and/or shared gene content to be clustered with known
                      [28]
               phages;  ). Following these criteria, 15 groupings (comprised of 30 genomes total) of MP genomes were
               removed from the dataset prior to VIRIDIC analysis, as detailed in the Results. To ensure the removal of
               these 30 genomes was not likely to influence the outcome of the analyses, each genome was analysed with
               B L A S T N           ( h t t p s : / / b l a s t . n c b i . n l m . n i h . g o v /
               Blast.cgi?PROGRAM=blastn&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome)to verify that they were
               not closely related to the remaining clusters.


               Manual curation of the taxonomic information related to the MP dataset
               The complete taxonomy of the shortlisted MP (721 genomes) from the RefSeq database was obtained from
               the NCBI (https://www.ncbi.nlm.nih.gov/; accessed 10-14 December 2022) and later cross-referenced with
               the most up-to-date information available from the (ICTV) Master Species List for conformational purposes
               (https://talk.ictvonline.org/files/master-species-lists/m/msl/12314, accessed 10 December 2022).

               Genomic analysis of the MP genomes using VIRIDIC and comparison of the VIRIDIC predicted
               taxonomy with the currently accepted taxonomic and subcluster classifications to identify novel
               groups
               The whole genome sequences of the MP associated with each cluster were curated in FASTA format for the
               VIRIDIC analyses. The FASTA files were obtained from NCBI Virus (https://www.ncbi.nlm.nih.gov/labs/
               virus/vssi/#/virus?SeqType_s=Genome&SourceDB_s=RefSeq). The input format for VIRIDIC requires that
               all the genomic FASTA sequences included in the analysis are compiled into a single file. Once a file had
               been created for each cluster, they were individually inputted into the VIRIDIC web portal (http://
               rhea.icbm.uni-oldenburg.de/VIRIDIC/;  ). The resulting heatmaps and predicted genus groupings were
                                                  [10]
               compared with the current taxonomy for each cluster to identify potentially novel genera. The VIRIDIC
               outputs were also compared to the Actinobacteriophage database assigned subclusters to verify the genus-
               subcluster hypothesis and indicate where the creation of novel subclusters may be beneficial. Clusters that
               featured genera and/or subclusters that appeared novel were selected for further proteomic analysis using
               Gegenees and VICTOR, as outlined below [Figure 1].


               Proteomic analysis using Gegenees to provide support for the creation of novel genera and
               subclusters
               The MP genomes belonging to the VIRIDIC-assigned genera of interest were entered into Gegenees for
               TBLASTX (i.e., proteome) analysis using the default settings of a 200 bp fragmentation and a 100 bp step-
               size . The fragmented approach is used alongside a multithread BLAST control engine to provide higher
                  [21]
               resolution of alignments . The resulting values (presented in heatmap format) are the average BLAST
                                    [21]
               scores expressed as a percentage of the score each genome would obtain when BLASTed against itself (i.e.,
               100 % identity;  [21,22] ). The resulting heatmaps were expected to somewhat reflect the VIRIDIC heatmaps if
               the existence of the genera and subclusters were supported at a proteomic level [10,21] .
   5   6   7   8   9   10   11   12   13   14   15