Page 12 - Read Online
P. 12

Page 6 of 21               O’Connell et al. Microbiome Res Rep 2023;2:21  https://dx.doi.org/10.20517/mrr.2023.17

               Data presentation
               The heatmaps generated by VIRIDIC and Gegenees were created and/or edited using Microsoft Office Excel
               365 (version 2203, Microsoft Corporation, Redmond, Washington, USA). “Conditional formatting” was
               used to introduce a colour scale to the VIRIDIC outputs, ranging from red (lowest score) to green (highest
               score). This type of conditional formatting is automatically applied by Gegenees when the software
               produces the alignments, so the heatmap outputs were simply exported as .html files and copied into Excel
               for formatting purposes. Specific ranges of similarity scores for each figure are indicated in the figure
               legends. At least one established genus was included in each analysis for comparative purposes, i.e., a
               “positive control” for how a genus should look in each output. Novel genera are indicated in bold black
               squares. The phylogenetic trees produced by VICTOR and ViPTree were digitally captured and edited in
               Microsoft PowerPoint (version 2203, Microsoft Corporation, Redmond, Washington, USA). The
               monophyletic branches representing genus groups are indicated with blue boxes and the proposed
               nomenclature is entered into each box. The subcluster groups are indicated in coloured boxes, while
               proposed subclusters are highlighted in red squares.

               RESULTS
               The initial selection of MP genomes from established bioinformatic databases
               As outlined in the Materials and Methods, initial shortlisting of the Genbank genomes involved limiting the
               scope of the study to those included in the RefSeq database, which totalled 752 MP. The MP genomes were
               then organised into their respective clusters to identify additional phage that could be removed from this
               stage of the analysis based on whether the cluster appeared to be underrepresented, as described in the
               Materials and Methods [Table 1].


               There were three phages that were not yet entered into the Actinobacteriophage Database, and their
               associated publications did not indicate their cluster or taxonomy, so these MP were not included in the
               initial analyses. Singleton phage and clusters represented by a single RefSeq genome (clusters Q, U, V, X,
               AA, and AD) were also not included due to their apparent underrepresentation in the dataset (> three
               genomes), as explained in the Materials and Methods. For this same reason, seven other clusters were
               removed - R, S, T, Y, Z, AB, and AC [Table 1]. BLASTN results indicated that the inter-cluster similarity of
               the removed clusters to those collated for the analysis is minimal and bears little to no influence on the
               results presented. By eliminating these clusters, the resulting dataset comprised 721 MP genomes grouped
               within 16 clusters (A-P), each represented by at least five genomes.


               Identification of novel genera and subcluster assignments by comparison of the current NCBI/ICTV
               taxonomy and Actinobacteriophage database subclusters with the VIRIDIC outputs
               Following the taxonomic analysis summarised in Figure 1 in the Materials and Methods, it was determined
               that there were 20 potentially novel genera within clusters A, J and K, while the analysis of cluster G
               suggested that a single genus (and its respective subcluster) classification may not be strictly necessary.
               Regarding the proposed genus-subcluster hypothesis, it was noted that 83.3% of subclusters included in the
               dataset support the proposed relationship between genus and subcluster proposed in the Introduction. It is
               likely, based on the identification of 20 novel genera, that a greater percentage of agreement could be
               obtained if some of the remaining 16.7% of subclusters were reorganised into smaller subclusters to reflect
               the proposed genera while changing as few of the existing subclusters as possible to avoid confusion. In
               total, thirteen novel subcluster assignments were deemed robust enough based on the VIRIDIC, Gegenees
               and VICTOR analyses to be proposed in the following sections, and their creation increases the support for
               the genus-subcluster hypothesis from 83.3% to 97.6% when the 20 novel genera are also considered. The
               data regarding proposed changes to MP classifications are presented below in alphabetical order based on
               cluster. In each case, the VIRIDIC results are presented, followed by the Gegenees analyses and the
   7   8   9   10   11   12   13   14   15   16   17