Page 24 - Read Online
P. 24

Page 18 of 21              O’Connell et al. Microbiome Res Rep 2023;2:21  https://dx.doi.org/10.20517/mrr.2023.17

                                      [12]
               recently ratified taxonomy . Typically this was supported by the Gegenees output as the application of a ~
               50% proteome similarity threshold clarified the boundaries between genera, which was the case of cluster J.
               For this cluster, a nucleotide similarity threshold of ≥ 66% was noted to be the minimum similarity value
               required for genus inclusion when the Gegenees and VICTOR outputs were considered in parallel. While
               the decision to change the threshold may seem arbitrary, the analyses presented in this study demonstrate
               that ≥ 50% proteome similarity and monophyletic groups are also strong indicators of genus groups and can
               be considered robust validations of genus assignments. It would therefore be recommended that this
               threshold of ~ 50% proteome similarity be applied to Gegenees analyses as additional supporting evidence
               for VIRIDIC predictions, which will allow researchers the opportunity to identify the most robust groups in
               circumstances where many nucleotide similarity values are approaching but not equal to 70%. Similarly, the
               VIRIDIC/Gegenees proposed genera should demonstrate a monophyletic nature (which proved especially
               important for clarifying the boundaries between the proposed novel genera within cluster K) [Figures 7 and
               8].


               As the taxonomic information regarding phages now reflects greater diversity than previously thought, it is
               decided to also explore the Actinobacteriophage database cluster and subcluster assignments to determine if
               they reflected the diversity described by the ICTV/NCBI taxonomy. It is pertinent that this classification
               system is reliable, as there have been studies designed with the understanding that MP belonging to the
               same subcluster may have similar attributes like codon usage bias or host range (e.g.,   [30,31] ) and
               approximated groups could lead to hypotheses that falsely include or exclude phages. Such hypotheses may
               therefore mischaracterise the molecular and/or phenotypic functionality of MP, which could have a knock-
               on effect on the design of phage-based therapies and diagnostics. While it is largely accepted that the
               current MP clusters are assigned based on the most convenient groups, as opposed to genetic or
               evolutionary accuracy [6,18,20] , importantly, it has been previously noted that organisation of MP into
                                                                                                       [18]
               appropriate clusters/subclusters has become difficult as the number of available sequences increases .
               Regarding subcluster classifications, there is currently no formal demarcation of “subcluster” and subcluster
               assignments are based on “recognisable divisions” in nucleotide similarity, as addressed in the
               Introduction . It was noted during the initial curation of taxonomic and cluster information for the 721
                          [18]
               MP genomes that there was a potential link between genus and subcluster assignments (Figure 1; the genus-
               subcluster hypothesis), i.e., that each subcluster comprised a single genus. This link appeared sound
               following the comparison of the existing taxonomy and the novel genera proposed in this study.

               In some cases, like in cluster A, novel genera were assigned to a single subcluster, which further supports the
               suggestion that a single genus can be assigned to each subcluster. In other cases, like cluster J which
               currently lacks subclusters, the creation of genera appeared to warrant the creation of subclusters based on
               the “recognisable divisions” within the nucleotide alignments  that were subsequently supported by the
                                                                    [19]
               Gegenees and VICTOR analyses. Notably, clusters A and K featured inconsistent alignment of subclusters
               and VIRIDIC highlighted distinct groups that would suggest smaller subclusters could be formed from
               larger ones. The Gegenees analyses broadly paralleled the VIRIDIC outputs and supported reorganising the
               subclusters into smaller groups; however, it was quite difficult to clearly and distinctly define the groups
               within cluster K as described above and more consideration was given to the VICTOR output in that case.
               Interestingly, the Gegenees threshold (broadly speaking, cluster K appears to be an exception) for subcluster
               classifications appeared to be ~ 50% proteome similarity, which is the threshold proposed for genus
               identification within Gegenees outputs presented in this study. This creates additional reassurance in the
               genus-subcluster relationship hypothesis, as the threshold for genus and subcluster recognition is the same.
               Finally, the VICTOR-generated phylogeny illustrated all proposed subclusters as monophyletic, even the
               minority of those comprised of more than one genus. In total, 13 novel subclusters were proposed and the
   19   20   21   22   23   24   25   26   27   28   29