Page 9 - Read Online
P. 9
O’Connell et al. Microbiome Res Rep 2023;2:21 https://dx.doi.org/10.20517/mrr.2023.17 Page 3 of 21
screening efforts to identify novel MP; however, the scope has broadened to include phages targeting other
bacteria [18,19] . Notably, the majority of MP identified thus far have been isolated using a single host strain,
2
Mycobacterium smegmatis mc 155, which creates a degree of bias with regard to the types of phages isolated
[20]
and characteristics such as host range and this bias may obscure the true diversity of MP .
In this investigation, VIRIDIC was used to group MP into genera. The subsequent VIRIDIC-defined
taxonomy was compared to the existing taxonomy within the National Centre for Biotechnology
Information (NCBI) database and the International Committee on Taxonomy of Viruses (ICTV) Master
Species List. The potentially novel genera suggested by VIRIDIC were further supported by proteomic
analyses performed using Gegenees and VICTOR. The purpose of Gegenees is to fragment complete
genomes to the default set length and search (using tBLASTx) the appropriate BLAST database for “seeds”
of each fragment against the other genomes. These results can then be used to infer phylogenetic
distances [21,22] . VICTOR visualises phylogeny based on comparisons of genome or proteome sequences to
generate dendrograms extrapolated from the genome-BLAST distance phylogeny method with branch
support . If the novel VIRIDIC predicted genera are supported, it was hypothesised that the Gegenees
[23]
output would mirror the VIRIDIC alignments, although it should be noted that comparing DNA-based to
proteomic-based similarity values is complicated by complex evolutionary patterns, genetic exchange events
and the mosaic nature of MP . Each genus was also anticipated to be represented by a monophyletic
[24]
branch (or clade) within the VICTOR-generated dendrogram.
The VIRIDIC-assigned groups were also compared to the existing cluster/subcluster assignments of the
phages to begin investigating a hypothesis that was generated during the initial curation of the MP into their
respective subclusters. Essentially, this study proposes a link between subclusters and genera. The cluster-
based classification system was initially established to aid the organisation of the outputs from the SEA-
PHAGES program and later broadened into a large public database featuring phages targeting a variety of
hosts, i.e., the Actinobacteriophage Database (https://phagesdb.org/). Originally, cluster assignment
required all members to share 50% nucleotide similarity of their total genomes [18,19] , but now requires that
phages share 35% of their gene content based on a bioinformatic pipeline involving the Phamerator
program, which assigns genes into groups of related sequences [19,25] . This cluster demarcation (≥ 50%
nucleotide similarity) was noted to have been the minimum similarity required for genus assignment until
[12]
recently . Therefore, one would expect each cluster to consist of a single genus. However, if the latest
genus demarcation requires a minimum of 70 % genome similarity [10,12] , it can be hypothesised that more
than one genus may exist within a single cluster using this threshold. As subcluster division within clusters
is largely based on subgroups of genomes having evidently higher nucleotide similarities to each other than
the cluster as a whole, it may be possible for subclusters to reflect the most up-to-date demarcation of genus
(i.e., > 70% similarity; [12,18,20,26] ). Therefore, similar sequential analyses using VIRIDIC, Gegenees, and
VICTOR could likely identify novel subcluster groups. Should this hypothesis prove correct, it could lead to
a formalisation of the criteria for subcluster creation, which is currently arbitrarily based on “recognisable
divisions” within comparisons of average nucleotide identity in each cluster (Hatfull, 2022). It is also widely
noted that subcluster thresholds vary between clusters, including those of phages that infect bacteria other
than mycobacteria . Therefore, the overall purpose of this analysis was to not only identify novel genera,
[27]
but to establish a more consistent method of subcluster assignment amongst MP.
METHODS
Initial selection of MP genomes from established bioinformatic databases
When this research began, 2,096 MP genomes had been fully sequenced, and the majority of this data was
available through Genbank. In order to generate a more manageable dataset and to create a “snapshot” of

