Page 99 - Read Online
P. 99
Ambros et al. Microbiome Res Rep 2023;2:34 https://dx.doi.org/10.20517/mrr.2023.18 Page 7 of 19
intact L. curvatus prophages and occasionally misinterpreted one intact prophage as two sequences (e.g., as
an intact prophage with an overlapping incomplete one, like in L. curvatus strain TMW 1.2272).
Furthermore, bacterial genome regions with a high density of transposases were sometimes misjudged as
(intact) prophages (e.g., in L. curvatus strain TMW 1.411).
To avoid these oversights, we manually checked the presence of all essential gene modules and determined
attachment sites (att-sites) by aligning the outer regions of each prophage (including a bacterial genome
part of a view thousand base pairs) onto one another. The att-site near the phage integrase was annotated as
“attL”, and near the lysin as “attR”. Mostly, attL and attR are identical, or almost identical, short (10-25 bp)
nucleotide sequences, which indicate the boundaries of the prophage and therefore mark its exact
integration locus. Incomplete prophages often lack genome parts (including att-sites), so assigning an exact
integration locus is difficult and inaccurate. With this in mind, only the integration sites of putatively intact
prophages were analysed [Supplementary Table 3].
Following manual evaluation, 50 putatively intact L. curvatus prophages were found across 33 strains (intact
prophage incidence of 73% in 45 analysed genomes) after manual evaluation. Notably, this includes
transposase containing prophages, described below, and phages with only partially sequenced genomes.
This revealed multiple putatively intact prophages in strains where no intact prophages were previously
detected by PHASTER (e.g., L. curvatus strain ZJUNIT8). The identified prophages start (attL site) with
lysogeny-related genes, followed by genes for replication, packaging, head and tail construction, including
fibre and receptor genes, and lysis genes at the end (attR site). The integration locus within the host
chromosome and the predicted att-sites for each intact predicted prophage are listed in Supplementary
Table 3. General features like length, number of coding regions (CDS), GC content, and number of tRNA
genes present in each prophage (including partial tRNA genes used for integration) are listed in
Supplementary Table 4.
To display the genetic diversity of those prophages, we pre-sorted them based on a neighbour joining tree
(Figure 1 left side; depicted as cladogram), which was constructed using a pairwise comparison table
[Supplementary Figure 1], originating from a whole genome alignment using the phage genomes (attL to
attR). The BLAST analysis revealed higher nucleotide similarities between closer related prophages, mostly
in the replication and tail gene modules (compare Figure 1; grey bars). Notably, the two phage pairs TMW
1.706 P1/DSM 20019 P2 (96.31% nucleotide similarity over 100.00% aligned nucleotides) and TMW 1.706
P2/DSM 20019 P1 (97.69% nucleotide similarity over 100.00% aligned nucleotides) are almost identical. For
this analysis, the split genome parts of prophages TMW 1.706 P1, TMW 1.706 P2 and NRIC0822 P1 have
been joined. A correlation between phylogenetic groups of the bacterial hosts [Supplementary Figure 2] and
the grouping of their respective prophages [Figure 1] could not be detected.
The shortest complete prophage detected was phage WiKim38 P2, with a length of 29.3 kb. Phage DRD-164
P2 was slightly shorter with 29.1 kb, but lacked lysis genes and could therefore be incomplete. Nevertheless,
we included it in our analyses, as it contained all other gene modules and therefore relevant sequence
information. The longest prophage was phage MRS6 P1 with 51.0 kb. The average intact prophage size was
37.8 ± 4.8 kb (median ± interquartile range). The GC content of the detected prophages ranged from 37.9%
to 43.5%. The number of CDS ranged from 40 (phage KG6 P1, and phage TMW 1.595 P1) to 71 (phage
MRS6 P1, and phage TMW 1.1365 P3). Each prophage harboured between 0 and 2 tRNA genes, not
counting partial tRNA-sequences used for integration [Supplementary Table 4].

