Page 40 - Read Online
P. 40
Page 6 of 19 Klaassens et al. Microbiome Res Rep 2024;3:38 https://dx.doi.org/10.20517/mrr.2024.13
[27]
We built the Kraken2 V2.1.1 database as described in the Kraken reference manual provided online
(https://ccb.jhu.edu/software/kraken/MANUAL.html). Inspecting the database after building showed that
99.98% of the minimizers were mapped to the lowest common ancestor (LCA) (see graphical abstract for a
graphical overview of the database).
For the bioinformatics analysis pipeline, we used Kraken 2, together with Bracken. We did not use the
standard settings for several steps in these tools, but optimized them for our specific question here. The
adaptions were set as follows: for Kraken: Confidence: 0.05; For Bracken: minimum reads (-t) 100; read
length (-r) 150; level (-l) D, P, C, O, F ,G ,S, S1.
The number of taxa is smaller than the number of added genomes because many taxa have multiple
[36]
genomes representing them. A Bracken V2.5 database was constructed based on the finished Kraken2
database with a read length of 150 bp.
Taxonomic profiling pipeline
[37]
Nextflow was used to write the infant taxonomic profiling pipeline. A schematic overview of the entire
workflow can be found in the graphical abstract. The computational part of the workflow combines the
following parts: Read in the paired-end reads in FASTQ format; Paired-end read classification down to the
LCA with Kraken2 using a specific infant gut microbiome database; Abundance profiling on different
taxonomic levels (from Domain to Strain) using Bracken. Bracken is designed to redistribute reads classified
by Kraken2. By default, Bracken only redistributes down to the species level. However, it is possible to
redistribute all the species-level reads down to the subspecies and strain levels. Redistribution at the strain
level is performed based on the assigned reads and available kmers per strain by Kraken2 using Bracken
[36]
(bayesian estimation of abundance with Kraken) .
Bioinformatics analysis of metagenomic sequencing data
After sequencing and demultiplexing the samples, the FASTQ files were analyzed in the following way: The
bioinformatics pipeline described above was run using a Kraken2 confidence of 0.1 and a Bracken threshold
of 100 reads.
The relative abundance was adjusted for the number of reads classified by Kraken2 and accepted by
Bracken. The relative read abundances were then used to analyze the composition of every sample at the
genus, species and strain levels. Figures were created within R3.6.0 [R Core Team (2021) using vegan
(version 2.5-5)] package and ggplot2 (v3.1.1) . Permanova statistics (adonis, vegan package) was
[38]
[39]
performed to investigate whether the centroids and dispersion of the two groups are equivalent.
Statistical analysis
To assess whether differences between treatment effects in terms of the investigated endpoints were
statistically relevant, paired two-sided t-tests within a given donor group [(1) vaginal, (2) C-section born,
and (3) all donors] were performed. To control the proportion of false discoveries when conducting a high
number of comparisons, the Benjamini-Hochberg false discovery rate (FDR) was applied. Differences
between treatment effects were considered significant when the obtained P-value (obtained through the
paired two-sided t-test) was smaller than a reference value (ref). This reference value was obtained by
ranking of the obtained P-values in ascending order within the donor group. The rank of a given P-value
was termed (i) and varied between 1 and the total amount of P-values (m = 9, as mentioned below). The
reference value was calculated by multiplying the FDR with the rank of the P-value, divided by the total
amount of comparisons made (ref = FDR*i/m). To compare the differences between treatment effects in
terms of changes in pH levels, gas pressures, and microbial metabolite productions (SCFA, lactate and

