Page 41 - Read Online
P. 41
Page 6 of 17 Wang et al. Microbiome Res Rep 2024;3:39 https://dx.doi.org/10.20517/mrr.2024.21
DIA-NN search for identification and quantitation of DIA-PASEF data
[16]
DIA-PASEF data were processed with DIA-NN (v1.8.1) using either spectral libraries generated through
MSFragger or a library-free mode using reduced FASTA databases generated through pFind or MSFragger
searches as described above, with a precursor FDR threshold of 1%. Default settings were used for all the
DIA-NN searches.
MSFragger search for identification and quantitation of DDA-PASEF data
To obtain identification and quantitation results of the DDA-PASEF dataset, the DDA data were processed
with the FragPipe (v21.1) MSFragger node using the LFQ-MBR workflow. The full mouse gut microbial
gene catalog and host database were used. Similar to the spectral library generation workflow, a split
database strategy was used to improve the sensitivity of peptide-spectrum match (PSM) with a database split
factor of 10. MaxLFQ intensity of identified protein groups was used in this study with a minimum ion of 1.
Taxonomic and functional annotation and analysis
Taxonomic and functional annotation and analysis for both DIA-NN and MSFragger outputs were
performed using MetaLab (version 2.3). Briefly, the DIA-NN outputs report_pg.tsv and report_pr.tsv files
[36]
were used for functional and taxonomic annotations, respectively. Similarly, the MSFragger outputs
combined_proteins.tsv and combined_peptides.tsv were used for functional and taxonomic annotations,
respectively. Unipept option in MetaLab was used to generate taxonomic annotations for both database
search results, and a minimum of three distinct peptides was used for confident identification of taxa.
Database search for small proteins and antimicrobial peptides
A small protein database derived from the human gut microbiome was downloaded from the supplemental
data of a previous study by Sberro et al. . We used the protein cluster data table, which contained 444,054
[25]
entries, to generate the small protein database. The AMPsphere AMP database was downloaded on
September 1, 2024, from AMPsphere (https://ampsphere.big-data-biology.org/home), containing 863,498
entries . To enable the calculation of the relative abundance of small protein or antimicrobial peptides
[37]
(AMPs) to total proteins in a sample, all identified peptide sequences from previous search using the
combined database (gut microbial gene catalog and mouse proteome) were concatenated with the AMP or
small protein databases for DIA-NN search or MSFragger search for DIA-PASEF data or DDA-PASEF data,
respectively. DDA and DIA data were quantitatively analyzed using the methods described above with the
respective alternate databases.
Data visualization and statistical analysis
Experimental flowcharts were generated using BioRender (https://www.biorender.com/). Graphs were
generated using R package ggplot2 and ggpubr.
RESULTS
Evaluating bioinformatics workflows for DIA-PASEF metaproteomics
To evaluate fecal sample preparation workflows, a pooled fecal sample from C3H/HeN female mice was
crushed into a powder, homogenized and aliquoted for either microbial enriched protein extraction through
DC or direct protein extraction with NC workflow [Figure 1A]. To assess the potential impacts of protein
digestion methods, both DC and NC protein lysates were subjected to in-solution trypsin digestion
(following acetone precipitation for detergent removal), FASP with 10kDa molecular weight cut-off filter
(FASP10), or 3kDa filter (FASP3). All comparisons were conducted with five replicates with a total of 30
peptide samples for MS analysis on a timsTOF Pro 2 mass spectrometry system using both DDA- and DIA-
PASEF acquisition modes [Figure 1A].

