Page 45 - Read Online
P. 45
Page 10 of 17 Wang et al. Microbiome Res Rep 2024;3:39 https://dx.doi.org/10.20517/mrr.2024.21
study [Supplementary Figure 3]. The most abundant host proteins identified in the fecal samples include
those involved in digestion function and antimicrobial activity, such as regenerating islet-derived protein
3-beta (Reg3β) and immunoglobins, which may bind tightly to the microbe surface and thereby become
enriched with microbial cells.
Next, we evaluated whether different sample preparation methods and MS data acquisition modes could
impact the identification of small proteins that are commonly overlooked in classical metaproteomics.
Among all the identified proteins in this dataset, there were 31 protein groups with ≤ 50 amino acids and
845 with ≤ 100 amino acids. With the application of recent more sensitive mass spectrometers, > 2,000 small
proteins (< 100 amino acids) can be identified in a similar experimental system . We selected a 100 amino
[38]
acids cut-off in this study to facilitate comparison between groups [28,39] . There were 200-600 small proteins
identified per sample, with DIA-PASEF yielding higher identifications than DDA-PASEF mode
[Figure 3A]. Differential centrifugation reduced the number of small protein identifications when the
extracted proteins were digested using in-solution and FASP-10kDa methods, but not for those with FASP-
3kDa method, which may be due to the lower overall performance of FASP-3kDa for NC samples. When
considering the sum relative abundance of those small proteins, the type of digestion, FASP column size and
MS acquisition mode all present minimal impacts [Figure 3B]. We then performed functional and
taxonomic annotation using GhostKOALA for these identified small proteins, which showed that they are
significantly enriched in genetic information processing and lipid metabolism pathways [Figure 3C and D,
Supplementary Figure 4]. In addition, the identified small proteins in both DDA and DIA datasets were
significantly enriched in undefined taxa (adjust P value of 1.59E-68 and 1.08E-45 for DDA and DIA
datasets, respectively) from the microbiomes [Supplementary Figure 5]. This might be due to the fact that
small protein sequences have lower information content for taxonomic annotation and are being under-
represented in current knowledge databases. Altogether, this study showed that the use of differential
centrifugation depleted the abundance of small proteins within sample sets across both DDA and DIA
acquisition modes and protein digestion methods, suggesting that the NC method is superior when
targeting and investigating small proteins specifically. It is worth mentioning that the NC method also has
the advantage of shortened sample preparation steps and time, and thereby reduces the sample-to-sample
variations introduced during sample preparation.
Previous metagenomics data mining identified small open reading frames (sORF) encoding > 4,000 small
proteins (≤ 50 amino acids) in the human microbiome, with the majority having no known function .
[25]
Although the human microbiome may not fully encapsulate the small proteins that would be present in a
mouse microbiome, the overlap that exists can help reinforce the evaluation of the impacts of different
sample preparation methods and MS data acquisition mode on small protein identification. To enable
calculation of the relative abundance and mitigate false discovery, all identified mouse gut microbial peptide
sequences in this study were combined with the predicted small protein sequences for database search for
both DDA and DIA-PASEF data using MSFragger and DIAN-NN, respectively. A consistent and large
increase in small protein identifications in all sample preparation methods was observed when samples were
run using a DIA-PASEF mode than with DDA-PASEF regardless of upstream sample preparation
workflows [Supplementary Figure 6]. NC combined with FASP-3kDa digestion workflow tends to identify
the lowest number of small proteins, which may be due to the overall low protein identifications in this
group; however, the small proteins identified represented the largest percentage in abundance in samples. In
fact, for both in-solution and FASP digestion workflows, NC leads to an increased abundance of small
proteins regardless of protein digestion methods and MS acquisition modes, suggesting that differential
centrifugation depletes small proteins within the sample. Altogether, the findings in this study suggest that
direct fecal lysis for protein extraction, followed by digestion with either in-solution or FASP-10kDa

