Page 45 - Read Online
P. 45

Page 10 of 17                 Wang et al. Microbiome Res Rep 2024;3:39  https://dx.doi.org/10.20517/mrr.2024.21

               study [Supplementary Figure 3]. The most abundant host proteins identified in the fecal samples include
               those involved in digestion function and antimicrobial activity, such as regenerating islet-derived protein
               3-beta (Reg3β) and immunoglobins, which may bind tightly to the microbe surface and thereby become
               enriched with microbial cells.


               Next, we evaluated whether different sample preparation methods and MS data acquisition modes could
               impact the identification of small proteins that are commonly overlooked in classical metaproteomics.
               Among all the identified proteins in this dataset, there were 31 protein groups with ≤ 50 amino acids and
               845 with ≤ 100 amino acids. With the application of recent more sensitive mass spectrometers, > 2,000 small
               proteins (< 100 amino acids) can be identified in a similar experimental system . We selected a 100 amino
                                                                                  [38]
               acids cut-off in this study to facilitate comparison between groups [28,39] . There were 200-600 small proteins
               identified per sample, with DIA-PASEF yielding higher identifications than DDA-PASEF mode
               [Figure 3A]. Differential centrifugation reduced the number of small protein identifications when the
               extracted proteins were digested using in-solution and FASP-10kDa methods, but not for those with FASP-
               3kDa method, which may be due to the lower overall performance of FASP-3kDa for NC samples. When
               considering the sum relative abundance of those small proteins, the type of digestion, FASP column size and
               MS acquisition mode all present minimal impacts [Figure 3B]. We then performed functional and
               taxonomic annotation using GhostKOALA for these identified small proteins, which showed that they are
               significantly enriched in genetic information processing and lipid metabolism pathways [Figure 3C and D,
               Supplementary Figure 4]. In addition, the identified small proteins in both DDA and DIA datasets were
               significantly enriched in undefined taxa (adjust P value of 1.59E-68 and 1.08E-45 for DDA and DIA
               datasets, respectively) from the microbiomes [Supplementary Figure 5]. This might be due to the fact that
               small protein sequences have lower information content for taxonomic annotation and are being under-
               represented in current knowledge databases. Altogether, this study showed that the use of differential
               centrifugation depleted the abundance of small proteins within sample sets across both DDA and DIA
               acquisition modes and protein digestion methods, suggesting that the NC method is superior when
               targeting and investigating small proteins specifically. It is worth mentioning that the NC method also has
               the advantage of shortened sample preparation steps and time, and thereby reduces the sample-to-sample
               variations introduced during sample preparation.


               Previous metagenomics data mining identified small open reading frames (sORF) encoding > 4,000 small
               proteins (≤ 50 amino acids) in the human microbiome, with the majority having no known function .
                                                                                                        [25]
               Although the human microbiome may not fully encapsulate the small proteins that would be present in a
               mouse microbiome, the overlap that exists can help reinforce the evaluation of the impacts of different
               sample preparation methods and MS data acquisition mode on small protein identification. To enable
               calculation of the relative abundance and mitigate false discovery, all identified mouse gut microbial peptide
               sequences in this study were combined with the predicted small protein sequences for database search for
               both DDA and DIA-PASEF data using MSFragger and DIAN-NN, respectively. A consistent and large
               increase in small protein identifications in all sample preparation methods was observed when samples were
               run using a DIA-PASEF mode than with DDA-PASEF regardless of upstream sample preparation
               workflows [Supplementary Figure 6]. NC combined with FASP-3kDa digestion workflow tends to identify
               the lowest number of small proteins, which may be due to the overall low protein identifications in this
               group; however, the small proteins identified represented the largest percentage in abundance in samples. In
               fact, for both in-solution and FASP digestion workflows, NC leads to an increased abundance of small
               proteins regardless of protein digestion methods and MS acquisition modes, suggesting that differential
               centrifugation depletes small proteins within the sample. Altogether, the findings in this study suggest that
               direct fecal lysis for protein extraction, followed by digestion with either in-solution or FASP-10kDa
   40   41   42   43   44   45   46   47   48   49   50