Page 43 - Read Online
P. 43

Page 8 of 17                  Wang et al. Microbiome Res Rep 2024;3:39  https://dx.doi.org/10.20517/mrr.2024.21

               conducted to see which pipeline for DIA data processing yields the best identification results. Briefly,
               protein sequences were extracted from both search outputs to generate reduced FASTA databases for DIA-
               NN search with library-free mode (pFind-LibF for using pFind-generated protein database, and MSF-LibF
               for using MSFragger-generated database). Spectral libraries were generated by MSFragger with either the
               full database or a pFind-generated reduced protein database, and were used for DIA data analysis (MSF-Lib
               and pFind-Lib, respectively). As shown in Figure 1D, all DIA-NN search inputs yielded similar protein
               group numbers, falling between 12,647 and 14,001. The main difference in identification was found in the
               peptide level, where all workflows tested resulted in approximately 58,000-60,000 peptides except for MSF-
               LibF which identified 72,979 peptides in total for the DIA-PASEF dataset. Based on these evaluations, we
               then chose MSF-LibF workflow for DIA-PASEF data processing, and to be consistent, the MSFragger LFQ-
               MBR quantification using the full database (with 10 splits) was used for the analysis of the DDA-PASEF
               dataset.

               DIA-PASEF metaproteomics achieved better identification and quantification
               We first compared the DDA-PASEF and DIA-PASEF data acquisition modes in terms of peptide and
               protein identification, as well as protein quantification in fecal metaproteomics. Across all fecal sample
               preparation methods, there was a consistent increase in peptides and protein groups identified per sample
               when samples were run using DIA-PASEF mass spectrometry compared to those run with DDA-PASEF
               mode [Figure 2A and B]. This finding is in agreement with a previous study by Gómez-Varela et al. . For
                                                                                                    [21]
               the identification of peptides and protein groups, the digestion method using a FASP-10kDa approach had a
               slight edge over other digestion methods tested in DC preparations, and in-solution digestion had a slight
               edge in NC samples [Figure 2A and B]. DIA-PASEF runs for DC samples achieved the highest number of
               precursor identifications with nearly 50,000 precursors per sample (identifications ranging from 36,912 to
               49,928 precursors, 35,324 to 46,665 peptides, 9,598 to 10,705 protein groups), while nearly 21,000 peptides
               (identifications ranging from 14,029 to 20,852 peptides, to 5,068 to 6,891 protein groups) were identified for
               the same samples with DDA-PASEF runs. DIA-PASEF runs for NC samples achieved a competitive number
               of precursor identifications with nearly 45,000 precursors per sample (identifications ranging from 17,681 to
               44,915 precursors, 17,457 to 42,646 peptides, 7,002 to 11,031 protein groups), while up to almost 20,000
               peptides (identifications ranging from 6,334 to 19,389 peptides, 3,292 to 7,492 protein groups) were
               identified for the same samples with DDA-PASEF runs. The filter-assisted sample preparation, using the
               FASP-3kDa columns, showed comparable identification rates for DC samples compared to other digestion
               methods, but it led to a decrease in peptide and protein group identifications when applied to NC samples
               [Figure 2A and B]. The identified peptides in the NC with the FASP-3kDa group exhibited smaller average
               peptide lengths and fewer missed cleavage sites compared to the other groups [Supplementary Figure 1].
               These differences might be due to the potential undesirable interactions between the FASP column matrix
               and non-protein components in NC samples and the resulting low efficiency for eluting large peptides.

               The impacts of MS data acquisition and sample preparation on quantification were then assessed with the
               amount of missing values across samples in each group. In Figure 2C, Q0 indicates a protein group with
               quantified intensity in at least 1/5 replicates in the group, Q50 is present in at least 3/5 replicates, and Q100
               is present in all replicates. Along with an increased number of protein group identifications in DIA-PASEF
               data across all sample preparation methods, there is a distinctly higher proportion of protein groups
               quantified in all sample replicates (Q100) for DIA-PASEF data compared with DDA-PASEF, in particular
               for samples prepared with in-solution digestion and FASP-10K column (73%-80% in DIA vs. 49%-53% in
               DDA; Figure 2C). Evaluation of the intra-group sample-wise Pearson’s correlation of the quantified protein
               intensities indicates high quantitative reproducibility in both DDA and DIA datasets for all proteins
               (Pearson’s r of 0.89-0.96 and 0.90-0.97, respectively) and for low abundant small proteins (100 amino acid
               length threshold; Pearson’s r of 0.82-0.96, 0.84-0.97, respectively) [Supplementary Figure 2]. Therefore,
   38   39   40   41   42   43   44   45   46   47   48