Page 86 - Read Online
P. 86
Yu et al. Carbon Footprints 2025, 4, 17 https://dx.doi.org/10.20517/cf.2025.12 Page 3 of 23
2. What categories of emission models exist, and how do they account for traffic-related variables?
3. How do different models vary in terms of data requirements, complexity, interpretability, and
transferability?
4. In which traffic management scenarios are these models most effectively applied?
RESEARCH METHODS
Systematic literature review process
A systematic literature review was conducted to comprehensively collect, screen, and critically evaluate
existing research on urban road traffic CO emission models. Unlike narrative reviews, a systematic
2
[10]
literature review adheres to a rigorous, replicable, and verifiable protocol . In this study, the PRISMA
checklist was adopted to guide the review process, ensuring comprehensive coverage of relevant literature
and methodological transparency.
Keyword extraction
Given the extensive volume of retrieved publications, natural language processing (NLP) techniques, such
as RAKE, TF-IDF, TextRank, and YAKE, were applied to automated keyword extraction and topic
classification . In this study, YAKE was selected for its lightweight, unsupervised design, with proven
[11]
strengths in computational efficiency and domain adaptability on small-to-medium-scale corpora.
Benchmark comparisons by Campos et al. showed that YAKE achieved competitive or superior
[12]
performance relative to baseline methods, with F1 scores ranging from 0.086 to 0.500 across eleven datasets,
while requiring less computational overhead. In addition, YAKE also performs generally well for different
domains and types of documents (e.g., agricultural papers, news articles, and scientific literature).
Text classification
To enhance classification accuracy and contextual understanding, advanced language models (e.g.,
ChatGPT and DeepSeek-R1) were employed for downstream text processing tasks . These models have
[13]
consistently outperformed traditional methods across a range of NLP benchmarks. In particular,
DeepSeek-R1 demonstrated strong generalization capabilities, achieving benchmark scores such as 90.8% on
MMLU, 92.2% F1 on DROP, 87.6% on AlpacaEval 2.0, and 71.5% on GPQA Diamond . These results
[14]
support its applicability to complex, high-stakes tasks such as scientific text classification within this review.
RESEARCH PROCESS
To enhance the efficiency and objectivity of the literature selection process, we adopted an AI-assisted
systematic review approach, drawing on the methodology proposed by Noroozi et al. . The approaches
[15]
comprise three main steps: (1) search query formulation, (2) document screening via large language models
(LLMs), and (3) quality assessment. Limiting the results to studies published between 2020 and 2024, the
review focuses on recent methodological advancements in CO emission modeling. To provide historical
2
context, Supplementary Table 1 compares foundational (pre-2020) and contemporary (2020-2024)
approaches in terms of modeling principles, data needs, and application scenarios.
Search query formulation
The core dataset was obtained from the Web of Science Core Collection, which includes over 21,100
peer-reviewed journals, conference papers, and books. An initial search query [Figure 1] was developed
through expert consultation, internal discussions, and AI assistance. This query specifically targeted urban
road traffic studies while excluding publications related to rail, air, and maritime transportation. To ensure

