Page 55 - Read Online
P. 55
Wang et al. Carbon Footprints 2024;3:14 https://dx.doi.org/10.20517/cf.2024.19 Page 7 of 19
Table 2. The proxy data and related data used in this study and their sources.
Data Name Data source Description
Population counts chn_ppp_2020 WorldPop Hub [43] GeoTIFF, 100 m
[44]
Nighttime light data VNL_v2_npp_2020 VIIRS Nighttime Light GeoTIFF, 15 arc second (~500 m at the Equator)
[45]
Road networks osm_roads OpenStreetMap Shapefile
[46]
Road traffic volume - Code for Design of Urban Road Engineering Average daily traffic for roads at all levels
[47]
Technical Standard of Highway Engineering
Waterways osm_waterways OpenStreetMap [45] Shapefile
[48]
Ship tracking intelligence - MarineTraffic Global ship real-time position information
Land cover MCD12Q1.A2020 NASA MODIS production [49] Hierarchical data format 4 file, 500 m
(1) Spatial allocation of point source emissions
For point source emissions, the spatial allocation process involved several key steps, as shown in Figure 4A.
Latitude and longitude information was directly used to pinpoint the location of the emission source and
allocate the emissions to the corresponding target grids. However, during the spatial allocation process,
challenges arose with several large industrial facilities that occupy extensive areas, which made it difficult to
accurately allocate specific emission points. Therefore, further processing was required for such emissions
from large-scale and broad-footprint factories. To address this, the top 500 point sources mentioned above
were manually screened to ensure accurate location identification for high-emission sources.
Given that most large enterprises had multiple emission sources, it was complex to determine the exact
location of each emission source. Therefore, it was assumed that emissions were uniformly distributed
within the factory area. Based on this assumption, for factories spanning multiple grids, the process began
with identifying point source locations on the map using latitude and longitude coordinates, followed by
delineating enterprise boundaries with satellite imagery, as shown in Step 1 in Figure 4A. Then, spatial data
overlay analysis was used to disaggregate the enterprise across the corresponding grids, determining the
location and quantity of grids in which the enterprise was located. Finally, emissions were evenly allocated
across these intersecting grids, as shown in Step 2 in Figure 4A. If there were multiple point sources within a
single grid, their emissions would be summed up. This method effectively solved the problem of spatial
allocation posed by multiple point sources in large enterprises, ensuring the accurate distribution of
emissions to target grids.
(2) Spatial allocation of non-point source emissions
Non-point source emissions were allocated using spatial proxy data, including population distribution, road
traffic density, and agricultural land distribution as weight factors. To ensure effective spatial allocation of
these emissions, the proxy data needed to be preprocessed.
For instance, this study focused on preprocessing population distribution data, which served as a proxy for
resident emissions. The WorldPop population data was utilized initially; however, it displayed noticeable
inconsistencies, particularly in city centers, where the data showed jagged distribution patterns that did not
accurately reflect the actual population distribution. To address this issue, nighttime light data, which could
partially reflect population distribution, was employed as a correction variable. The preprocessing involved
setting the population data as the dependent variable and nighttime light data as the independent variable.
A random forest model was then applied to fit these variables and develop an appropriate model for
correction. The random forest model yielded a coefficient of determination (R ) of 0.8665, indicating a
2

