Page 55 - Read Online
P. 55

Wang et al. Carbon Footprints 2024;3:14  https://dx.doi.org/10.20517/cf.2024.19  Page 7 of 19

               Table 2. The proxy data and related data used in this study and their sources.
                Data            Name         Data source                   Description
                Population counts  chn_ppp_2020  WorldPop Hub [43]         GeoTIFF, 100 m
                                                           [44]
                Nighttime light data  VNL_v2_npp_2020 VIIRS Nighttime Light  GeoTIFF, 15 arc second (~500 m at the Equator)
                                                       [45]
                Road networks   osm_roads    OpenStreetMap                 Shapefile
                                                                        [46]
                Road traffic volume  -       Code for Design of Urban Road Engineering    Average daily traffic for roads at all levels
                                                                        [47]
                                             Technical Standard of Highway Engineering
                Waterways       osm_waterways  OpenStreetMap [45]          Shapefile
                                                      [48]
                Ship tracking intelligence -  MarineTraffic                Global ship real-time position information
                Land cover      MCD12Q1.A2020  NASA MODIS production [49]  Hierarchical data format 4 file, 500 m

               (1) Spatial allocation of point source emissions

               For point source emissions, the spatial allocation process involved several key steps, as shown in Figure 4A.
               Latitude and longitude information was directly used to pinpoint the location of the emission source and
               allocate the emissions to the corresponding target grids. However, during the spatial allocation process,
               challenges arose with several large industrial facilities that occupy extensive areas, which made it difficult to
               accurately allocate specific emission points. Therefore, further processing was required for such emissions
               from large-scale and broad-footprint factories. To address this, the top 500 point sources mentioned above
               were manually screened to ensure accurate location identification for high-emission sources.


               Given that most large enterprises had multiple emission sources, it was complex to determine the exact
               location of each emission source. Therefore, it was assumed that emissions were uniformly distributed
               within the factory area. Based on this assumption, for factories spanning multiple grids, the process began
               with identifying point source locations on the map using latitude and longitude coordinates, followed by
               delineating enterprise boundaries with satellite imagery, as shown in Step 1 in Figure 4A. Then, spatial data
               overlay analysis was used to disaggregate the enterprise across the corresponding grids, determining the
               location and quantity of grids in which the enterprise was located. Finally, emissions were evenly allocated
               across these intersecting grids, as shown in Step 2 in Figure 4A. If there were multiple point sources within a
               single grid, their emissions would be summed up. This method effectively solved the problem of spatial
               allocation posed by multiple point sources in large enterprises, ensuring the accurate distribution of
               emissions to target grids.

               (2) Spatial allocation of non-point source emissions


               Non-point source emissions were allocated using spatial proxy data, including population distribution, road
               traffic density, and agricultural land distribution as weight factors. To ensure effective spatial allocation of
               these emissions, the proxy data needed to be preprocessed.


               For instance, this study focused on preprocessing population distribution data, which served as a proxy for
               resident emissions. The WorldPop population data was utilized initially; however, it displayed noticeable
               inconsistencies, particularly in city centers, where the data showed jagged distribution patterns that did not
               accurately reflect the actual population distribution. To address this issue, nighttime light data, which could
               partially reflect population distribution, was employed as a correction variable. The preprocessing involved
               setting the population data as the dependent variable and nighttime light data as the independent variable.
               A random forest model was then applied to fit these variables and develop an appropriate model for
               correction. The random forest model yielded a coefficient of determination (R ) of 0.8665, indicating a
                                                                                    2
   50   51   52   53   54   55   56   57   58   59   60