CSO statistical release, , 11am
Machine learning (ML) is becoming more widely adopted among producers of official statistics, as it offers the possibility of producing reliable, up-to-date statistics with a high degree of accuracy, efficiency, and reproducibility. In some cases, ML allows us to more effectively extract data from novel data sources, such as images or unstructured documents.
ML is a fundamental tool used to produce land cover maps from satellite imagery as it is not feasible to annotate images manually, especially when the imagery covers an entire country.
ML was a key tool used in the production of the NLCM itself, whereas more advanced ML algorithms are used to produce the CLMS land cover products, including its next generation land cover product CLCplus Backbone.
ML was applied in this instance to classify each piece of land in Ireland as either peatland or non-peatland to enable mapping and accounting of ecosystem type (ET) 7.2 Mires, Bogs & Fens. This was done to provide for reporting of ecosystem extent accounts to Eurostat under EU Regulation 691/2011.
This classification was applied to update the extent of the peatland land cover class from the 2018 National Land Cover Map (NLCM), which has direct correspondence to ET7.2. This was done for two reasons:
The NLCM has an infrequent update cycle which does not satisfy the three-year reporting cycle of ecosystem extent accounts.
Although a range of high-resolution and frequently updated pan-European land cover products are provided by the Corine Land Monitoring Service (CLMS), they do not include a land cover class which corresponds to Irish peatlands.
Sentinel satellite data required to update the peatland land cover class every three years is updated every five days and is accessed through the Copernicus Data Space Ecosystem (CDSE). The CDSE is the European Commission's official platform for accessing Sentinel satellite data. To map peatlands from Sentinel satellite data, it is necessary to use ML.
We used a ‘random forest’ (RF) model to infer, with high probability, areas of peatland from Sentinel-2 satellite imagery. Essentially, the imagery is a spatially referenced collection of pixels, each of which is 10 × 10 m in size and contains intensity values for several wavelengths (or bands) of light. The RF algorithm is given a sample of pixels with known labels (i.e. peatland and non-peatland) and uses them to create a model which describes what peatland and non-peatland pixels should look like in the Sentinel-2 imagery. The model is essentially a series of digital decision trees which vote on whether each pixel is peatland or not peatland. If the training process has been effective, the model will be able to classify every pixel for the whole country with a high level of accuracy.
Data security: There is a low risk level associated with the data used for this process, as the input data is publicly available.
Accuracy: Accuracy was considered in relation to model relevance (i.e. suitability of selected predictor variables), model accuracy, and accuracy of the predicted map.
Accuracy was assessed and ensured throughout the process in the following ways:
Data quality – the data used to train and map peatlands was the atmospherically-corrected Sentinel-2 Level-2A imagery, which has already been pre-processed to a high, analysis-ready standard by the European Space Agency.
Training data – the model was trained primarily from the NLCM, which is comprehensive, accurate, very-high-resolution and regarded as providing the best national mapping of peatlands by various national experts. The pipeline also includes an NLCM update step, which is intended to prevent model training using outdated sample points (for example, where an area of peatland in 2018 has since been converted to forestry).
Model validation – this was tested using a robust spatial-cross validation method (blockCV). This demonstrated that model accuracy was consistent across different splits of training and testing datasets.
Data post-processing – the predicted map was passed through a range of cleaning steps. This primarily served to remove areas of peatland which were classified by the ML model but with a low probability. Note, two key steps at this stage included (1) the ‘adding back’ of areas mapped as peatland in the NLCM and (2) removal of newly predicted peatland areas which were not predicted by the model with a high probability. Both steps mean that only around 0.5% of the final updated peatland area for 2018 was mapped using ML.
Consultation with national experts – the CSO met with a range of national peatland ecologists and remote sensing experts to showcase the mapping pipeline and obtain feedback. The final methodology is the product of multiple, iterative revisions with each round of that feedback.
| Statistic | Value | What does it tell us? |
|---|---|---|
| Precision | 0.887 | 11.3% of identified peatland pixels are not peatland |
| Recall | 0.938 | 6.3% of peatland pixels were missed |
| F1 score1 | 0.912 | Harmonic mean of precision and recall |
| 1An F1 value of more than 0.81 indicates almost perfect agreement for categorical data (Landis and Koch, 1977. Biometrics, 33, 1), whereas F1 greater than 0.85 is considered particularly good for remote sensing (Foody, 2008. International Journal of Remote Sensing, 29, 11) | ||
The limitations relate to the sources of the observed error (Table 1). These sources include:
Training data: The ML model is trained using areas which have already been identified using ML in the 2018 NLCM. However, these areas are subjected to a quality check whereby administrative data or newer land cover data (e.g. from the Corine Land Monitoring Service) is used to update the training data, thereby reducing the risk that the ML model is trained with outdated information.
The above limitations should also be understood with consideration of the following:
The final updated peatland maps are not subjected to in-situ field visits as part of their accuracy assessments. Instead, they are checked and validated against available high resolution satellite imagery. This approach is that which is best available to the CSO and was recommended by national experts in the absence of more granular, field-based ground truth datasets.
The peatland maps have a resolution of 10 × 10 m. They are subsequently used as an input for classifying Mires, Bogs & Fens for ecosystem extent mapping to produce CSO Ecosystem Extent Accounts. The ecosystem extent maps use a 1-hectare (100 × 100 m) resolution, which functions, in part, to reduce errors introduced from training and source data. The ecosystem extent maps are also subjected to their own classification accuracy assessment.
The CSO is committed to using innovation and technology to enhance the quality, efficiency, and value of the statistics and insights we provide. As set out in the CSO Statement of Strategy, innovation is essential to meeting evolving user needs, improving services, reducing burden on respondents, and ensuring the continued delivery of high-quality official statistics.
Trust is the foundation of the CSO’s work. As the producer of official statistics in Ireland, the CSO recognises that public confidence depends on maintaining the highest standards of quality, objectivity, transparency, and data protection.
The ML method described above is a basic type of Artificial Intelligence (AI). The CSO’s use of machine learning or other AI tools is guided by our core principles and the Government’s commitment to embrace the full potential of emerging technology to deliver better public services and deliver better outcomes. We also strive to utilise new and secondary data sources for improved insights and reduce the need to collect new primary data. The CSO is also aligned with the guidelines for the Responsible Use of AI in the Public Service, which emphasise innovation, human oversight, public trust, and the protection of people’s rights.
The use of AI within the CSO is overseen by an internal AI Working Group, which considers proposals for AI adoption and ensures that any use of AI is appropriate, responsible, and supports the CSO’s mission to provide insight on Ireland, its people, society, economy, and environment. Human expertise and verification remain central to all statistical processes, with AI serving as a tool to assist and enhance, rather than replace, professional judgement.
Where AI is used in the production of statistical outputs, the CSO will be transparent with users about how AI has been applied. Through this balanced approach, the CSO aims to harness the benefits of AI to improve insight and service delivery while maintaining the trust, transparency, and quality standards that underpin our work.
Our methods might change, but our commitment to confidentiality, accuracy, quality, transparency, and independence remains the same.
Learn about our data and confidentiality safeguards, and the steps we take to produce statistics that can be trusted by all.