Statistical methodsObservational constraint

    This study relies on the Kriging for Climate Change (KCC) statistical method, first introduced to constrain Global mean Surface Temperature (GST)22, and subsequently extended to tackle regional to local temperatures23,24. This method seeks to estimate the forced warming, both in the past (1850 to present) and the future (present to 2100, for a given emission scenario), through an observational constraint approach that combines model data and observations. Here we provide a summary of the key elements of this method, but refer to the corresponding publications22,24 for a comprehensive description.

    The KCC method is basically a Bayesian estimation method, and as such involves the formulation of an a priori distribution of the forced response. It works in three steps, and is described here for a given emission scenario.

    First, the forced response of a range of climate models is estimated for the period 1850–2100. This calculation is based on the response to natural forcings, as estimated using a simple climate model58, as well as a smoothing procedure for the remaining human-induced forced response.

    Second, the sample of the forced responses from available CMIP6 climate models is used to formulate a Gaussian prior of the real-world forced response, under the assumption that “models are statistically indistinguishable from the truth”. Denoting the time-series of the real-world forced response over 1850–2100 as x (i.e., x is a vector of length 251), the prior distribution for x is:

    $${{\rm{\pi }}}({{\bf{x}}})={{\rm{N}}}(\,{{\boldsymbol{\mu }}}_{{\rm{m}}},{{\boldsymbol{\Sigma }}}_{{\rm{m}}})$$

    (1)

    where N stands for the multivariate Gaussian distribution, and μm and Σm are taken as the empirical mean and covariance of the CMIP6 models’ forced response. As a result, this prior includes physical information from climate models and is representative of model uncertainty (Σm).

    Third, observations are used to derive a posterior distribution of the past and future forced response given observations, in a Bayesian way. We assume that

    $${{\bf{y}}}={{\bf{Hx}}}+{{\rm{\varepsilon }}},\ {{\rm{with}}}\ {{\rm{\varepsilon }}}\sim {{\rm{N}}}({{\bf{0}}},{{\boldsymbol{\Sigma }}}_{{\rm{y}}})$$

    (2)

    where y denotes observations (a vector from, e.g., 1850 to 2023), x is still the forced response from 1850 TO 2100, H denotes an observational operator (not all years in x have a corresponding observation in y), ε denotes random noise (a vector) corresponding to internal variability and measurement errors, and Σy denotes its covariance matrix. In this way, the raw observations y are considered to be observations of the forced response, subject to some uncertainty (mainly arising from internal variability and measurement errors). Given these assumptions, the posterior can be derived as

    $${{\rm{p}}}({{\bf{x}}}|{{\bf{y}}}) \sim {{\rm{N}}}(\,{{\boldsymbol{\mu }}}_{{\rm{p}}},{{\boldsymbol{\Sigma }}}_{{\rm{p}}})$$

    (3)

    with the posterior mean μp and the posterior covariance matrix Σp available in closed forms.

    As a result, this method provides observationally constrained estimates of past, present and future forced warming (i.e., the entire vector x), including projections.

    The analyses presented here are identical to those previously published for GST22 (except for the larger set of CMIP6 models used), and France23. This approach differs from the Global Warming Index (GWI)7 in terms of both the physical models used (CMIP models here vs a 2-box Energy Balance Model), and the statistical procedure used to fit observations (Bayesian inference here vs linear regression as used in optimal fingerprinting).

    It is important to notice that, despite the use of recent observations, KCC (like other observational constraint methods) does not produce “initialised” projections. It is not intended to account for or predict a specific state of internal variability, such as ENSO variability. Observations are only used to improve the estimation of the forced response.

    Method evaluation

    The reliability of the KCC method has been evaluated in several ways.

    First, for GST only, a perfect model framework was used to evaluate the method for future projections22. This evaluation suggested that the method was not overly confident, as the coverage probability (i.e., the frequency with which the estimated constrained range contains the true value of interest) was estimated to be 91% in 2100 over a sample of 66 CMIP6 simulations. This is very consistent with the expected nominal value of 90%. The evaluation also illustrated the fact that models with a particularly high (eg, hot models) or low warming were correctly predicted. This result was obtained using all CMIP6 models to derive the prior, and so suggests that the presence of hot models does not bias the outputs of our method.

    Second, a similar perfect model framework was used to assess the reliability of the method for regional temperature24. Again, the reported coverage probabilities suggested the method was reliable and not overly confident.

    Third, the specific calculations performed in the current study provide additional evidence regarding the reliability of KCC in estimating both the forced warming to date (see Figs. S1 and S3 for the global and regional scales, respectively) and the projected forced warming (see Figs. S4 and S5, again for the global and regional scales, respectively). These figures show a perfect model evaluation: one simulation (historical and SSP2-4.5) from one CMIP6 model is used as a pseudo-observation, and the KCC method is applied to it. They illustrate how our estimates change over time as new pseudo-observed years become available. The estimated range is then compared to the best possible estimate of the particular model’s forced response – derived using all available members (ensemble mean), the pseudo-observations for the entire 1850–2100 period, and a smoothing method. Consistent with previous evaluations, the results show that the KCC method is able to correctly estimate both the current and future forced responses, as the “true” value lies within the estimated confidence range in most cases. This finding applies to a broad range of CMIP6 models, including some with a relatively low or high sensitivity, and some with relatively low or high responses to the aerosol forcing. Additionally, this finding was made using all CMIP6 models in the prior, including the hottest ones.

    Finally, these various lines of evidence show that KCC is a reliable method to estimate the present and future forced response, provided that key assumptions are satisfied, in particular “models are statistically indistinguishable from the truth”, meaning that all models are not consistently biased.

    Bias and variance of the warming to date estimators

    The statistical properties (bias, standard deviation, root mean square error) of the 1-year and 10-year estimators of the warming to date are shown in Tables S1 and S2. This simple calculation is based on the analysis of the perfect model framework and on all the model runs analysed in Fig. 2a and S1. First, we derive a best estimate of the forced response in each CMIP model. This is done by averaging over all available members (the ensemble size varies across models, but all of the models used have 7 members or more of historical simulations and SSP2-4.5 scenarios), and then applying a filtering procedure22 in which the anthropogenic response is smoothed over time to eliminate most of the remaining internal variability. This best-estimate is then considered as the true forced response, and is compared to the estimated warming to date using the pseudo-observations from one specific member (global scale, as illustrated in Fig. 2a and S1) or 4 members (regional scale, as illustrated in Fig S3 with one single member; the number of members used to calculate the bias and variance is increased in order to ensure the robustness of the results, despite the lower signal-to-noise ratio at the regional scale). The difference between these two quantities is considered as the error in the warming to date estimator. From the sample of errors over the period 2015–2100, we can easily derive indicators such as bias, standard deviation, and RMSE.

    Observed data

    The observed global temperature data are taken from the infilled version of the HadCRUT535 dataset; only the annual and global means are used. This dataset covers the period 1850-2023. HadCRUT5 measurement errors are estimated using the set of 200 members provided by this dataset, assuming Gaussian error. Following recent studies11,12, we assume that the Global Surface Temperature (GST, mixing surface atmospheric temperature over land and sea surface temperature over oceans) warming assessed in this dataset is representative of GSAT (surface atmosphere temperature everywhere) changes, although this is still a matter of debate.

    In Fig. 2c shows the results of a retrospective calculation based on the data available for each year. We use various earlier versions of this dataset:

    • From 2006 to 2012, we used the HadCRUT3v dataset, which is a statistically adjusted version of HadCRUT336. This dataset combines land (CRUTEM3) and ocean (HadSST2) data at a 5 × 5° resolution since 1850.

    • From 2012 to 2014, the HadCRUT4 dataset34. HadCRUT4 combines land surface temperature data from CRUTEM4 and sea surface temperature data from HadSST3. HadCRUT4 measurement errors are estimated from a set of 100 equiprobable realisations.

    • From 2014 to 2020, a new version of HadCRUT4 was used, based on a better reconstruction and fusion of land and ocean surface temperatures. This version is referred to as HadCRUT4-CW37. Similar to HadCRUT4, this dataset provides an ensemble of 100 equiprobable members to assess measurement uncertainty.

    • From 2020 onwards, the HadCRUT5 dataset, i.e., the current version, was used. As this version was published in 2021, it is used here from the year 2020 onwards. This is based on the assumption that the analysis is carried out in the year following the last year of observations used, similar to the approach of Forster et al. (2023).

    The results shown in Fig. 2c suggest that revisions of these observed datasets all resulted in small breaks in the estimated warming. Over the last 20 years, all revisions have been upwards.

    The observed temperature data for France comes from Météo-France23. These are monthly temperatures over mainland France since 1899, derived as the average of observations from 30 homogenised stations evenly distributed across the country.

    Model data

    All analyses but Fig. 2c are based on a sample of 45 CMIP6 models40—see the detailed list below. To improve the estimation of the forced response, we use historical and SSP2-4.5 scenario simulations and take the average over all available members. For Fig. 2c, we additionally use the previous CMIP338 and CMIP539 generations. In each case, we use intermediate emission scenarios to extend historical simulations (consistent with the use of SSP2-4.5 in CMIP6): A1B for CMIP3, and RCP4.5 for CMIP5. Results in Fig. 2c suggest that the transitions from CMIP5 to CMIP6 was very smooth in terms of estimating the current warming, unlike the previous transition (from CMIP3 to CMIP5).

    The processing of model data involves the following steps. The mean temperature of each model is calculated by considering the annual mean surface air temperature (SAT). For global analyses, we compute the Global mean SAT (GSAT) and assume that it is equal to GST. To illustrate regional analyses, we use temperature data over France, which are interpolated onto a 0.25° x 0.25° grid covering mainland France, and then averaged over continental grid-points (consistent with Ribes et al., 2022). At both the global and regional scales, the subsequent processing of the model data is consistent with that described in the reference papers describing the statistical method22,23, which involves in particular a smoothing procedure to reduce internal variability.

    The following CMIP6 models (SSP2-4.5 scenario) were used for GST analyses (Figs. 1, 2, 3a,b, S1 and S4; 45 models): ACCESS-CM2, ACCESS-ESM1-5, AWI-CM-1-1-MR, BCC-CSM2-MR, CAMS-CSM1-0, CAS-ESM2-0, CESM2, CESM2-WACCM, CIESM, CMCC-CM2-SR5, CMCC-ESM2, CNRM-CM6-1, CNRM-CM6-1-HR, CNRM-ESM2-1, CanESM5, CanESM5-CanOE, EC-Earth3, EC-Earth3-CC, EC-Earth3-Veg, EC-Earth3-Veg-LR, FGOALS-f3-L, FGOALS-g3, FIO-ESM-2-0, GFDL-CM4, GFDL-ESM4, GISS-E2-1-G, GISS-E2-1-H, HadGEM3-GC31-LL, IITM-ESM, INM-CM4-8, INM-CM5-0, IPSL-CM6A-LR, KACE-1-0-G, KIOST-ESM, MCM-UA-1-0, MIROC-ES2L, MIROC6, MPI-ESM1-2-HR, MPI-ESM1-2-LR, MRI-ESM2-0, NESM3, NorESM2-LM, NorESM2-MM, TaiESM1, UKESM1-0-LL.

    The following CMIP6 models (SSP2-4.5 scenario) were used for regional analyses (Fig. 3c, d, S2, S3 and S5; 27 models; same models as in Ribes et al23.): ACCESS-CM2, ACCESS-ESM1-5, AWI-CM-1-1-MR, CAMS-CSM1-0, CanESM5-CanOE, CanESM5, CESM2, CESM2-WACCM, CMCC-CM2-SR5, CNRM-CM6-1-HR, CNRM-CM6-1, CNRM-ESM2-1, EC-Earth3-Veg, FGOALS-f3-L, FGOALS-g3, GISS-E2-1-G, INM-CM4-8, IPSL-CM6A-LR, MIROC6, MIROC-ES2L, MPI-ESM1-2-HR, MPI-ESM1-2-LR, MRI-ESM2-0, NorESM2-LM, NorESM2-MM, TaiESM1, UKESM1-0-LL.

    The following CMIP5 models were used (Fig. 2c; RCP4.5 scenario; 31 models): BNU-ESM, CESM1-CAM5, FGOALS-s2, MRI-CGCM3, MIROC-ESM, MIROC5, MIROC-ESM-CHEM, GISS-E2-R, GISS-E2-H, MPI-ESM-MR, MPI-ESM-LR, CanESM2, CNRM-CM5, EC-EARTH, FIO-ESM, bcc-csm1-1-m, bcc-csm1-1, NorESM1-ME, NorESM1-M, CSIRO-Mk3-6-0, CCSM4, IPSL-CM5A-LR, IPSL-CM5A-MR, ACCESS1-3, ACCESS1-0, CESM1-BGC, GISS-E2-R-CC, GISS-E2-H-CC, inmcm4, CMCC-CMS, IPSL-CM5B-LR.

    The following CMIP3 models were used (Fig. 2c; AIB scenario; 23 models): BCCR, CCCMA0, CCCMA, CNRM, CSIRO0, CSIRO, GISSA, HCM, INM, IPSL, MRI, CCSM, ECHO, FGOALS, GFDL0, GFDL, GISSH, GISSR, HGEM, INGV, MIROC0, MIROC, MPI, PCM.

    Share.

    Comments are closed.