CL
Climate: Past, Present & Future

Challenges in accessibility and code quality of the climate models that underpin our understanding of climate change

Challenges in accessibility and code quality of the climate models that underpin our understanding of climate change

Climate models help us understand how the climate system works, project future climate change, and provide evidence that supports international assessments such as exposed on the IPCC. At this point, there is a fundamental question that is rarely discussed and solved: Can we access and reproduce the climate models used throughout the history of the Coupled Model Intercomparision Project (CMIP)?
In our recent study, we explored this question by assessing the accessibility and software quality of climate models that participated in CMIP.

Why reproducibility matters
Reproducibility is one of the foundations of science. In principle, independent researchers should be able to examine the methods used in a study and verify the results. These results must be exactly the same obtained by other researcher groups. For computational science, and especially climate science, this means that source code, configuration files, documentation, and model settings should ideally remain accessible over time. Without them, reproducing simulations becomes extremely difficult, even when the scientific publications themselves are available.  Also, this is very important for educating new generations of climate modelers, as it is essential to have clear, reproducible code and experiments to understand why certain decisions and implementations were made in the past when developing the models.

Climate models are particularly challenging because they are developed over decades by large international teams and contain millions of lines of code. As models evolve, older versions can become difficult to locate, maintain, or even identify.

Looking back across the history of CMIP
To investigate how accessible climate models actually are, we examined models from all historical CMIP phases; from CMIP1 to CMIP6. CMIP7 still under development nowadays. Our approach was focused on search for publically available model code, contact institutions and modelling centres when code was not available, evaluate licensing conditions, assess documentation and usability, and analyse software quality in the models that could be obtained.
Despite the central role of CMIP in climate science, we were able to obtain only a fraction of the models that have participated through its history. Many early models have simply not been preserved, while others remain inaccessible because of licensing restrictions, institutional policies, or uncertainties regarding intellectual property rights. One of the most striking findings was the complete absence of recoverable source code from CMIP1 and CMIP2. These models played an important role in the development of modern climate science, yet many have effectively become part of a lost digital heritage. The scientific publications describing them still exists, but the software itself is often unavailable. Preserving model source code is not only about reproducibility today. It is also about preserving scientific knowledge for future generations.

Are climate models following software engineering best practices?
Access to code is only one part of the reproducibility challenge. Once source code is available, another question emerges: How maintainable and understandable is it?
To explore this issue, we analyse the climate models using FortranAnalyser, a toll specifically designed to assess scientific software written in Fortran (the most widely used programming language in the climate models). Newer models achieved higher software quality scores than older models. The results suggest that climate modelling groups are increasingly adopting better software development practices and paying greater attention to maintainability and reproducibility.

Beyond climate science: preserving reproducible research
Although this study focuses on climate models, its implications extend far beyond the climate science community. Modern research increasingly relies on complex software systems, and when source code is unavailable, poorly documented, or inadequately preserved, reproducibility becomes difficult to achieve regardless of the discipline. The challenges identified in this work are therefore not only technical, but also organisational and cultural. To address these challenges, we argue that final versions of scientific software, including climate models, should be preserved in long-term public repositories together with their documentation, configuration files, and licensing information. Adopting open-source practices and routinely evaluating software quality can strengthen transparency, improve trust in scientific results, and make research more accessible to future generations.
Climate models represent decades of scientific investment, collaboration, and accumulated knowledge. Ensuring that future researchers can understand, inspect, reproduce, and build upon these systems is essential for the continued advancement of climate science and computational research more broadly. After all, preserving scientific knowledge does not end with publishing a paper. Sometimes, it begins with preserving the code behind it.

 

This blog post is based on a manuscript accepted for publication in Geoscientific Model Development.

This post has been edited by the editorial board

Michael García-Rodríguez is a researcher at the Universidade de Vigo (Spain), affiliated with EPhysLab and the School of Computer Science. His work lies at the intersection of climate modelling and software engineering, with a particular interest in reproducibility, code accessibility, and software quality assessment. He develops tools and methodologies to improve transparency and maintainability in scientific software, especially within the Coupled Model Intercomparison Project (CMIP) modelling framework.


Javier Rodeiro-Iglesias is a researcher and lecturer in Computer Science at the Universidade de Vigo (Spain), where he is affiliated with the Department of Computer Science and collaborates with EPhysLab. His research interests include scientific computing, software engineering, cloud computing, sustainability, and the development of digital tools to support scientific research. He works on the application of computational methodologies and software quality practices to improve the reliability, accessibility, and reproducibility of scientific software and environmental research workflows.


Juan A. Añel is Professor of Earth Physics at the Universidade de Vigo (Spain) and a member of the Environmental Physics Laboratory (EPhysLab). His research spans climate change, atmospheric dynamics, renewable energy, computational science, and scientific reproducibility. His work focuses on advancing the transparency, reliability, and societal relevance of climate research through the integration of Earth system science and modern computational approaches.


Leave a Reply

Your email address will not be published. Required fields are marked *

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>

*