HS
Hydrological Sciences

What machine learning can and cannot tell us about floods

Collage based on a picture by Imad Clicks via Pexel

Floods are shaped by complex and nonlinear interactions between weather patterns, rainfall, soil properties, topography, land cover, and how wet a catchment already is. This complexity is one reason machine learning (ML) is increasingly used in hydrology. ML models can learn patterns from large and complex datasets with many variables. 

Early in my PhD, while exploring machine learning models for estimating flood magnitudes, I became interested in feature importance. I hoped it could help reveal which flood drivers mattered most. However, I quickly noticed that feature importance was highly sensitive, and the most “important” features could change when I altered the model setup, added new predictors, changed the spatial scale, or used a different interpretation method. 

This led to a broader question:

Once a machine learning model performs well, what exactly has it learned? 

This question motivated our recent study, published in Hydrology and Earth System Sciences (Ford et al., 2026). We developed an interpretable machine learning framework to estimate winter flood magnitudes across near-natural UK catchments. Here, interpretability means understanding how a trained model uses different predictors to make its estimates. Our aim was not only to build predictive models, but also to examine how different predictors contributed to modelled flood magnitudes, and how those interpretations changed depending on the information available to the model and the spatial scale of the analysis. 

We used Random Forest regression models and applied SHapley Additive exPlanations, or SHAP, to investigate how the models used different predictors. SHAP is an explainable artificial intelligence method that breaks down a model prediction into contributions from each input variable. In simple terms, it asks how much each predictor pushed an individual model estimate up or down. This allowed us to examine how the model used different types of information

Opening the black box: a feature incorporation framework

However, when we talk about feature importance, we must remember we are not talking about the true physical importance of a flood-generating process. Feature importance describes how useful a variable is to a particular model, trained on a particular dataset, under a particular set of assumptions. A predictor may represent a hydrological process directly, act as a proxy for something not measured, or reflect shortcut learning by the model. 

A key contribution of our study is the feature incorporation framework. Rather than placing all predictors into one model and interpreting a final ranked list of feature importance, we progressively added groups of predictors. These included spatial identifiers, weather pattern information, catchment characteristics, and precipitation variables. This allowed us to examine how model performance and interpretation changed as new information was introduced. 

This sequential approach helps answer a more nuanced question: which predictors remain influential once overlapping information is already available? Interpretation becomes a process of comparison rather than a single final ranking. 

Our findings show that there is no simple answer to which variables matter most for flood magnitude modelling. The answer depends on the data included, the model structure, the interpretation method, and the spatial scale of the experiment. 

Static versus dynamic predictors  

Our framework also distinguishes between dynamic event-scale predictors and static catchment characteristics. Dynamic predictors, such as event rainfall, describe conditions directly linked to a flood event. Static predictors, such as aridity, baseflow index, or catchment location, help explain why the same rainfall can produce different responses in different places. 

Static predictors are often powerful, but they require careful interpretation. Some, such as baseflow index, summarise meaningful hydrological behaviour. Others, such as latitude and longitude, may help the model encode broad regional patterns without representing direct physical controls on flooding. This does not make them useless, but it does mean their importance should not be overinterpreted as physical causation. 

Main takeaway  

The main message from our work is that interpretable machine learning can support flood research, but only when used carefully. Methods such as SHAP can help reveal how a trained model uses information, while our feature incorporation framework helps test how those interpretations change as new predictor groups are added. Together, these methods provide insights into model behaviour, not a definitive ranking of real world flood drivers.  

Feature importance should therefore be understood as evidence of model behaviour under a specific setup, not as a direct explanation of the real world. Used critically, interpretable machine learning can help hydrologists ask better questions about floods, models, and the limits of what data-driven methods can tell us. 

 

Note from the editorial team: This is a guest blogger contribution received following a recent innovation in the Copernicus Publishing System: upon acceptance of your paper in one of our journals, authors are invited to consider if they would like to turn their science paper into a blog post.

 Bibliography 

Emma Ford., Manuela I Brunner., Hannah Christensen., and Louise Slater; Interpretable feature incorporation machine-learning framework for flood magnitude estimation, Hydrol. Earth Syst. Sci., 30, 2135–2160, https://doi.org/10.5194/hess-30-2135-2026, 2026. 

Emma Ford is a final year PhD candidate at the University of Oxford, studying machine learning for large sample hydrology. Her research uses a data driven lens to investigate flood generating processes, while also examining the tools themselves, including machine learning and explainable artificial intelligence, and what they can and cannot tell us about flood systems.


Leave a Reply

Your email address will not be published. Required fields are marked *

You may use these HTML tags and attributes: <a href="" title=""> <abbr title=""> <acronym title=""> <b> <blockquote cite=""> <cite> <code> <del datetime=""> <em> <i> <q cite=""> <s> <strike> <strong>

*