Dear EGU Hydrology community,
Over the past few years, machine learning (ML) and artificial intelligence have fundamentally changed what is possible in large-scale hydrological modeling. At Google, our team has worked to develop machine learning river forecast models that operate globally, extending reliable forecasting to ungauged basins and covering almost two billion people through the Google Flood Hub and related services.
From academic research to flood forecasting
Much of the published academic research that supported this ongoing work, especially in the early stages, was encapsulated within the NeuralHydrology open source modeling framework developed and maintained by our colleagues Drs. Frederik Kratzert and Martin Gauch. NeuralHydrology is a powerful and flexible rainfall-runoff modeling package that supports academic research through community extensions including new models, new datasets, and new training algorithms.
Operational flood forecasting is different from academic research, and importantly is traditionally a local effort. Primary reasons why flood forecasting requires local expertise are because traditional hydrology models perform best when calibrated locally, and because of the need for local, real time knowledge about factors like reservoir operation, flood protection infrastructure, basin and station behavior, local model behavior and error patterns, rating curves, and flood wave behaviors. Global forecast models like Google’s Flood Hub and ECMWF’s Global Flood Awareness System can help fill gaps in the ~40% of countries that lack early warning systems, and can augment national systems by filling coverage gaps and by providing high level overviews and third party confirmation, but these systems cannot directly replace local forecast infrastructure or the Indigenous and Local Knowledge (ILK) behind those systems.
Flood forecasting is somewhat different from weather forecasting in this regard. ILK is critical in weather forecasting, of course, but atmospheric fluid dynamics have an inherent planetary scale component while river dynamics do not, at least not when simulated with models that use atmospheric forcings (e.g., precipitation) as input data. Because of this, operational hydrology and flood forecasting agencies are used to working only with local models, instead of starting with planetary scale models at some point upstream in the modeling chain. Practically, this means that agencies with the responsibility and mandate for official flood forecasting typically expect to run their own models. Professional operational flood forecasters are used to calibrating, tuning, and running models, running counterfactuals and what-if scenarios, performing real-time data assimilation, and generally directly interacting with forecast tools.
Consequently, research to operations (R2O) in flood forecasting requires putting interactive models in the hands of experts. This is a related but fundamentally different project than developing and maintaining open source tools (like NeuralHydrology) to support hydrology research.
The Google Flood Forecasting team developed OpenHydroNet to support R2O and operational use cases. OpenHydroNet is a heavily modified fork of the NeuralHydrology Github repository. The main difference is that OpenHydroNet supports real forecasting, including using meteorological forecast data (e.g., quantitative precipitation forecasts). OpenHydroNet is designed specifically to work with the Caravan Multimet dataset. Caravan is, to our knowledge, the largest open, public streamflow dataset in the world and the Multimet extension provides historical training data for more than 20,000 watersheds globally. OpenHydroNets contains model weights pre-trained on this dataset.
Earlier, we said that traditional hydrology models perform best when calibrated locally. This is an important difference between traditional and machine learning models. ML-based models perform best when trained on large-sample datasets. They can be improved with local fine tuning, but best practice is to start with a large, pre-trained model. OpenHydroNet supports fine tuning with local data for local catchments, which we personally recommend in an operational setting. Fine tuning should be done carefully using data or knowledge from local experts, which is why we do not use fine tuning in the global-scale Flood Hub.
OpenHydroNet is public and truly open source in the sense that community contributions are welcome. We have accepted several community contributions to date. However, unlike NeuralHydrology, OpenHydroNet is not intended to support a plethora of models or datasets. The goal is to keep this software package simple and as easy to use as possible. One of the main barriers to developing and operating national flood forecast systems is complexity and resources. The ML models are inexpensive to run, and can be operated by forecasters with less expertise and experience than traditional models. Of course, more domain and professional expertise is always better, but not always possible everywhere in reality.
What’s next?
This modeling package is expanding. We are currently adding capabilities that make training and forecasting tasks easier, for example, through automated training and real-time data extraction for new watersheds. We would love to hear from the operational hydrology community about how to make ML models more accessible.
Edited by B. Schaefli