Skip to content
Artwork for Earthly Machine Learning
ScienceEarth Sciences

Earthly Machine Learning

Amirpasha

“Earthly Machine Learning (EML)” offers AI-generated insights into cutting-edge machine learning research in weather and climate sciences. Powered by Google NotebookLM, each episode distils the essence of a standout paper, helping you decide if it’s worth a deeper look. Stay updated on the ML innovations shaping our understanding of Earth.
It may contain hallucinations.

Play
  • 23 episodes
  • weekly
  • Avg 16 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • S2 · E14
    Sunday · 20 min

    Machine learning is revolutionizing weather forecasting – the next step is a change in how we work

    Citation: Dueben, P., Bauer, P., Fuhrer, O., Koldunov, N., & Kristiansen, J. (2026). Machine learning is revolutionizing weather forecasting – the next step is a change in how we work. arXiv preprint arXiv:2606.25076v1. Key Takeaways: A Shift in the Forecasting Value Chain: While machine learning has rapidly achieved competitive skill in weather predictions, the next critical phase is a complete evolution of working practices and operating models. This transition will fundamentally reshape how models are coded, how observational data is exploited, and how forecasts are verified and turned into public services. The Rise of Agentic AI and Automated Workflows: The traditional, slow-paced manual approach to Earth-system model development is giving way to AI-assisted and agentic workflows. Large Language Models (LLMs) and software agents will increasingly handle the writing, testing, optimizing, and porting of code, shifting the role of human scientists from active programmers to supervisors, test designers, and monitors. Modernized Software and Open Data Stewardship: To benefit from rapid AI innovation, meteorological centres must adopt industry-standard software frameworks (like Python, PyTorch, and JAX) and modular code structures. This runs parallel to a major shift in data stewardship, moving toward highly compressed, cloud-native, open-access datasets that can be efficiently searched and streamed by both humans and AI agents. On-the-Fly Generative Emulation: Emerging generative machine learning foundation models will enable interactive, real-time "what-if" simulations and climate scenario exploration. Instead of moving or storing massive datasets, models can recreate specific atmospheric states on demand, though this introduces a critical need for rigorous verification techniques to distinguish physical realism from AI hallucinations. Organizational Risks and Preserving Expertise: Adapting to mixed human-AI environments presents risks, including the potential erosion of expert scientific knowledge if critical workflows are fully offloaded to AI. Weather and climate centres must proactively design transparent model diagnostics, continuous testing, and educational practices to maintain human operational understanding.

  • S2 · E13
    September 13 · 19 min

    CORDEX-ML-Bench: A Benchmark for Data-Driven Regional Climate Downscaling—Experiment Design and Overview

    Citation: Rampal, N., González-Abad, J., Addison, H., Baño-Medina, J., Bettolli, M. L., Blasone, V., Booth, B., Coppola, E., Di Gioia, S., Oldham-Dorrington, J., Doury, A., Engelbrecht, F., Fuentes-Franco, R., Gibson, P. B., Glawion, L., Hardy, C., Ivanov, M., Lee, H. K., Legasa, M. N., Olmo, M., Orr, A., Polz, J., Rogers, M. S. J., Schillinger, M., Sharma, S., Soares, P. M. M., Sobolowski, S., Steinkopf, J., Tang, W., Tian, J.-B., Tomé, R., Wang, K.-C., Wang, Y.-C., Watson, P. A. G., Wetherell, T., Widmann, M., & Gutiérrez, J. M. (2026). CORDEX-ML-Bench: A Benchmark for Data-Driven Regional Climate Downscaling—Experiment Design and Overview. WCRP-CORDEX Machine Learning Task Team. Key Takeaways: Establishing a Standardized Global Benchmark: The paper introduces CORDEX-ML-Bench, the first coordinated multi-domain, multi-architecture benchmarking framework explicitly designed to standardize machine learning (ML) models for regional climate downscaling. The framework enables researchers to evaluate and compare models consistently using open-source datasets and metrics, focusing on a 20× spatial resolution increase (from ~200 km down to ~10 km grids) for daily precipitation and maximum temperature. The initial benchmark targets three highly diverse geographical regions: the European Alps, New Zealand, and Southern Africa. The Critical Extrapolation Gap of Historical Training: A foundational finding of the benchmark is that ML models trained strictly on historical climate data systematically underestimate future climate change signals, including extreme warming and intensive precipitation. This reveals a major vulnerability in traditional historical-only statistical downscaling methods and proves that incorporating future climate projections into training datasets (the emulator approach) is vital for producing physically credible long-term projections. Generative AI Outperforms for Precipitation Extremes: In an evaluation of 40 independently developed ML configurations, generative approaches (such as diffusion models, flow matching, and Generative Adversarial Networks) consistently outperform deterministic regression models at downscaling precipitation. They excel at capturing highly localized spatial variability and heavy-tailed extreme events, whereas deterministic models suffer from spatial "oversmoothing". However, deterministic architectures remain highly competitive for predicting daily maximum temperature. Balancing Accuracy and Computational Cost: While highly complex diffusion models (like the top-ranked RCMGEM-mv-orog) achieve outstanding accuracy, they are computationally intensive. In contrast, flow-matching and GAN-based models achieve a highly favorable skill-to-compute ratio, lowering inference costs by one to two orders of magnitude. This makes them highly practical solutions for generating large, multi-model ensemble climate projections under restricted computational budgets.

  • S2 · E12
    September 6 · 19 min

    AIMIP Phase 1: Systematic Evaluations of AI Weather and Climate Models

    Citation: Henn, B., Bretherton, C. S., Kodunov, N., Lessig, C., Molina, M. J., Arcomano, T., Watt-Meyer, O., Couairon, G., Singh, R., Brunstein, R., Hasson, Y., Jost, A., Brenowitz, N., Manshausen, P., Cresswell-Clay, N., Durran, D., Hall, K. J. C., Yuval, J., Kochkov, D., Hoyer, S., & Lopez-Gomez, I. (2026). AIMIP Phase 1: systematic evaluations of AI weather and climate models. arXiv preprint. A New Benchmarking Era for AI Climate Models: AIMIP Phase 1 establishes the first systematic intercomparison framework for artificial intelligence weather and climate models (AIWCMs). It defines a common experimental protocol, standardizes CMIP-compatible output formats, and provides an open dataset to evaluate how different AI architectures influence long-term climate simulation behaviors. Standardized Historical Simulation Protocol: Under the Phase 1 protocol, participating models are trained exclusively on historical ERA5 atmospheric reanalysis data from 1979 to 2014 and run through a 10-year out-of-sample test period (2015–2024). To prevent overfitting, models are forced only by specified sea surface temperatures (SST) and sea ice concentrations (SIC), with direct greenhouse gas inputs (like CO2 concentrations) strictly excluded. Excellent Representation of Baseline Climate and ENSO: The initial evaluations of the eight participating AI models demonstrate that they represent time-mean climate averages and natural variability patterns—such as the El Niño-Southern Oscillation (ENSO)—just as well as, or in some cases with lower systematic biases than, conventional physically-based climate models like the NOAA GFDL-CM4. The Out-of-Sample Warming Gap: A primary weakness identified across several AI models is their struggle to accurately replicate global warming trends during the out-of-sample test period (2015–2024). Because greenhouse gases like CO2 are omitted as direct predictors to avoid overfitting, some models fail to translate rising ocean temperatures into the full magnitude of observed atmospheric warming. Extreme Extrapolation Remains a Challenge: When subjected to extreme, highly out-of-sample sensitivity experiments where sea surface temperatures are uniformly raised by +2 K and +4 K, the AI models diverge significantly. They produce highly inconsistent and sometimes physically implausible responses (such as simulated cooling over land), highlighting that projecting unseen future climates remains a key development challenge for the AI climate modeling community.

  • S2 · E11
    August 30 · 20 min

    WV-Net: A Foundation Model for SAR Ocean Satellite Imagery

    Citation: Glaser, Y., Stopa, J. E., Wolniewicz, L. M., Foster, R., Vandemark, D., Mouche, A., Chapron, B., & Sadowski, P. (2025). WV-Net: A Foundation Model for SAR Ocean Satellite Imagery. Artificial Intelligence for the Earth Systems, e250003. DOI: 10.1175/AIES-D-25-0003.1 Key Takeaways First Foundation Model for Open-Ocean SAR Imagery: WV-Net represents the first-ever foundation model designed specifically for open-ocean sea surface images, utilizing a massive dataset of nearly 10 million unannotated C-band synthetic aperture radar (SAR) wave mode images collected globally by the Sentinel-1 satellite mission. Overcoming the Annotation Bottleneck: By leveraging contrastive self-supervised learning (SimCLR), the model learns highly robust, general-purpose representations of complex geophysical signatures directly from raw, unlabeled imagery—bypassing the traditional bottleneck of expensive manual expert annotation. Outperforming General-Purpose Computer Vision Models: The model's specialized ocean-domain embeddings consistently beat standard models pretrained on natural images (like ImageNet) across key downstream tasks, including estimating wave height, predicting air-sea temperature differences, and identifying 12 distinct atmospheric and oceanic phenomena. High Data Efficiency & Fine-Tuning Stability: WV-Net scales exceptionally well in data-constrained settings, delivering strong performance with as few as 100 labeled training examples. Additionally, it exhibits greater robustness to hyperparameter selections during fine-tuning, dramatically reducing the need for broad, computationally heavy optimization sweeps. Optimizing Domain-Specific Augmentations: The researchers discovered that standard computer vision augmentations adapted for radar data (such as mixup, color inversions, rotations, and sharpness adjustments) were crucial to bridging the domain gap, while complex, domain-specific signal filtering (such as random notch filtering) actually degraded model performance.

  • S2 · E10
    August 27 · 17 min

    Toward Skillful Forecasting of Super El Niño Events Using a Diffusion-Based Westerly Wind Burst Parameterization

    Citation: Ji, C., Mu, M., Qin, B., Lian, T., Yuan, S., Feng, J., Song, S., Wei, Y., Dai, G., Wang, J., & Fang, X. (2025). Toward skillful forecasting of super El Niño events using a diffusion-based westerly wind burst parameterization. npj Climate and Atmospheric Science (Published in partnership with CECCR at King Abdulaziz University). https://doi.org/10.1038/s41612-025-01158-x Key Takeaways: Innovative Generative AI Parameterization: The study introduces a state-of-the-art Denoising Diffusion Probabilistic Model (DDPM) to parameterize westerly wind bursts (WWBs). These wind bursts are critical, episodic atmospheric events that inject wind energy into the Pacific, playing a pivotal role in triggering super El Niños. This new generative AI framework successfully captures the complex, joint modulation of wind bursts by both slow-varying oceanic states and rapid atmospheric processes. Superior Representation of Wind Burst Physics: Traditional schemes rely heavily on ocean-state indicators like the warm pool eastern edge, which fails to capture high-frequency atmospheric noise. By incorporating multiple physical conditions—Sea Surface Temperature Anomalies (SSTA), Outgoing Longwave Radiation Anomalies (OLRA), and Sea Level Pressure Anomalies (SLPA)—the DDPM-based scheme dramatically improves the simulated frequency, intensity, duration, and spatial distribution of wind bursts compared to observational data. Drastic Improvements in Super El Niño Intensity Predictions: When coupled online with the Community Earth System Model (CESM), the DDPM scheme significantly outperforms both standard control runs and traditional parameterization schemes. It accurately predicts the absolute amplitude of historic super El Niño events—specifically the 1982/83, 1997/98, and 2015/16 events—by correcting the severe underestimations found in baseline climate models. Mitigation of Seasonal Phase-Locking Bias: A persistent challenge in climate modeling is "seasonal phase-locking" prediction bias, where models incorrectly project a double-peak warming cycle (peaking in summer, weakening, then re-intensifying in winter). The DDPM scheme overcomes this issue by generating stronger and more realistically eastward-shifted wind stress anomalies, which correctly trigger the positive dynamical feedbacks (such as the Bjerknes feedback) necessary to sustain a steady, natural warming progression toward a single December peak.

  • S2 · E9
    May 9 · 19 min

    Aligning artificial intelligence with climate change mitigation

    Citation: Kaack, L. H., Donti, P. L., Strubell, E., Kamiya, G., Creutzig, F., & Rolnick, D. (2022). Aligning artificial intelligence with climate change mitigation. Nature Climate Change, 12, 518–527. https://doi.org/10.1038/s41558-022-01377-7 Main Takeaways: Three Layers of AI's Climate Footprint: The authors propose a framework that splits machine learning's climate impact into three distinct categories — the energy and hardware emissions of computing itself, the immediate effects of specific ML applications, and the broader system-level changes that ML induces across society. The categories that are easiest to measure (like the electricity used to train a model) are likely not the ones with the largest effects, which is why most current discussions of "AI and climate" capture only a sliver of the real picture. Computing Is a Small Slice — For Now: The entire global ICT sector accounts for roughly 1.4% of global greenhouse gas emissions, and AI workloads are only a fraction of that. But the trajectory is steep: at Facebook, ML training compute has been growing about 150% per year and inference compute about 105% per year, far outpacing efficiency gains. Even striking efficiency wins — like Google's TPU being 30–80 times more energy-efficient than contemporary CPUs or GPUs — can be swamped by raw growth in demand. The "Internet of Cows" Problem: ML is a general-purpose tool, which means it's just as good at accelerating oil and gas exploration or scaling up cattle farming (an industry already responsible for about 9% of global emissions) as it is at forecasting solar power or optimizing data center cooling. Whether AI is net-positive or net-negative for the climate is genuinely undetermined, and depends on which applications get funded, deployed, and regulated. System-Level Effects May Dwarf Everything Else: The largest climate impacts of AI may come not from training runs or even individual applications, but from how ML reshapes society — through rebound effects (efficiency gains that drive more consumption), technological lock-in (autonomous cars entrenching private vehicle travel over transit and rail), and ML-powered recommender systems that boost demand for emissions-intensive goods. These effects are the hardest to quantify but potentially the most consequential, and the authors argue they need to be built into climate scenario modeling — something the IEA, EIA, and IPCC's Shared Socioeconomic Pathways largely don't do today.

  • S2 · E8
    May 3 · 19 min

    Machine learning for the physics of climate

    Machine learning for the physics of climate Citation: Bracco, A., Brajard, J., Dijkstra, H. A., Hassanzadeh, P., Lessig, C., & Monteleoni, C. (2025). Machine learning for the physics of climate. Nature Reviews Physics, 7, 6–20. https://doi.org/10.1038/s42254-024-00776-3 Main Takeaways: Breaking the El Niño Spring Barrier: For decades, forecasts of the El Niño Southern Oscillation hit a hard wall at roughly 6 months lead time — a limit known as the spring predictability barrier. Convolutional neural networks trained on a mix of climate model and reanalysis data have shattered this ceiling, delivering skillful forecasts at 17 months out, with newer architectures pushing to 21–24 months. ML models can also now anticipate which type of El Niño will develop (eastern vs. central Pacific), which matters enormously because the two flavors produce very different regional impacts around the world. Weather Forecasting at a Fraction of the Cost: A new generation of ML weather emulators — Pangu-Weather, GraphCast, FourCastNet, FuXi, NeuralGCM — now match or beat the European Centre's flagship physics-based forecasting system on most variables, including hurricane tracks, while running orders of magnitude faster. They achieve this with surprisingly compressed state representations: roughly 10 vertical atmospheric levels and 0.25° horizontal resolution, compared to 100+ levels and 0.1° in conventional models. The catch is that these models can violate basic physics — geostrophic balance, energy conservation, the butterfly effect — which currently blocks naive extension to climate timescales. Hybrid Models Are Eating the Climate Stack: Pure ML works for short-range forecasts, but for climate-length runs the field is converging on hybrid architectures that pair a traditional dynamical core with neural-network parameterizations of sub-grid processes like clouds, turbulence, and gravity waves. Google's NeuralGCM exemplifies the approach and already reduces biases in tropical cyclone frequency and tracks. A telling case study on the quasi-biennial oscillation showed that an offline-trained neural network produced unstable, unphysical results — but retraining just two layers online, coupled to the model, recovered the correct physics. Offline-only or online-only training each fail in characteristic ways; the mix is what works. The Data Wall Is the Real Bottleneck: Climate ML has less than 50 years of dense satellite-era observations to work with, and those observations are heavily biased toward the atmosphere and ocean surface — a single, spatiotemporally correlated realization of one climate. This limits how confidently ML models can extrapolate to warmer, unseen climates, which is exactly what climate projection requires. The path forward involves three parallel bets: hybrid physics-ML models that bake in conservation laws, large-scale "foundation models" for weather and climate trained across simulations and observations together (efforts like ClimaX and AtmoRep are early examples), and rare-event sampling strategies to handle the extremes that matter most for adaptation policy but are by definition underrepresented in any training set.

  • S2 · E7
    April 27 · 20 min

    Atmospheric Transport Modeling of CO2 With Neural Networks

    Citation: Benson, V., Bastos, A., Reimers, C., Winkler, A. J., Yang, F., & Reichstein, M. (2025). Atmospheric transport modeling of CO2 with neural networks. Journal of Advances in Modeling Earth Systems, 17, e2024MS004655. https://doi.org/10.1029/2024MS004655 Main Takeaways: A New Benchmark for AI Carbon Tracking: The authors introduce CarbonBench, the first systematic benchmark dataset designed specifically for training and evaluating machine learning emulators of Eulerian atmospheric transport. Built from CarbonTracker CT2022 inversions and ObsPack station observations, it ships at three resolutions (the coarsest being 5.625° × 10 vertical levels × 6h) and is engineered to plug directly into modern deep learning pipelines — opening atmospheric carbon modeling to the broader ML community. SwinTransformer Wins, Decisively: Of the four architectures tested (UNet, GraphCast, SFNO, and SwinTransformer), the SwinTransformer reaches near-perfect emulation with a 90-day R² above 0.99 and stays stable in physically plausible forward runs for over three years — a regime where neural PDE solvers typically blow up. At measurement stations, it actually captures the seasonal cycle in Svalbard better than TM5, the conventional model it was trained to emulate, possibly due to differences in boundary layer transport near the poles. Physics Tricks Were the Unlock: Out of the box, the neural networks were unstable — especially the mesh-based UNet and GraphCast. Two simple physics-aware adjustments fixed this across all four architectures: centering the CO2 input field at each timestep to remove the covariate shift from steadily rising atmospheric CO2 (called CentFlux), and a post-hoc mass fixer that rescales predicted mass to match the surface flux budget. The result is mass conservation with RMSE of just 0.00058 PgC against a total atmospheric carbon mass of ~865 PgC — effectively negligible. Speed Isn't the Selling Point (Yet): Unlike AI weather models, which famously outpace numerical forecasting by orders of magnitude, the SwinTransformer is not significantly faster than TM5 at this resolution — about 1.5 seconds for a 30-day run on an A40 GPU versus a few minutes for TM5 on 24 CPUs. The real promise lies elsewhere: the networks are fully differentiable (useful for inverse modeling of surface fluxes), natively support batched ensembles, and scale better to high resolution where conventional solvers become prohibitively expensive — exactly the regime where current CO2 inversions struggle most.

  • S2 · E6
    April 20 · 17 min

    On the foundations of Earth foundation models

    Citation: Zhu, X. X., Xiong, Z., Wang, Y., Stewart, A. J., Heidler, K., Wang, Y., Yuan, Z., Dujardin, T., Xu, Q., & Shi, Y. (2026). On the foundations of Earth foundation models. Communications Earth & Environment, 7, 103. https://doi.org/10.1038/s43247-025-03127-x Main Takeaways: Current Earth AI Models Are Missing the Point: Researchers have identified eleven features that an ideal Earth foundation model must have — including geolocation awareness, multi-sensor integration, physical consistency, and carbon minimization — yet no existing model comes close to checking all eleven boxes. Most models focus on only one or two features, leaving a major gap between what we have and what we actually need to tackle real-world climate and environmental challenges. The Data Situation Is More Lopsided Than You'd Think: There are now over 1,000 active remote sensing satellites generating nearly 100 petabytes of open satellite data — but labeled datasets used to train AI models account for less than 0.1% of that archive. This massive imbalance is precisely why self-supervised foundation models, which can learn from unlabeled data, are so critical for Earth science going forward. Weather AI Is Already Dramatically More Efficient — But Incomplete: Models like FourCastNet can generate a week-long global weather forecast in under two seconds on a single GPU, using roughly 12,000 times less energy than traditional forecasting systems. Despite this leap in efficiency, major gaps remain: models struggle beyond two-week forecasts, long-term climate projections drift due to incomplete energy balance, and connecting fine-scale satellite imagery with coarse climate models remains largely unsolved. What Comes After the Ideal Model: Once a true Earth foundation model exists, the authors argue the most exciting frontier is using it to build an "Earth Embedding" — a compact, unified representation of our entire planet that researchers worldwide could query without ever touching raw satellite data. Beyond that, challenges like machine unlearning (making models forget sensitive imagery), adversarial defenses, and continual learning as the climate itself changes will define the next generation of Earth AI research.

  • S2 · E5
    April 14 · 21 min

    Whose weather is it? A fairness framework for data-driven weather forecasting

    Citation: Olivetti, L., & Messori, G. (2025). Whose weather is it? A fairness framework for data-driven weather forecasting. Environmental Research Letters, 20, 121006. https://doi.org/10.1088/1748-9326/ae21f5 Main Takeaways: AI Weather Models Aren't Fair to Everyone: The latest generation of AI-powered weather forecasts improves predictions globally — but not equally. Using ECMWF's AIFS model as a case study, the authors show that wealthier and more densely populated areas consistently receive a higher share of forecast improvements compared to poorer and more rural regions, violating basic fairness criteria borrowed from the algorithmic fairness literature. Two Measurable Fairness Tests — Both Failed: The paper proposes two concrete criteria: statistical parity(improvement rates should be similar across income groups) and conditional independence (a region's GDP or population density should not predict whether it benefits from the new model). Across nearly all tested variables and forecast lead times, AIFS fails both tests at the 0.01 significance level — meaning the disparity is not a statistical fluke. Extreme Weather Is Where the Gap Hurts Most: For standard temperature and wind forecasts, gaps between rich and poor regions are modest. But for cold extremes, the fairness gap is especially pronounced — precisely the events where accurate early warnings matter most for vulnerable populations with fewer resources to adapt. Fixing It Is Technically Feasible: Unlike traditional physics-based models, AI weather models offer genuine design levers for fairness. The authors describe two practical approaches: adding penalty terms to the loss function (such as the Hilbert–Schmidt Independence Criterion) to reduce associations with protected variables, and using geographically adaptive weighting that iteratively compensates for emerging performance gaps — without necessarily sacrificing global accuracy.

  • S2 · E4
    March 7 · 18 min

    Learning predictable and informative dynamical drivers of extreme precipitation using variational autoencoders

    Citation: Spuler, F. R., Kretschmer, M., Balmaseda, M. A., Kovalchuk, Y., & Shepherd, T. G. (2025). Learning predictable and informative dynamical drivers of extreme precipitation using variational autoencoders. Weather and Climate Dynamics, 6, 995–1014. https://doi.org/10.5194/wcd-6-995-2025 Main Takeaways: Innovative Machine Learning Approach: The study introduces the Categorical Mixture Model Variational Autoencoder (CMM-VAE), a novel generative machine learning method designed to identify probabilistic atmospheric circulation regimes by combining targeted dimensionality reduction and probabilistic clustering into a single model. Resolving a Major Forecasting Trade-off: Traditionally, atmospheric regimes are either highly predictable globally but locally uninformative, or highly informative for local impacts but lacking in subseasonal predictability. CMM-VAE resolves this trade-off, successfully identifying patterns that predict local extremes without sacrificing forecast skill at subseasonal lead times. Targeted Application for Moroccan Rainfall: When applied to extreme winter precipitation in Morocco, the CMM-VAE method successfully disentangled a distinct, highly impactful weather pattern—a Scandinavian blocking coupled with a localized cut-off low—that traditional linear clustering methods failed to isolate. Linkages to Global Climate Drivers: The weather regimes identified by the model remain physically interpretable and show clear, predictable teleconnections to large-scale, low-frequency climate drivers, notably the Madden-Julian Oscillation (MJO) and the Stratospheric Polar Vortex (SPV). Enhancing Early Warning Systems: By providing a better representation of regional dynamical drivers, this framework offers significant potential to improve subseasonal-to-seasonal (S2S) forecasts, statistical downscaling, and early-warning systems for severe, localized weather impacts.

  • S2 · E3
    February 28 · 18 min

    Green and intelligent: the role of AI in the climate transition

    Green and intelligent: the role of AI in the climate transition Citation: Stern, N., Romani, M., Pierfederici, R., Braun, M., Barraclough, D., Lingeswaran, S., Weirich-Benet, E., & Niemann, N. (2025). Green and intelligent: the role of AI in the climate transition. https://doi.org/10.1038/s44168-025-00252-3. Main Takeaways: Five Key Areas for Climate Action: Artificial Intelligence can accelerate the net-zero transition across five primary avenues: transforming complex economic systems, innovating technology discovery and resource efficiency, nudging consumer behavior toward sustainable choices, modeling climate systems for better policy, and managing adaptation and resilience. Significant Emissions Reduction Potential: By applying AI to just three major sectors—power, food (specifically meat and dairy), and mobility (light road vehicles)—global emissions could be reduced by 3.2 to 5.4 GtCO2e annually by 2035. Net-Positive Climate Impact: The emissions savings generated by AI in these three sectors alone would more than offset the projected 0.4 to 1.6 GtCO2e increase in emissions caused by the energy consumption of all global AI activities and data centers. Closing the Emissions Gap: Harnessing AI to improve the efficiency and market adoption of low-carbon solutions could push global progress 36% closer to aligning with an ambitious emissions reduction trajectory by 2035. The Critical Role of Government: Relying solely on market forces to govern AI is risky; an "active state" is essential to direct AI toward public goods, regulate its environmental footprint (like mandating renewable energy for data centers), and ensure equitable deployment so the Global South is not left behind.

  • S2 · E2
    January 26 · 11 min

    Climate Knowledge in Large Language Models

    Climate Knowledge in Large Language Models Kuznetsov, I., Grassi, J., Pantiukhin, D., Shapkin, B., Jung, T., & Koldunov, N. (2025). Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research. LLMs have an internal "map" of the climate, but it is fuzzy: Without access to external tools, Large Language Models (LLMs) can recall the general structure of Earth’s climate—correctly identifying that the tropics are warm and high latitudes are cold. However, their specific numeric predictions are often inaccurate, with average errors ranging from 3°C to 6°C compared to historical weather data. Location names matter more than coordinates: The study found that providing geographic context—such as the country, region, or city name—alongside coordinates reduced prediction errors by an average of 27%. This suggests models rely heavily on text associations with place names rather than possessing a precise spatial understanding of latitude and longitude. Performance struggles with altitude and local trends: Models perform significantly worse in mountainous regions, with errors spiking sharply at elevations above 1500 meters. Furthermore, while LLMs can estimate the global average magnitude of warming, they fail to accurately reproduce the specific local patterns of temperature change that are essential for understanding regional climate dynamics. Caution is needed for scientific use: The results highlight that while LLMs encode a static snapshot of climatological averages, they lack true physical understanding and struggle with dynamic trends. Consequently, they should not be relied upon as standalone climate databases; reliable applications require connecting them to external, authoritative data sources.

  • S2 · E1
    January 11 · 13 min

    Artificial Intelligence for Atmospheric Sciences: A Research Roadmap

    Artificial Intelligence for Atmospheric Sciences: A Research Roadmap Citation: Zaidan, M. A., Motlagh, N. H., Nurmi, P., Hussein, T., Kulmala, M., Petäjä, T., & Tarkoma, S. (2025). Artificial Intelligence for Atmospheric Sciences: A Research Roadmap. Revolutionizing Environmental Monitoring: The paper illustrates how AI is transforming atmospheric sciences by bridging the gap between computer science and environmental research. It details how AI processes massive datasets generated by diverse sources—including satellite imagery, ground-based research stations, and low-cost IoT sensors—to improve our understanding of air quality, extreme weather events, and climate change. Optimizing Infrastructure and Prediction: Current AI applications are already enhancing operational meteorology and Earth system modeling. By utilizing techniques like deep learning and neural networks, researchers can automate sensor calibration, detect anomalies in real-time, and simulate complex climate scenarios with greater speed and efficiency than traditional physical models allow. A Roadmap for Future Hardware: To handle the escalating demand for data, the authors propose a hardware roadmap that includes self-sustaining and biodegradable sensor networks, CubeSat constellations for high-resolution monitoring, and the adoption of cutting-edge computing paradigms like quantum, neuromorphic, and DNA-based molecular computing. Next-Generation AI Methodologies: The paper argues for the adoption of advanced AI techniques such as Foundation Models and Generative AI (including Digital Twins of Earth) to predict complex atmospheric phenomena. Crucially, it emphasizes the need for Explainable AI (XAI) and Physics-Informed Machine Learning to solve the "black box" problem, ensuring that AI predictions abide by physical laws and are transparent enough for scientists and policymakers to trust. From Data to Action: Beyond observation, the research highlights the shift toward actionable insights. This includes automated feedback loops (such as smart HVAC systems responding to air quality data), the integration of citizen science to augment data collection, and the establishment of robust ethical frameworks to manage data privacy and governance in global monitoring networks.

  • S1 · E43
    Dec 19, 2025 · 13 min

    Differentiable and accelerated spherical harmonic and Wigner transforms

    Differentiable and accelerated spherical harmonic and Wigner transforms Matthew A. Price, Jason D. McEwen *Journal of Computational Physics (2024)* * This work introduces novel algorithmic structures for the **accelerated and differentiable computation** of generalized Fourier transforms on the sphere ($S^2$) and the rotation group ($SO(3)$), specifically spherical harmonic and Wigner transforms. * A key component is a **recursive algorithm for Wigner d-functions** designed to be stable to high harmonic degrees and extremely parallelizable, making the algorithms well-suited for high throughput computing on modern hardware accelerators such as GPUs. * The transforms support efficient computation of gradients, which is critical for machine learning and other differentiable programming tasks, achieved through a **hybrid automatic and manual differentiation approach** to avoid the memory overhead associated with full automatic differentiation. * Implemented in the open-source **S2FFT** software code (within the JAX differentiable programming framework), the algorithms support various sampling schemes, including equiangular samplings that admit exact spherical harmonic transforms. * Benchmarking results demonstrate **up to a 400-fold acceleration** compared to alternative C codes, and the transforms exhibit **very close to optimal linear scaling** when distributed over multiple GPUs, yielding an unprecedented effective linear time complexity (O(L)) given sufficient computational resources.

  • S1 · E42
    Dec 11, 2025 · 12 min

    Score-based diffusion nowcasting of GOES imagery

    Score-based diffusion nowcasting of GOES imagery *Randy J. Chase, Katherine Haynes, Lander Ver Hoef, Imme Ebert-Uphoff, a Cooperative Institute for Research in the Atmosphere, Colorado State University, Fort Collins, CO, b Electrical and Computer Engineering, Colorado State University, Fort Collins, CO* * The research explored score-based diffusion models to perform short-term forecasts (nowcasting) of GOES geostationary infrared satellite imagery (zero to three hours). This newer machine learning methodology combats the issue of **blurry forecasts** often produced by earlier neural network types, enabling the generation of clearer and more realistic-looking forecasts. * The **residual correction diffusion model (CorrDiff)** proved to be the best-performing model, quantitatively outperforming all other tested diffusion models, a traditional Mean Squared Error trained U-Net, and a persistence forecast by one to two kelvin on root mean squared error. * The diffusion models demonstrated sophisticated predictive capabilities, showing the ability to not only advect existing clouds but also to **generate and decay clouds**, including initiating convection, despite being initialized with only the past 20 minutes of satellite imagery. * A key benefit of the diffusion framework is the capacity for **out-of-the-box ensemble generation**, which enhances pixel-based metrics and provides useful uncertainty quantification where the spread of the ensemble generally correlates well to the forecast error. * However, the diffusion models are computationally intensive, with the Diff and CorrDiff models taking approximately five days to train on specialized hardware and about 10 minutes to generate a 10-member, three-hour forecast, compared to just 10 seconds for the baseline U-Net forecast.

  • S1 · E41
    Dec 4, 2025 · 15 min

    FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution

    FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution *Qiusheng Huang, Yuan Niu, Xiaohui Zhong, Anboyu Guo, Lei Chen, Dianjun Zhang, Xuefeng Zhang, Hao Li* --- * **First Data-Driven Sub-Daily Global Forecast:** FuXi-Ocean is the first deep learning-based global ocean forecasting model to achieve six-hour temporal resolution at an eddy-resolving 1/12° spatial resolution, with vertical coverage extending up to 1500 meters. This capability addresses a crucial need for high-frequency predictions that traditional numerical models struggle to deliver efficiently. * **Adaptive Temporal Modeling Innovation:** A key component of the model is the **Mixture-of-Time (MoT) module**, which adaptively integrates predictions from multiple temporal contexts based on variable-specific reliability. This mechanism is crucial for accommodating the diverse temporal dynamics of different ocean variables (e.g., fast-changing surface variables vs. slowly evolving deep-ocean processes) and effectively mitigates the accumulation of forecast errors in sequential prediction. * **Superior Performance and Efficiency:** The model demonstrates superior skill in predicting key variables (temperature, salinity, and currents) compared to state-of-the-art operational numerical forecasting systems (like HYCOM, BLK, and FOAM) at sub-daily intervals. Furthermore, it achieves this high performance with remarkable data efficiency, requiring only approximately 9 years of training data and relying solely on ocean variables (T, S, U, V, SSH) as input, without external data dependencies like atmospheric forcing. * **High-Impact Applications:** By providing accurate, high-resolution, sub-daily forecasts, FuXi-Ocean creates critical opportunities for maritime operations, including improved navigation, search and rescue, oil spill trajectory tracking, and enhanced marine resource management, particularly due to its comprehensive vertical coverage (0-1500 m).

  • S1 · E40
    Nov 28, 2025 · 15 min

    Beyond the Training Data: Confidence-Guided Mixing of Parameterizations in a Hybrid AI-Climate Model

    Beyond the Training Data: Confidence-Guided Mixing of Parameterizations in a Hybrid AI-Climate Model *By Helge Heuer, Tom Beucler, Mierk Schwabe, Julien Savre, Manuel Schlund, and Veronika Eyring* * This paper presents a **successful proof-of-concept for transferring a machine learning (ML) convection parameterization**—trained on the ClimSim dataset—to the ICON-A climate model. The resulting hybrid ML-physics model achieved stable and accurate simulations in long-term AMIP-style runs lasting at least 20 years. * A core innovation is the **confidence-guided mixing scheme**, which allows the Neural Network (NN) to predict its own error. When the NN's predicted confidence is low (e.g., in moist, unstable regimes or high-variability areas), its prediction is mixed with the conventional Tiedtke convection scheme. This mechanism improves reliability, prevents unphysical outputs by detecting potential extrapolation beyond the training domain, and makes the hybrid model tunable against observations. * The scheme's robustness and accuracy were further enhanced through the **use of a physics-informed loss function**—which encourages adherence to conservation laws like enthalpy and mass—and **noise-augmented training**. These techniques mitigate stability issues commonly faced by ML parameterizations and significantly improve physical consistency compared to purely data-driven models. * In evaluation against observational data, several hybrid configurations **outperformed the default Tiedtke scheme**, demonstrating improved precipitation statistics and showing a better representation of global climate variables. The confidence-guided approach demonstrated a fundamental change in the model's behavior, with the ML component contributing approximately 67% of the convective tendencies on average.

  • S1 · E39
    Nov 23, 2025 · 13 min

    Climate in a Bottle: Towards a Generative Foundation Model for the Kilometer-Scale Global Atmosphere

    Climate in a Bottle: Towards a Generative Foundation Model for the Kilometer-Scale Global Atmosphere (By Noah D. Brenowitz, Tao Ge, Akshay Subramaniam, Peter Manshausen, Aayush Gupta, David M. Hall, Morteza Mardani, Arash Vahdat, Karthik Kashinath, Michael S. Pritchard, NVIDIA * The paper introduces **Climate in a Bottle (cBottle)**, a generative diffusion-based AI framework capable of synthesizing full global atmospheric states at an unprecedented $\mathbf{5 \text{ km resolution}}$ (over 12.5 million pixels per sample). Unlike prevailing auto-regressive paradigms, cBottle samples directly from the full distribution of atmospheric states without requiring a previous time step, thereby avoiding issues like drifts and instabilities inherent to time-stepping models. * cBottle utilizes a **two-stage cascaded diffusion approach**: a global coarse-resolution generator conditioned on minimal climate-controlling inputs (such as monthly sea surface temperature and solar position), followed by a patch-based 16x super-resolution module. * The model demonstrates **foundational versatility** by being trained jointly on multiple data modalities, including ERA5 reanalysis and ICON global cloud-resolving simulations. This enables various zero-shot applications such as climate downscaling, channel infilling for missing or corrupted variables, bias correction between datasets, and translation between these modalities. * cBottle proposes a new form of **interactive climate modeling** through the use of guided diffusion. By training a classifier alongside the generator, users can steer the model to conditionally generate physically plausible **extreme weather events, such as Tropical Cyclones**, at specified locations on demand, circumventing the need to sift through petabytes of output to find rare events. * The model exhibits **high climate faithfulness** across a battery of tests, including reproducing diurnal-to-seasonal scale variability, large-scale modes of variability (like the Northern Annular Mode), and tropical cyclone statistics. Furthermore, it achieves **extreme distillation** by encapsulating massive datasets into a few GB of neural network weights, offering a 256x compression ratio per channel.

  • S1 · E38
    Nov 7, 2025 · 13 min

    Probabilistic Measures for Fair AI and NWP Model Comparison

    Probabilistic measures afford fair comparisons of AIWP and NWP model output (Tilmann Gneiting, Tobias Biegert, Kristof Kraus, Eva-Maria Walz, Alexander I. Jordan, Sebastian Lerch, June 10, 2025) Introduction of a New Fair Comparison Metric: The paper introduces the Potential Continuous Ranked Probability Score (PC), a new measure designed to allow fair and meaningful comparisons between single-valued output from data-driven Artificial Intelligence based Weather Prediction (AIWP) models and physics-based Numerical Weather Prediction (NWP) models. This approach addresses concerns that traditional loss functions (like RMSE) may unfairly favor AIWP models, which often optimize their training using these metrics. Methodology Based on Probabilistic Postprocessing: PC is calculated by applying the same statistical postprocessing technique—specifically Isotonic Distributional Regression (IDR), also known as Easy Uncertainty Quantification (EasyUQ)—to the deterministic output of both AIWP and NWP models. PC is then defined as the mean Continuous Ranked Probability Score (CRPS) of these newly generated probabilistic forecasts. Measure of Potential Skill and Invariance: PC quantifies potential predictive performance. A key property of PC is that it is invariant under strictly increasing transformations of the model output, treating both forecasts equally and facilitating comparisons where the pre-specification of a loss function might otherwise place competitors on unequal footings. AIWP Outperformance and Operational Proxy: When applied to WeatherBench 2 data, the PC measure demonstrated that the data-driven GraphCast model outperforms the leading physics-based ECMWF high-resolution (HRES) model. Furthermore, the PC measure for the HRES model was found to align exceptionally well with the mean CRPS of the operational ECMWF ensemble, confirming that PC serves as a reliable proxy for the performance of real-time operational probabilistic products.

Showing 1–20 of 23 episodes