Skip to content

Research

A forecast is a claim. Advice is a claim. We grade both.

Source-model skill is live. Whether the weather agent makes the right operational call — that bench is still being built.

Live · updated daily

Forecast analysis

How the models we ingest score against NOAA analysis.

We do not train a forecast model. We ingest the National Blend of Models, HRRR, RRFS, and GFS, serve a priority blend (NBM first), and publish each source model's skill against NOAA analysis — URMA for the surface, MRMS Multi-Sensor QPE for rain. The blend we serve is not a row on the page.

CONUS forecast skill map on the public forecast analysis page
Neighborhood skill, not a national average

What skill means here

For temperature, wind, and other continuous fields the headline is a mean-squared-error skill score: one minus the model's squared error over the variance of the analysis values we actually scored. The reference is this sample's local average, not a 30-year climate normal. Chance of rain uses a Brier skill score against how often it rained in the sample. Rain-or-not uses an equitable threat score, which does not get easier just because a climate is wet. Negative skill is published. A cell with fewer than 30 pairs is left blank, not smoothed.

Why the record has to be kept as it happens

The analysis we pair against expires in about two days. A valid hour we miss in that window is gone — we cannot reconstruct last month. That is why a public skill number is rare, and why starting the same measurement today still starts at zero.

How to read it

We accumulate by neighborhood, about 50 km across, because a national average mixes the Gulf Coast with the Front Range. We do not put two models side by side and declare a winner: those rows are not a matched sample. When NBM is stale and we served HRRR, that hour is not in the NBM card. Formulas, the 30-pair floor, and what we withhold are on the methodology page.

In progress

Weather Bench

How well the weather agent decides for real operations.

Forecast skill asks whether the air was right. Decision skill asks whether the call was right. A model can nail tonight's temperature and still tell a campground to stay open. Weather Bench, still being built, grades the weather agent we are building for operations.

Not a new kind of eval

Labs that bench coding and enterprise agents already solved the hard parts. The public writeups use the same bones: real tasks instead of invented prompts, a sealed run so the agent cannot look up the answer, and graders that are code — not another model reading the prose. Newest work is held out. A human expert sits the same exam. We take that method as-is. The domain is operational weather.

The job under test

The weather agent, for one customer. We replay it through real past weather rebuilt from public NOAA archives. It sees only what was knowable at the time — no web, no live data. Code then asks a few questions. Did it act when the operation needed a call? Did it stay quiet otherwise? Did it reach the right people, cite the official product, and put odds on the decision? Tasks are seeded from real National Weather Service warning decisions — expert calls that already happened — not from prompts we invented.

What we will not claim yet

No public score yet. The first full run has not been scored. When it is, the headline will be reliability, published with intervals. We will hold back the newest events so we cannot tune on them, and add a human meteorologist sitting the same exam. Until then this is a method, not a number.

Sources

Design we adapted. We are not a participant in these benches.

Results publish when the first full run is scored.