What we measured — including what does not work. Every number here comes from a script you can run. A validation page that only showed wins would be marketing, not validation.
The brief asks for RMSE against a persistence baseline. Persistence (“tomorrow ≈ today”) is excellent short-term and decays with horizon; the model does not. Measured OUT OF SAMPLE: the served model is scored only on hours after its training cutoff (30 Nov 2025 – 14 Jan 2026), with climatology truncated at the same cutoff so it cannot leak. A dash means that city has not been scored out of sample — its training runs to Aug 2026, leaving too short a window to measure honestly, so we report nothing rather than a number from data the model has already seen.
| Horizon | Delhi | Chennai | Bengaluru | Mumbai | Kolkata | Hyderabad | Pune | Ahmedabad |
|---|---|---|---|---|---|---|---|---|
| +24h | — | — | — | — | — | — | — | — |
| +48h | — | — | — | — | — | — | — | — |
| +72h | — | — | — | — | — | — | — | — |
The direction is the finding, not the decimal.Persistence tends to win at 24h; the model wins at 72h in all three cities — and 48–72h is the enforcement-scheduling window (“stagnant winds Thursday, act before”). The exact % moves with how many stations OpenAQ serves on a given day.
We claimed this cut error ~36% vs a naive station-mean. On real cities it does not. We report it rather than quietly dropping it.
| City | Model RMSE | Naive city-mean | Verdict |
|---|---|---|---|
| Delhi | — | — | — |
| Chennai | — | — | — |
| Bengaluru | — | — | — |
| Mumbai | — | — | — |
| Kolkata | — | — | — |
| Hyderabad | — | — | — |
| Pune | — | — | — |
| Ahmedabad | — | — | — |
With ~24 stations we cannot demonstrate spatial skill. Detection is the contribution, not the fusion field. Detection runs on satellite contrast + fire persistence and never touches a station.
Reproduce: python scripts/eval_detection.py · eval_attribution.py · eval_hotspot_recovery.py · eval_station_sensitivity.py — currently viewing delhi.