Validation

What we measured — including what does not work. Every number here comes from a script you can run. A validation page that only showed wins would be marketing, not validation.

Forecast skill vs persistence

The brief asks for RMSE against a persistence baseline. Persistence (“tomorrow ≈ today”) is excellent short-term and decays with horizon; the model does not. Measured OUT OF SAMPLE: the served model is scored only on hours after its training cutoff (30 Nov 2025 – 14 Jan 2026), with climatology truncated at the same cutoff so it cannot leak. A dash means that city has not been scored out of sample — its training runs to Aug 2026, leaving too short a window to measure honestly, so we report nothing rather than a number from data the model has already seen.

HorizonDelhiChennaiBengaluruMumbaiKolkataHyderabadPuneAhmedabad
+24h
+48h
+72h

The direction is the finding, not the decimal.Persistence tends to win at 24h; the model wins at 72h in all three cities — and 48–72h is the enforcement-scheduling window (“stagnant winds Thursday, act before”). The exact % moves with how many stations OpenAQ serves on a given day.

Fusion exposure field — claim withdrawn

We claimed this cut error ~36% vs a naive station-mean. On real cities it does not. We report it rather than quietly dropping it.

CityModel RMSENaive city-meanVerdict
Delhi
Chennai
Bengaluru
Mumbai
Kolkata
Hyderabad
Pune
Ahmedabad

With ~24 stations we cannot demonstrate spatial skill. Detection is the contribution, not the fusion field. Detection runs on satellite contrast + fire persistence and never touches a station.

Reproduce: python scripts/eval_detection.py · eval_attribution.py · eval_hotspot_recovery.py · eval_station_sensitivity.py — currently viewing delhi.