Methodology
How we grade, what each number means, which station settles each city, where we have gaps, and how you can verify every figure yourself. Record start date: to be announced; all performance figures are Sample data until then.
Slot types: forecast vs live confirmation
We snapshot the daemon's current forecast at fixed local-time slots and grade them separately — never blended into one headline:
| Slot | Local window | Class |
|---|---|---|
| eve_before (high) | 18:00–21:00, day before | day-ahead forecast |
| morning_of (high) | 06:00–09:00, market day | forecast (peak not yet reached) |
| low_eve_before (low) | 20:00–22:00, day before | day-ahead forecast |
| pre_entry (high) | 16:00–17:30, market day | live confirmation — includes same-day METAR |
| low_pre_dawn (low) | 04:00–06:00, market day | live confirmation |
pre_entry and low_pre_dawn carry the day-0 live-METAR override, so they measure near-real-time confirmation, not forecast skill. They are labelled as such everywhere.
How grading works
The settled result is the venue's resolved winning bin (read from the market's own rules source). We also take an independent cross-check: the full station-local-day maximum from IEM ASOS (a mirror of the same NOAA observations). Per graded snapshot we report hit rate (our most-likely bin == the venue winning bin), signed error and MAE in °F, each beside persistence (yesterday's value) and climatology (same calendar-week, prior years) baselines with a skill score. Bin bounds come from the venue ladder, not fixed widths. Small samples (n<40) are flagged; confidence intervals are shown with the caveat that consecutive days are autocorrelated, so the effective sample is smaller than n. We do not publish PMF-based calibration (Brier/log-loss) until the daemon's real PMF is logged.
Chains are sealed and Bitcoin-anchored (OpenTimestamps) hourly, with full chain verification in the daily seal. So a forecast's precedence over its outcome is provable to within one hour — honest resolution, which matters for cities whose market day ends well before the daily seal (e.g. Tokyo).
Settlement source per city
Verified from each venue's own rules text (Polymarket). data_feed is how we read the observation; settlement_source is the authority the market settles on.
| City | Station (ICAO) | Settlement source | Unit | Bin |
|---|---|---|---|---|
| New York | LaGuardia (KLGA) | NOAA weather.gov | °F | 2°F |
| Denver | Buckley SFB (KBKF) | NOAA weather.gov | °F | 2°F |
| Dallas | Love Field (KDAL) | NOAA weather.gov | °F | 2°F |
| Chicago | O'Hare (KORD) | NOAA weather.gov | °F | 2°F |
| Atlanta | Hartsfield-Jackson (KATL) | NOAA weather.gov | °F | 2°F |
| Austin | Austin-Bergstrom (KAUS) | NOAA weather.gov | °F | 2°F |
| Houston | Hobby (KHOU) | NOAA weather.gov | °F | 2°F |
| Los Angeles | LAX (KLAX) | NOAA weather.gov | °F | 2°F |
| Miami | Miami Intl (KMIA) | NOAA weather.gov | °F | 2°F |
| San Francisco | SFO (KSFO) | NOAA weather.gov | °F | 2°F |
| Seattle | Sea-Tac (KSEA) | NOAA weather.gov | °F | 2°F |
All eleven carry a conditional fallback: Weather Underground Daily Observations is used only if NOAA data is unavailable by 11:59 PM ET the day after the observation date. More cities are added as each venue's rules text confirms them.
Known gaps
We publish our own gaps. Rendered from known_gaps.jsonl:
| When | Gap | Status |
|---|---|---|
| before 2026-10-07 04:36Z | logger not yet running (setup week) | historical |
| 2026-10-07 | partial slot coverage (early/absent slots) | historical |
| 2026-10-07 | Wellington pre_entry missed (logger first run 6 min after the window) | resolved 2026-10-08 |
Verify it yourself
Every row is row_hash = sha256(prev_hash + content_hash + timestamp), genesis prev_hash = "GENESIS". The daily seal commits each chain's head hash and is Bitcoin-anchored (OpenTimestamps). To check a published export against a seal:
import json, hashlib, sys
def sha(s): return hashlib.sha256(s.encode()).hexdigest()
prev, ok, head = "GENESIS", True, "GENESIS"
for line in open(sys.argv[1]): # e.g. settlement_registry.jsonl
r = json.loads(line)
if r["prev_hash"] != prev: ok = False
if r["row_hash"] != sha(r["prev_hash"] + r["content_hash"] + r["ts"]): ok = False
prev = head = r["row_hash"]
print("chain_ok =", ok, " head =", head)
# compare head to the chain's head in that day's seal_<day>.txt
# then: ots verify seal_<day>.txt.ots (confirms the seal existed at that Bitcoin block)
Public export (one-day delay)
Each UTC day, after that day's markets have settled, we publish to a public hash repository, on a one-day delay, three independent append-only chains as JSONL (each row carrying its payload_json + content/prev/row hashes), plus that day's seal: the dedicated official_forecasts chain, the settlement registry, and outcomes.
What you can verify from public data: all three chains verify standalone — recompute each row's content_hash from its payload_json and each row_hash from prev_hash+content_hash+timestamp, walk to the head, and match the head to the Bitcoin-anchored seal. The official_forecasts chain is written live at capture time and is self-contained, so its export is a verifiable prefix (not a filtered subset).
What is NOT in public data: per-model values (model_pmf_json, models_json are never exported), adhoc (non-official) snapshots, and any market whose outcome is not yet final. The main full snapshot chain (which includes adhoc rows) is not published; official_forecasts is the public, verifiable view. The main_row_id/main_row_hash pointers inside official_forecasts are internal cross-references — they are not independently verifiable from public data; the official chain stands entirely on its own hashes and seal. Backfilled rows are marked provenance=backfill with their creation date, distinct from live-written rows, and are excluded from headline accuracy.