Zenith Flow
Baseline: Eulerian persistence · latest layer pending · 0 settled forecasts · 0 fully paired · 0 pending
Rolling 90-day scorecards compare public lanes and explicitly labeled shadow challengers to declared baselines. Public claims only publish when completeness is at least 80% and the bootstrap 95% confidence interval excludes zero.
Methodology visible
Live scorecard uses an older verification contract; v2 data is not displayed as Flow evidence.
This separates the public methodology from the live scoring feed. A lane can show methodology and baselines before it has enough settled samples to publish a claim.
Baseline: Eulerian persistence · latest layer pending · 0 settled forecasts · 0 fully paired · 0 pending
Baseline: HRRR · latest layer pending
Baseline: Raw NBM · latest layer pending
Baseline: GFS · latest layer pending
Baseline: Eulerian persistence · Metric: Precipitation rate MAE (mm/h)
verifying
≥0.5 mm/h
CSI — · POD — · FAR —
Persistence CSI —
≥2 mm/h
CSI — · POD — · FAR —
Persistence CSI —
≥10 mm/h
CSI — · POD — · FAR —
Persistence CSI —
Baseline: HRRR · Metric: Temperature MAE (°F)
verifying
Baseline: Raw NBM · Metric: Temperature MAE (°F)
verifying
Baseline: GFS · Metric: Temperature MAE (°F)
verifying
Flow compares to Eulerian persistence; Middle to HRRR; Regional Lite to raw NBM; Global FCN3 to GFS. Missing baseline samples keep the lane in verifying state.
Flow T+15 is scored only after the nearest archived NOAA MRMS frame (within eight minutes) has settled for at least two hours. Samples store forecast, persistence, and truth layer identities plus truthSettledAt.
Flow is event-driven, not continuously fresh. Every produced event run creates 5 fixed-station scoring opportunities; missed samples stay missing. Execution is capped at 4 runs per UTC day, and public claims require at least 80% opportunity coverage.
Flow uses model-QC reflectivity converted to precipitation rate. Covered dry pixels are zero; outside or unknown coverage fails closed as missing. Display-smoothed radar is never verification input.
Daily lane-vs-baseline skill deltas are bootstrapped into a 95% interval. If the interval crosses zero, the card shows verifying instead of a marketing claim.
The public claim gate requires 30 distinct days of Zenith Flow against Eulerian persistence; current proof is 0 days, so unproven claims remain verifying.
A statistically settled loss is shown in the same card style as a win, with the baseline, window, metric, and completeness still visible.
ECMWF AIFS is sampled against Global FCN3 as an internal challenger. Its artifacts stay research-only and routing-ineligible regardless of measured skill until a separate promotion decision.