How we score it
Every score on this site is tracked data, not opinion. One rubric runs against every forecaster, every day. Here is what it measures, and where the actual weather comes from.
The 100-point model
Each day we compare a forecast to the actual recorded conditions across four fields. Closer to the truth earns more points, out of 100:
| Field | Max | Full credit when… |
|---|---|---|
| High temp | 30 | within 1°F of actual, then −3 pts per °F beyond. |
| Low temp | 30 | within 1°F of actual, then −3 pts per °F beyond. |
| Wind | 20 | within 3 mph of actual, then −2 pts per mph beyond. A forecast given as a rangeis scored at its midpoint, then taxed by half the range width for vagueness. A 5–15 mph range scores lower than a precise 10 mph for the same midpoint. A single number has zero width and pays no tax. |
| Precip | 20 | One field for what fell and how much — type and amount are graded together so the same wet-or-dry fact is never scored twice. On a dry day the whole 20 measures the predicted amount against zero: predict no precip and none falls, full marks. On a wetday it splits into 10 for identifying the form (rain, snow, or mixed) and 10 for the amount — rain within 0.1″, then −2 pts per extra 0.1″; snow within 1″ or 20% of the actual, whichever is larger, because snow totals are noisier. Naming the wrong form, say rain when it snowed, caps the amount half at 5. A miss inside the trace band still earns 6 of the 10 identification points: when the only disagreement between “none” and “precip” is an amount small enough to count as zero (rain of 0.1″ or less, snow of 1″ or less), the call was nearly right, and no forecast is graded fully wrong on the form while fully right on the total. |
Saying “no rain” is a forecast too
Every forecast is scored out of a fixed 100. The one wrinkle is the rain total. A forecast of no rain is a zero-inch prediction, so on a dry day it earns those points like any other right call. A forecast of rain with no stated total leaves the amount blank and earns nothing for it. The trace-band partial credit above follows the same logic in reverse, and it also demands the number: a source that predicted rain but never said how much cannot claim its miss was only a trace.
That matters for Ray's Weather, who never publishes a numeric rain total. He earns the amount points on the days he forecasts dry, and forfeits them on the days he forecasts rain. Staying quiet on the hard number cannot beat answering it. A blank is never worth more than a real forecast.
When the forecast is in words
Ray's often describes wind in words instead of numbers. We map those to the National Weather Service scale, then score the result as a range like any other:
- calm: 0–1 mph
- light: 1–7 mph
- breezy: 12–20 mph
- windy or gusty: 18–30 mph
A word only counts as a wind descriptor when it sits next to the word “wind,” so a phrase like “light rain” is not read as light wind. We also strip any “gusting to N” clause, so only sustained wind is scored.
Reading the overnight low
Most services hand us a finished daily low that already covers the whole day, including the pre-dawn hours. Two of them, Met.no and OpenWeather, don't. We rebuild their daily low ourselves from an hour-by-hour feed, and by the time we capture at midday that feed no longer reaches back to the overnight low.
So for those two we read the low from the forecast they published the morning before, when the full day was still ahead. That is a longer lead time than the same-day number every other source gets, so if anything it is a slightly harder test. Sources that give us a full-day low directly are scored on it as is.
Grades
The day's score becomes a verdict:
- 90–100: Right (5 rays)
- 75–89: Right (4 rays)
- 60–74: Meh (3 rays)
- under 60: Wrong (1–2 rays)
How the Dave's Sweater Index is built
The Dave's Sweater Index (DSI) is our own forecast — a consensus of the independent automated forecasters we track. It is not a black box, and it is graded by the same 100-point rubric as every source it draws on. Here is exactly how each day's number is made:
- Members. Every free, automated forecaster with a forecast that day. Ray's Weather is excluded (it's the forecast we grade against), as is the Apple slot when it's the Open-Meteo fallback (it would double-count Open-Meteo). The index forms only when at least two members report.
- High, low, wind, amount. A straight average of the members. No source is weighted above another — the point of a consensus is that independent errors cancel.
- Precip type — the one rule. A plain majority vote is the wrong tool here: rain days are the minority, and the costly miss is calling a wet day “dry.” So a dry majority does not get to veto a credible minority. If at least a quarter of the members (and never fewer than two) forecast precipitation, the index forecasts precipitation; rain versus snow follows the majority among those callers, and a genuine rain/snow split reads “wintry mix.”
That precip rule is the only place the index does anything other than average, and it uses no weighting and no memory of past performance — it's a fixed rule you can apply by hand to any day's forecasts. We adopted it because, measured across the record, it recovers the points a majority vote was throwing away on marginal-precipitation days. When we change how the index is built, it's to make it more accurate against what the sky actually did — and the change shows up here first.
What counts as “actual”
The conditions we grade against come from the Open-Meteo historical archive, a reanalysis of observed weather for Boone. We say this plainly because it cuts both ways. One forecaster we track, Open-Meteo, is graded partly against its own provider's archive. The same numbers apply to every source, and the free-versus-Ray's comparison that carries the thesis does not lean on that one source. To close the gap for good, we are standing up an independent weather station in Boone. Once it is live, its readings become the actual.
Every town, its own numbers
Boone is the flagship, but we grade a forecast for each town we track, and every one stands on its own. The same 100-point rubric runs unchanged from town to town — no per-town tuning — but each town's forecasts are graded against that town's own actuals: the Open-Meteo historical archive read at that town's own coordinates, not Boone's. A forecast for a 5,400-foot ridge is scored against what happened on that ridge.
We never blend towns into one average. A combined score would mix places of very different forecast difficulty, and the whole point is that each place is real. A new town launches provisional and stays that way until it crosses 9scored days — the same gate every forecast source on the site clears before its record is called established — then it ranks like the rest. The future Boone weather station upgrades Boone's actuals only; every other town keeps the archive as its ground truth.
Ray's Weather gives each town its own high, low and sky icon, so his town boards are graded on all three. We read the icon as his precipitation call — a dry icon counts as a no-rain forecast (and earns the amount points on days it stays dry), a rain or thunderstorm icon as rain — and the mapping is a fixed, source-blind lookup from his icon families (dry, rain, lightning, snow); an icon we don't recognize is forfeited, never guessed. His wind, by contrast, is a single regional sentence repeated verbatim on every town page, so we don't grade town wind against it: an honest forfeit beats a stamped number. Numeric precip amount he still never publishes, so on wet days he earns the form call and forfeits the amount, exactly as on the Boone board.
Grading forecasts by lead time
The daily scoreboard grades each morning's same-day forecast. We also grade every forecast at longer range: for a given day, the forecast each source published one to five days earlier, scored with the same 100-point model. Nothing about the rubric changes with distance, only the difficulty. Lead 0 is the same-day forecast; lead 3 is what a source said three days out.
We stop at five days, the longest horizon we can score consistently across sources. Ray's rows reliably reach about four to five days out, but his five-day sample is a single scored day so far, so the charts floor it out: a lead is not charted until it holds at least 10 scored days. One pattern already holds at every horizon we track: Ray's highs have averaged +3.1 to +3.5°F warm.
Two honesty notes. The comparison is at matched elevation: his Boone station sits at 3,240 feet, our grading point at 3,242. And on the high temperature alone, Open-Meteo beats Ray's at every horizon, but by days 3 and 4 that gap narrows to a few tenths of a degree. The score gap is the meaningful one: it prices the fields Ray's extended days leave unanswered.
The actuals behind every lead are the same Open-Meteo archive described above, and the caveat there applies to lead-time scoring too.
Grading the road-condition forecast
The road-condition forecast uses no new data. It reads the snow, ice, and temperature we already forecast and sorts each day onto one of five ordered surface levels, best to worst: Clear, Wet, Slushy, Icy, Hazardous. The exact thresholds, the same numbers the code runs:
- Hazardous — 2″ or more of forecast snow, or freezing rain (rain falling into sub-freezing air).
- Icy — any snow or rain with an overnight low at or below 30°F, when a wet surface can refreeze into black ice.
- Slushy — 0.1″ or more of snow above that refreeze line.
- Wet — 0.1″ or more of rain with no freeze expected.
- Clear — dry, or only a trace.
Scoring is ordinal, because the levels are ordered. An exact call scores 100, and every level of distance between the forecast and what happened costs 25 points: one level off is 75, two is 50, and so on to a floor of zero. Calling Icy when it was Slushy is a near miss; calling Clear when it was Hazardous is not.
The actual comes from NCDOT's DriveNC road-condition report for Division 11 (Watauga, Avery, and Ashe), reduced to the worst surface it lists that day. Two honesty notes, in the spirit of the weather caveat above. Those categories are entered by hand by people in the field, not a sensor, so the actual is softer than a measured number. And this is one public agency grading against another's report; a future phase adds our own roadside cameras as an independent ground truth, the same arc as the weather station. Off-season the report reads “No Report” and the forecast simply accrues no scored days until winter.
What each service reports
Not every forecaster publishes every field. Here is what each one gives us to score. A blank means the question went unanswered, not that they got it wrong.
Green cells are fully reported fields. Dark orange marks a deliberate gap, like Ray's, who never publishes a precip amount. Lighter cells are days a value was not available to scrape.
- Open-Meteo
- High temp508/508
- Low temp508/508
- Wind508/508
- Precip508/508
- Ray's Weather
- High temp142/142
- Low temp142/142
- Wind142/142
- Precip59/142
To see the model applied, read the 118-day review of Ray's Weather, the June 2026 report card, or what a 10-day forecast is actually good for.
