Limits
A benchmark that hides its limits is an ad. These are ours, global first, then per category as generated from each current bundle.
Global limits
- Results expire. Each bundle has
valid_untiland a freshness state (CURRENT, STALE, SUPERSEDED). A stale result is never served as a current recommendation. - Results are population-bound. A score on our frozen dataset does not transfer to a different language, market or workload; read
populationandcontextbefore reusing a number. - Statuses can refuse to conclude.
INDETERMINATEandINSUFFICIENT_EVIDENCEare first-class outcomes; deriving a ranking from them is a misuse of this data. - Categories prefixed
fixture-are synthetic demo data used to validate the machinery. Their limits say so in the data itself; they are not evidence about any real tool. - Cost figures use public prices at
observed_at; vendors change prices without notice. - We measure what the protocol froze, nothing else. Vendor features outside the frozen objective are out of scope.
Per-category limits (from the current bundles)
fixture-widgets OK
- Synthetic fixture: exercises the mechanics, measures no real tool.
- Success is a deterministic arithmetic pattern, not a market observation.
fixture-widgets-smalln INDETERMINATE
- Synthetic fixture: exercises the mechanics, measures no real tool.
- Success is a deterministic arithmetic pattern, not a market observation.
- Precision insufficient under the frozen protocol: no rank published.