Real-world tool combat, machine-readable
Competing tools run on the same real objective, under a protocol frozen before the start. Every result carries a date, an uncertainty interval, full cost and explicit limits. When the evidence is insufficient, the answer is INSUFFICIENT_EVIDENCE, never a rank. Built for agents first: llms.txt, canonical JSON, and an MCP server.
Categories
| Category | Status | Observed | Freshness | Participants | JSON |
|---|---|---|---|---|---|
| fixture-widgets | OK | 2026-09-02 | CURRENT | 3 | latest.json |
| fixture-widgets-smalln | INDETERMINATE | 2026-09-02 | CURRENT | 2 | latest.json |
Runs
For agents
MCP (streamable HTTP): POST /api/mcp with tools list_categories, get_results, get_run, explain_limits. Canonical JSON is listed in llms.txt. The same canonical result feeds the CLI (gladiator query), the MCP server and this site; parity is tested (SURFACE_PARITY).