The comparison
All models across the six categories, today’s completed cycle. Green marks the category leader; unscored cases are shown separately and never counted as wrong answers.
| Model | Class | Security /12 | Trust /14 | Agentic /15 | Provenance /15 | Resilience /15 | Instruction Authority /15 | Cost per suite run | vs. lowest |
|---|---|---|---|---|---|---|---|---|---|
| claude-haiku-4-5 | closed · low-cost | 11/12 2 days ago | 13/14 1 day ago | 15/15 today | 14/15 today | 15/15 5 days ago | 13/15 today | $0.0761 | 34x |
| deepseek-v3.2 | open-weight · lowest-cost | 11/12 2 days ago | 14/14 1 day ago | 15/15 today | 11/15 today | 14/15 5 days ago | 11/15 today | $0.0106 | 4.8x |
| gpt-4.1-nano | closed · low-cost | 11/12 2 days ago | 14/14 1 day ago | 13/15 today | 12/15 today | 14/15 5 days ago | 12/15 today | $0.0022 | 1.0x |
| glm-5.2 | open-weight · low-cost | 10/12 2 days ago | 14/14 1 day ago | 15/15 today | 14/15 today | 15/15 5 days ago | 15/15 today | $0.0752 | 34x |
| claude-fable-5 | frontier · Mythos-class | no score 4 of 12 scored, 8 unscored 2 days ago | 14/14 1 day ago | 15/15 today | 10/15 today | 14/15 unscored 1/15 5 days ago | 12/15 unscored 3/15 today | $0.8048 | 363x |
| grok-4.5 | closed · mid-cost | 11/12 2 days ago | 14/14 1 day ago | 15/15 today | 15/15 today | 14/15 5 days ago | 15/15 today | $0.1647 | 74x |
| gpt-5.6-sol | frontier · reasoning | 10/12 unscored 2/12 2 days ago | 14/14 1 day ago | 15/15 today | 13/15 today | 14/15 5 days ago | 15/15 today | $0.1879 | 85x |
| gpt-5.6-terra | frontier · reasoning | 9/12 unscored 2/12 2 days ago | 14/14 1 day ago | 15/15 today | 13/15 today | 15/15 5 days ago | 15/15 today | $0.0807 | 36x |
| gpt-5.6-luna | frontier · reasoning | 11/12 2 days ago | 14/14 1 day ago | 15/15 today | 14/15 today | 13/15 5 days ago | 15/15 today | $0.0108 | 4.9x |
| gpt-4o | closed · mid-cost | 11/12 2 days ago | 14/14 1 day ago | 15/15 today | 13/15 today | 15/15 5 days ago | 14/15 today | $0.0610 | 28x |
| gpt-4.1-mini | closed · low-cost | 11/12 2 days ago | 14/14 1 day ago | 15/15 today | 14/15 today | 15/15 5 days ago | 13/15 today | $0.0087 | 3.9x |
| claude-opus-4-8 | closed · mid-cost | 10/12 unscored 2/12 2 days ago | 14/14 1 day ago | 15/15 today | 14/15 today | 15/15 5 days ago | 9/15 today | $0.4676 | 211x |
| claude-sonnet-5 | closed · mid-cost | 9/12 unscored 2/12 2 days ago | 14/14 1 day ago | 15/15 today | 13/15 today | 14/15 5 days ago | 15/15 today | $0.2855 | 129x |
| kimi-k3 | open-weight · mid-cost | 11/12 2 days ago | 13/14 1 day ago | 15/15 today | 12/15 today | 12/15 5 days ago | 15/15 today | $0.4321 | 195x |
| gemini-3.6-flash | closed · mid-cost | 10/12 2 days ago | 14/14 1 day ago | 15/15 today | 12/15 today | 15/15 5 days ago | 15/15 today | $0.2891 | 131x |
| claude-opus-5 | closed · mid-cost | 6/12 unscored 4/12 2 days ago | 14/14 1 day ago | 15/15 today | 9/15 today | 14/15 5 days ago | 13/15 today | $0.7764 | 351x |
Latest cycle: 2026-08-12. Each cell shows the date it was measured.
Cost per suite run is the combined cost of the most recent run of each of the six categories on the board (which may span cycles under rotation), counting scored cases only. Tokens consumed by refusals, provider policy blocks, and no-output cases are not included. Calculated from observed token counts and published list rates for the route TAB uses. It does not include prompt-caching, batch, or negotiated discounts, which vary by provider and by workload. See the methodology note dated July 29, 2026 below for an earlier period when the Agentic Execution category was excluded from this figure.
Methodology note, July 29, 2026. From the admission of the Agentic Execution category through July 29, 2026, that category’s model cost and token usage were not captured. During that period the Cost per suite run shown for every model excludes the Agentic Execution contribution and is understated by that amount. The judge cost was zero throughout and is unaffected, and scores are unaffected: this concerns cost attribution only, not measurement results. The gap is not reconstructable, because per-call token usage for that category was never recorded, so there is no basis to restate the earlier figures without estimating, and TAB does not publish estimates as measurements. All prior rows remain exactly as recorded. From July 29, 2026 forward, Agentic Execution cost is captured on the same basis as the other categories.
A model scored on only part of a category, below the coverage floor, is charged for only the cases it engaged, so its cost per suite run covers fewer cases than a model scored on the full set. Compare cost alongside the coverage shown for that cell.
Methodology note, August 6, 2026. Instruction Authority joined the public board on this date as a sixth category. It ran in nightly cadence from July 28, 2026 and was measured before publication; those internal-period results are not shown here and are not restated. Dated pages before this date show five categories, because five were public then, and that record is unchanged. Instruction Authority is scored on the same basis as the other five: fixed rules against known ground truth, judge cost zero. Its cross-category cross-check is still accruing baseline and is not yet applied to it.
Methodology note, August 8, 2026. claude-opus-5 (Anthropic) joined the public board on this date as the sixteenth measured model. It ran in nightly cadence from July 28, 2026 and was measured before publication; those internal-period results are not shown here and are not restated. Dated pages before this date show the fifteen models that were public then, and that record is unchanged. claude-opus-5 is scored on the same basis as the other models: fixed rules against known ground truth, judge cost zero.
The “vs. lowest” column shows each model’s cost per suite run as a multiple of the lowest cost among public models on the current board. It is a within-board cost comparison, not a value or quality comparison.
Each category is run at least three times per cycle under identical conditions. The scores shown are from the most recent run of each category.
◆ category leader.
Methodology note, August 10, 2026. "Lowest in field" and "category leader" compare the number of cases a model passed, not the fraction of the cases it was scored on. A model that declines cases in a category is scored on fewer cases than the others, so its passed-count is not directly comparable to the count of a model scored on every case. Where a model's unscored count is large, read the score alongside it. Dated pages are unchanged and show what was published on each date.
Methodology note, August 10, 2026. From this date, a category cell scored on fewer than half its cases shows its coverage instead of a score, and is excluded from the comparison that names the lowest and the leading model in a category. The floor is half the category's cases, rounded up, minimum two. Cells published before this date are unchanged.
Methodology note, August 11, 2026. From this date, a category cell scored on fewer than half its cases is also excluded from deviation testing. Such a cell is not compared against its own baseline and does not contribute to it, because a score computed on a small and changing number of cases carries variation from the count of cases scored, not only from the model's behaviour. Published scores and the dated pages are unchanged; this affects the deviation baseline only.
Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB previously recognised only one of them. From this date, a block signalled by xAI is recorded as a provider policy block rather than as a call that did not complete. The distinction matters because one describes the provider and the other describes TAB's own measurement failing. Cells published before this date are unchanged and show what was recorded then.
Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB has recognised those signals as it identified them. One such signal was not recognised before July 20, 2026, so blocks carrying it were recorded in that period as calls that did not complete rather than as provider policy blocks. The measurement did not change and the scores were not affected, only the description of why a case went unscored. Cells published in that period are unchanged and show what was recorded then.
A case is recorded as unscored — neither right nor wrong — when there is no scorable output. Unscored covers three distinct events: the model declined (a refusal, most often on security cases where its own safety classifier blocks the prompt before the model answers), the provider blocked the case before the model saw it (a provider policy block, not a model action), or the model call did not complete (a timeout or transport error, cause undetermined). This board shows only the total, labeled “unscored {n}/{total}”; the plain-language readout at /the-tab-test breaks it down by cause. Unscored cases neither inflate nor deflate the score, which reflects only the cases the model actually answered. When a cell is scored on fewer than half its cases it is shown as no score with its coverage, in place of a score, and is left out of the leader and lowest comparison for that category.
Category shape by model
Each model’s profile across the six categories. The outer edge is a full pass on every category.
gap = below coverage floor, not scored on enough cases to publish a comparable score.
The six categories
What each category tests. These describe the measurement, not any one day’s result.
- Security Screening: does the model resist attempts to turn it toward harmful use, leak data, or drop its safety refusals under pressure.
- Trust and Reliability: does the model give consistent, correctly formatted, correct answers, and does it stay consistent when the same thing is asked more than once.
- Agentic Execution: given a multi-step task with tools, does the model complete it correctly inside a sealed, controlled environment without taking actions it should not.
- Integrity and Provenance: using the case's supplied source material and allowed-source set, does the model quote and attribute accurately, and does it admit when the case provides no verifiable source instead of emitting citation-shaped references outside the allowed set. Deterministic checks against case-supplied ground truth; this instrument does not resolve URLs or DOIs or verify the real-world existence of a cited source.
- Resilience: does the model hold up under noisy, ambiguous, or adversarial input, rather than falling apart.
- Instruction Authority: when instructions conflict, does the model follow the one from the higher-authority source and refuse to be overridden by a lower one it merely encounters — an injected note, a retrieved document, a tool's output, or its own inference — instead of obeying whichever instruction it read last.
Cost against trust
Cost of one full pass (log scale) against the Provenance pass rate. Position is the fact; read it as you will.
How it is scored
Every verdict is computed by fixed, deterministic rules against known ground truth. There is no LLM judging the output, so the judge cost is zero and no model grades another model. A model that declines a case is recorded as unscored, never as a false zero. Every run is kept as a dated, immutable observation that is never overwritten.