INDEPENDENT MEASUREMENT • 16 MODELS • 6 CATEGORIES • UPDATED DAILY

The TAB Index

Independent measurement of what AI models actually do.

The comparison

All models across the six categories, today’s completed cycle. Green marks the category leader; unscored cases are shown separately and never counted as wrong answers.

ModelClassSecurity /12Trust /14Agentic /15Provenance /15Resilience /15Instruction Authority /15Cost per suite runvs. lowest
claude-haiku-4-5closed · low-cost11/12
2 days ago
13/14
1 day ago
15/15
today
14/15
today
15/15
5 days ago
13/15
today
$0.076134x
deepseek-v3.2open-weight · lowest-cost11/12
2 days ago
14/14
1 day ago
15/15
today
11/15
today
14/15
5 days ago
11/15
today
$0.01064.8x
gpt-4.1-nanoclosed · low-cost11/12
2 days ago
14/14
1 day ago
13/15
today
12/15
today
14/15
5 days ago
12/15
today
$0.00221.0x
glm-5.2open-weight · low-cost10/12
2 days ago
14/14
1 day ago
15/15
today
14/15
today
15/15
5 days ago
15/15
today
$0.075234x
claude-fable-5frontier · Mythos-classno score
4 of 12 scored, 8 unscored
2 days ago
14/14
1 day ago
15/15
today
10/15
today
14/15
unscored 1/15
5 days ago
12/15
unscored 3/15
today
$0.8048363x
grok-4.5closed · mid-cost11/12
2 days ago
14/14
1 day ago
15/15
today
15/15
today
14/15
5 days ago
15/15
today
$0.164774x
gpt-5.6-solfrontier · reasoning10/12
unscored 2/12
2 days ago
14/14
1 day ago
15/15
today
13/15
today
14/15
5 days ago
15/15
today
$0.187985x
gpt-5.6-terrafrontier · reasoning9/12
unscored 2/12
2 days ago
14/14
1 day ago
15/15
today
13/15
today
15/15
5 days ago
15/15
today
$0.080736x
gpt-5.6-lunafrontier · reasoning11/12
2 days ago
14/14
1 day ago
15/15
today
14/15
today
13/15
5 days ago
15/15
today
$0.01084.9x
gpt-4oclosed · mid-cost11/12
2 days ago
14/14
1 day ago
15/15
today
13/15
today
15/15
5 days ago
14/15
today
$0.061028x
gpt-4.1-miniclosed · low-cost11/12
2 days ago
14/14
1 day ago
15/15
today
14/15
today
15/15
5 days ago
13/15
today
$0.00873.9x
claude-opus-4-8closed · mid-cost10/12
unscored 2/12
2 days ago
14/14
1 day ago
15/15
today
14/15
today
15/15
5 days ago
9/15
today
$0.4676211x
claude-sonnet-5closed · mid-cost9/12
unscored 2/12
2 days ago
14/14
1 day ago
15/15
today
13/15
today
14/15
5 days ago
15/15
today
$0.2855129x
kimi-k3open-weight · mid-cost11/12
2 days ago
13/14
1 day ago
15/15
today
12/15
today
12/15
5 days ago
15/15
today
$0.4321195x
gemini-3.6-flashclosed · mid-cost10/12
2 days ago
14/14
1 day ago
15/15
today
12/15
today
15/15
5 days ago
15/15
today
$0.2891131x
claude-opus-5closed · mid-cost6/12
unscored 4/12
2 days ago
14/14
1 day ago
15/15
today
9/15
today
14/15
5 days ago
13/15
today
$0.7764351x
tabverified.ai

Latest cycle: 2026-08-12. Each cell shows the date it was measured.

Cost per suite run is the combined cost of the most recent run of each of the six categories on the board (which may span cycles under rotation), counting scored cases only. Tokens consumed by refusals, provider policy blocks, and no-output cases are not included. Calculated from observed token counts and published list rates for the route TAB uses. It does not include prompt-caching, batch, or negotiated discounts, which vary by provider and by workload. See the methodology note dated July 29, 2026 below for an earlier period when the Agentic Execution category was excluded from this figure.

Methodology note, July 29, 2026. From the admission of the Agentic Execution category through July 29, 2026, that category’s model cost and token usage were not captured. During that period the Cost per suite run shown for every model excludes the Agentic Execution contribution and is understated by that amount. The judge cost was zero throughout and is unaffected, and scores are unaffected: this concerns cost attribution only, not measurement results. The gap is not reconstructable, because per-call token usage for that category was never recorded, so there is no basis to restate the earlier figures without estimating, and TAB does not publish estimates as measurements. All prior rows remain exactly as recorded. From July 29, 2026 forward, Agentic Execution cost is captured on the same basis as the other categories.

A model scored on only part of a category, below the coverage floor, is charged for only the cases it engaged, so its cost per suite run covers fewer cases than a model scored on the full set. Compare cost alongside the coverage shown for that cell.

Methodology note, August 6, 2026. Instruction Authority joined the public board on this date as a sixth category. It ran in nightly cadence from July 28, 2026 and was measured before publication; those internal-period results are not shown here and are not restated. Dated pages before this date show five categories, because five were public then, and that record is unchanged. Instruction Authority is scored on the same basis as the other five: fixed rules against known ground truth, judge cost zero. Its cross-category cross-check is still accruing baseline and is not yet applied to it.

Methodology note, August 8, 2026. claude-opus-5 (Anthropic) joined the public board on this date as the sixteenth measured model. It ran in nightly cadence from July 28, 2026 and was measured before publication; those internal-period results are not shown here and are not restated. Dated pages before this date show the fifteen models that were public then, and that record is unchanged. claude-opus-5 is scored on the same basis as the other models: fixed rules against known ground truth, judge cost zero.

The “vs. lowest” column shows each model’s cost per suite run as a multiple of the lowest cost among public models on the current board. It is a within-board cost comparison, not a value or quality comparison.

Each category is run at least three times per cycle under identical conditions. The scores shown are from the most recent run of each category.

category leader.

Methodology note, August 10, 2026. "Lowest in field" and "category leader" compare the number of cases a model passed, not the fraction of the cases it was scored on. A model that declines cases in a category is scored on fewer cases than the others, so its passed-count is not directly comparable to the count of a model scored on every case. Where a model's unscored count is large, read the score alongside it. Dated pages are unchanged and show what was published on each date.

Methodology note, August 10, 2026. From this date, a category cell scored on fewer than half its cases shows its coverage instead of a score, and is excluded from the comparison that names the lowest and the leading model in a category. The floor is half the category's cases, rounded up, minimum two. Cells published before this date are unchanged.

Methodology note, August 11, 2026. From this date, a category cell scored on fewer than half its cases is also excluded from deviation testing. Such a cell is not compared against its own baseline and does not contribute to it, because a score computed on a small and changing number of cases carries variation from the count of cases scored, not only from the model's behaviour. Published scores and the dated pages are unchanged; this affects the deviation baseline only.

Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB previously recognised only one of them. From this date, a block signalled by xAI is recorded as a provider policy block rather than as a call that did not complete. The distinction matters because one describes the provider and the other describes TAB's own measurement failing. Cells published before this date are unchanged and show what was recorded then.

Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB has recognised those signals as it identified them. One such signal was not recognised before July 20, 2026, so blocks carrying it were recorded in that period as calls that did not complete rather than as provider policy blocks. The measurement did not change and the scores were not affected, only the description of why a case went unscored. Cells published in that period are unchanged and show what was recorded then.

A case is recorded as unscored — neither right nor wrong — when there is no scorable output. Unscored covers three distinct events: the model declined (a refusal, most often on security cases where its own safety classifier blocks the prompt before the model answers), the provider blocked the case before the model saw it (a provider policy block, not a model action), or the model call did not complete (a timeout or transport error, cause undetermined). This board shows only the total, labeled “unscored {n}/{total}”; the plain-language readout at /the-tab-test breaks it down by cause. Unscored cases neither inflate nor deflate the score, which reflects only the cases the model actually answered. When a cell is scored on fewer than half its cases it is shown as no score with its coverage, in place of a score, and is left out of the leader and lowest comparison for that category.

Category shape by model

Each model’s profile across the six categories. The outer edge is a full pass on every category.

claude-haiku-4-5
50%75%100%11/12Sec13/14Trust15/15Agent14/15Prov15/15Resil13/15Instrtabverified.ai
deepseek-v3.2
50%75%100%11/12Sec14/14Trust15/15Agent11/15Prov14/15Resil11/15Instrtabverified.ai
gpt-4.1-nano
50%75%100%11/12Sec14/14Trust13/15Agent12/15Prov14/15Resil12/15Instrtabverified.ai
glm-5.2
50%75%100%10/12Sec14/14Trust15/15Agent14/15Prov15/15Resil15/15Instrtabverified.ai
claude-fable-5
50%75%100%gapSec14/14Trust15/15Agent10/15Prov14/15Resil12/15Instrtabverified.ai
grok-4.5
50%75%100%11/12Sec14/14Trust15/15Agent15/15Prov14/15Resil15/15Instrtabverified.ai
gpt-5.6-sol
50%75%100%10/12Sec14/14Trust15/15Agent13/15Prov14/15Resil15/15Instrtabverified.ai
gpt-5.6-terra
50%75%100%9/12Sec14/14Trust15/15Agent13/15Prov15/15Resil15/15Instrtabverified.ai
gpt-5.6-luna
50%75%100%11/12Sec14/14Trust15/15Agent14/15Prov13/15Resil15/15Instrtabverified.ai
gpt-4o
50%75%100%11/12Sec14/14Trust15/15Agent13/15Prov15/15Resil14/15Instrtabverified.ai
gpt-4.1-mini
50%75%100%11/12Sec14/14Trust15/15Agent14/15Prov15/15Resil13/15Instrtabverified.ai
claude-opus-4-8
50%75%100%10/12Sec14/14Trust15/15Agent14/15Prov15/15Resil9/15Instrtabverified.ai
claude-sonnet-5
50%75%100%9/12Sec14/14Trust15/15Agent13/15Prov14/15Resil15/15Instrtabverified.ai
kimi-k3
50%75%100%11/12Sec13/14Trust15/15Agent12/15Prov12/15Resil15/15Instrtabverified.ai
gemini-3.6-flash
50%75%100%10/12Sec14/14Trust15/15Agent12/15Prov15/15Resil15/15Instrtabverified.ai
claude-opus-5
50%75%100%6/12Sec14/14Trust15/15Agent9/15Prov14/15Resil13/15Instrtabverified.ai

gap = below coverage floor, not scored on enough cases to publish a comparable score.

The six categories

What each category tests. These describe the measurement, not any one day’s result.

Cost against trust

Cost of one full pass (log scale) against the Provenance pass rate. Position is the fact; read it as you will.

0%25%50%75%100%$0.01$0.1$1Cost per full pass, log scale →Provenance pass rate →grok-4.515/15 · $0.1647claude-haiku-4-514/15 · $0.0761gpt-5.6-luna14/15 · $0.0108gpt-4.1-mini14/15 · $0.0087claude-opus-4-814/15 · $0.4676glm-5.214/15 · $0.0752gpt-5.6-sol13/15 · $0.1879gpt-5.6-terra13/15 · $0.0807gpt-4o13/15 · $0.0610claude-sonnet-513/15 · $0.2855gpt-4.1-nano12/15 · $0.0022kimi-k312/15 · $0.4321gemini-3.6-flash12/15 · $0.2891deepseek-v3.211/15 · $0.0106claude-fable-510/15 · $0.8048claude-opus-59/15 · $0.7764tabverified.ai

How it is scored

Every verdict is computed by fixed, deterministic rules against known ground truth. There is no LLM judging the output, so the judge cost is zero and no model grades another model. A model that declines a case is recorded as unscored, never as a false zero. Every run is kept as a dated, immutable observation that is never overwritten.

Read the full methodology →