Independent Measurement • 9 Models • 5 Categories • One Identical Case Set
The frontier premium won nothing, and the rankings invert by property. Nine models, five property categories, one identical set of cases. The most expensive model tested led no category outright, no model leads everywhere, and on the properties that matter most the scores move from one run to the next.
One representative run per model on each of the five TAB Index categories. Scores are raw checks passed out of the checks available in each category. Category leaders are marked; ties are marked on every model that reached the top score.
| Model | Class | Security /12 | Trust /14 | Agentic /15 | Provenance /15 | Resilience /15 | Cost/run |
|---|---|---|---|---|---|---|---|
| claude-haiku-4-5 | closed · low-cost | 11 | 13 | 15 | 14 | 15 | $0.0028 |
| deepseek-v3.2 | open-weight · lowest-cost | 10 | 14 | 15 | 12 | 13 | $0.0005 |
| gpt-4.1-nano | closed · low-cost | 11 | 14 | 13 | 12 | 14 | $0.0002 |
| glm-5.2 | open-weight · low-cost | 10 | 14 | 15 | 15 | 15 | $0.0110 |
| claude-fable-5 | frontier · Mythos-class | refused 8/12 (unscored) | 14 | 15 | 11 | 14 | $0.0567 |
| grok-4.5 | closed · mid-cost | 10 | 14 | 15 | 15 | 15 | $0.0087 |
| gpt-5.6-sol | frontier · reasoning | 9 | 14 | 15 | 12 | 14 | $0.0166 |
| gpt-5.6-terra | frontier · reasoning | 9 | 14 | 15 | 14 | 13 | $0.0089 |
| gpt-5.6-luna | frontier · reasoning | 12 | 14 | 15 | 13 | 14 | $0.0061 |
◆ marks the leading score in a category; where models tie for the lead, each is marked. claude-fable-5's Security cell is a refusal: its safety classifier declined 8 of 12 cases, so the category is recorded as unscored, not as a zero the model never produced.
Each finding reads directly off the run above. The numbers are measured, not modeled.
The Mythos-class model, at more than 300x the cheapest model's per-run cost, led no category outright. Its single best result is a shared tie on Agentic that several cheaper models also reached.
DeepSeek ties for the top on Trust and Agentic yet sits near the bottom on Provenance. GLM 5.2 (open-weight, roughly one fifth the priciest model's cost) ties for the lead on four of five categories and is held back only by Security. No single model leads everywhere, and no model sits at the bottom everywhere. A composite "best model" number averages that inversion away and hands the buyer the wrong answer for the property their business runs on.
On Integrity and Provenance, Anthropic's Claude Fable 5, the priciest model we test at roughly $0.057 per run, scored lowest of all nine models. It was beaten on provenance by models a fraction of its cost, including gpt-4.1-nano at $0.0002 per run, roughly 280 times cheaper. Two models scored a flawless 15 out of 15 in this category, and neither was the most expensive. Provenance decides whether a model cites real, supporting sources instead of inventing them, the property that matters most for legal, financial, and medical work, and here price bought the opposite of what you would expect.
The frontier model's safety classifier declined most Security cases. TAB records that as unscored and reports it plainly, rather than assigning a number the model never produced.
Model outputs vary run to run, so a single score is a snapshot, not a verdict. Because every run is dated and kept rather than overwritten, that variance is itself visible in the record.
The measurement is reproducible and the scoring is inspectable. Nothing here depends on a model grading another model.