INDEPENDENT MEASUREMENT • 9 MODELS • 5 CATEGORIES • UPDATED DAILY

The TAB Index

Independent measurement of what AI models actually do.

ModelClassSecurity /12Trust /14Agentic /15Provenance /15Resilience /15Cost/pass
claude-haiku-4-5closed · low-cost11/1213/1415/1514/1515/15$0.0138
deepseek-v3.2open-weight · lowest-cost10/1213/1415/1512/1513/15$0.0028
gpt-4.1-nanoclosed · low-cost11/1214/1413/1512/1514/15$0.0009
glm-5.2open-weight · low-cost10/1214/1415/1514/15
refused 1/15 (unscored)
14/15$0.0548
claude-fable-5frontier · Mythos-class4/12
refused 8/12 (unscored)
14/1415/1512/1515/15$0.2719
grok-4.5closed · mid-cost10/12
refused 1/12 (unscored)
14/1415/1514/1514/15$0.0370
gpt-5.6-solfrontier · reasoning9/12
refused 3/12 (unscored)
14/1415/1514/1514/15$0.0894
gpt-5.6-terrafrontier · reasoning9/12
refused 3/12 (unscored)
14/1415/1513/1514/15$0.0416
gpt-5.6-lunafrontier · reasoning12/1214/1415/1514/1514/15$0.0278

Measured 2026-07-20

category leader. A refusal is recorded as unscored, never as a wrong answer.

The five categories

What each category tests. These describe the measurement, not any one day’s result.

How it is scored

Every verdict is computed by fixed, deterministic rules against known ground truth. There is no LLM judging the output, so the judge cost is zero and no model grades another model. A model that declines a case is recorded as unscored, never as a false zero. Every run is kept as a dated, immutable observation that is never overwritten.

Read the full methodology →