Historical TAB Snapshot  ·  Originally published August 2026  ·  Archived record — not current TAB data  ·  Archive

INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-08-09

Deviation from baseline detected in 7 cells across 7 models this cycle — 4 above baseline, 3 below baseline. Detail below.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

deepseek-v3.2

deepseek-v3.2 scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

deepseek-v3.2 scored 11/15 on instruction authority, lowest in field.

gpt-4.1-nano

gpt-4.1-nano scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2

glm-5.2 scored 13/15 on provenance, down from 14/15 its prior run.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2 scored 15/15 on instruction authority, tied for the category lead.

claude-fable-5

claude-fable-5 scored 11/15 on provenance, down from 12/15 its prior run.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-fable-5 engaged with 11 of 15 instruction authority cases, 4 refused. Prior run: 14 of 15. 7 days ago: n/a (insufficient comparable history).

grok-4.5

grok-4.5 scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

grok-4.5 scored 15/15 on instruction authority, tied for the category lead.

gpt-5.6-sol

gpt-5.6-sol scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-sol scored 15/15 on instruction authority, tied for the category lead.

gpt-5.6-terra

gpt-5.6-terra scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 15/15 on instruction authority, up from 14/15 its prior run, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-luna

gpt-5.6-luna scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 15/15 on instruction authority, tied for the category lead.

gpt-4o

gpt-4o scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4o scored 14/15 on instruction authority, up from 13/15 its prior run.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

gpt-4.1-mini

gpt-4.1-mini scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-4-8

claude-opus-4-8 scored 15/15 on provenance, the category leader.

claude-opus-4-8 scored 12/15 on instruction authority, up from 10/15 its prior run.

Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5

claude-sonnet-5 scored 14/15 on provenance, up from 12/15 its prior run.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5 scored 14/15 on instruction authority.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

kimi-k3

kimi-k3 scored 13/15 on provenance, down from 14/15 its prior run.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

kimi-k3 scored 15/15 on instruction authority, up from 14/15 its prior run, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

gemini-3.6-flash

gemini-3.6-flash scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gemini-3.6-flash scored 14/15 on instruction authority, down from 15/15 its prior run.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).

claude-opus-5

claude-opus-5 scored 10/15 on provenance, down from 11/15 its prior run, lowest in field.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-5 scored 13/15 on instruction authority, up from 11/15 its prior run.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).