Historical TAB Snapshot  ·  Originally published August 2026  ·  Archived record — not current TAB data  ·  Archive

INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-08-14

Deviation from baseline detected in 5 cells across 5 models this cycle — 5 above baseline. Detail below.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.
Measurement validity notice, September 1, 2026. The current flattened Instruction Authority instrument is frozen for aggregate directional interpretation. Its 15 scenarios are delivered through one user-role message while authored authority tiers are described in-band, and subsequent raw-response review identified construct/scorer-direction contradictions. Historical observations remain unchanged. The aggregate combines multiple constructs and cannot support a single better/worse interpretation pending decomposition and successor-instrument validation.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 14/15 on provenance, within repeat variation vs prior run 15/15, tied for the category lead.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).

deepseek-v3.2

deepseek-v3.2 scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-nano

gpt-4.1-nano scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-nano scored 12/15 on instruction authority, within repeat variation vs prior run 11/15.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

glm-5.2

glm-5.2 scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2 scored 15/15 on instruction authority, tied for the category lead.

claude-fable-5

claude-fable-5 scored 10/15 on provenance, within repeat variation vs prior run 12/15, lowest in field.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-fable-5 engaged with 14 of 15 instruction authority cases, 1 refused. Prior run: 15 of 15. 7 days ago: n/a (insufficient comparable history).

grok-4.5

grok-4.5 scored 14/15 on provenance, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

grok-4.5 scored 15/15 on instruction authority, tied for the category lead.

gpt-5.6-sol

gpt-5.6-sol scored 14/15 on provenance, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra

gpt-5.6-terra scored 14/15 on provenance, within repeat variation vs prior run 13/15, tied for the category lead.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 14/15 on instruction authority, within repeat variation vs prior run 15/15.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-luna

gpt-5.6-luna scored 13/15 on provenance, within repeat variation vs prior run 14/15.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 15/15 on instruction authority, tied for the category lead.

gpt-4o

gpt-4o scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-mini

gpt-4.1-mini scored 14/15 on provenance, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-4-8

claude-opus-4-8 scored 14/15 on provenance, within repeat variation vs prior run 15/15, tied for the category lead.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).

claude-opus-4-8 scored 10/15 on instruction authority, lowest in field.

Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5

claude-sonnet-5 scored 14/15 on provenance, within repeat variation vs prior run 13/15, tied for the category lead.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5 scored 14/15 on instruction authority.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

kimi-k3

kimi-k3 scored 11/15 on provenance.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

kimi-k3 scored 15/15 on instruction authority, tied for the category lead.

gemini-3.6-flash

gemini-3.6-flash scored 13/15 on provenance, within repeat variation vs prior run 12/15.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gemini-3.6-flash scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

claude-opus-5

claude-opus-5 scored 10/15 on provenance, within repeat variation vs prior run 11/15, lowest in field.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-5 scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).