Historical TAB Snapshot  ·  Originally published August 2026  ·  Archived record — not current TAB data  ·  Archive

INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-08-15

Deviation from baseline detected in 7 cells across 7 models this cycle — 4 above baseline, 3 below baseline. Detail below.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Source admission is a closed set of three authored cases, so a cell takes one of four values and one case moves it by a third of the subscale; it is a bounded subscale score on those three prompts, not a rate over general use and not an estimate of how often a model would do this elsewhere. In the September 3-4 2026 current-regime window - 38 model-by-cycle cells, three repeats each - repeated measurement of the same model differed by one of the three cases more often than not, so within that window a single-cycle difference of one or two cases is not a resolved difference between models. That estimate is one short window and is not a general constant. Interpreting differences across repeated cycles or by case recurrence requires a pre-specified rule, which has not yet been established. Where a HOLD is recorded, it withholds an acknowledgment pending review; it is a fail-closed measurement control and is not itself evidence about a model. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.
Measurement validity notice, September 1, 2026. The current flattened Instruction Authority instrument is frozen for aggregate directional interpretation. Its 15 scenarios are delivered through one user-role message while authored authority tiers are described in-band, and subsequent raw-response review identified construct/scorer-direction contradictions. Historical observations remain unchanged. The aggregate combines multiple constructs and cannot support a single better/worse interpretation pending decomposition and successor-instrument validation.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 15/15 on provenance, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

deepseek-v3.2

deepseek-v3.2 scored 13/15 on provenance, within repeat variation vs prior run 12/15.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-nano

gpt-4.1-nano scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-nano scored 12/15 on instruction authority, within repeat variation vs prior run 11/15.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

glm-5.2

glm-5.2 scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2 scored 15/15 on instruction authority, tied for the category lead.

claude-fable-5

claude-fable-5 engaged with 14 of 15 provenance cases, 1 refused. Prior run: 15 of 15. 7 days ago: n/a (insufficient comparable history).

claude-fable-5 engaged with 13 of 15 instruction authority cases, 2 refused. Prior run: 15 of 15. 7 days ago: n/a (insufficient comparable history).

grok-4.5

grok-4.5 scored 15/15 on provenance, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

grok-4.5 scored 15/15 on instruction authority, tied for the category lead.

gpt-5.6-sol

gpt-5.6-sol scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-sol scored 14/15 on instruction authority, within repeat variation vs prior run 15/15.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-terra

gpt-5.6-terra scored 13/15 on provenance, within repeat variation vs prior run 14/15.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-luna

gpt-5.6-luna scored 14/15 on provenance, within repeat variation vs prior run 12/15.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 15/15 on instruction authority, tied for the category lead.

gpt-4o

gpt-4o scored 13/15 on provenance.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-mini

gpt-4.1-mini scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-4-8

claude-opus-4-8 scored 15/15 on provenance, tied for the category lead.

claude-opus-4-8 scored 9/15 on instruction authority, within repeat variation vs prior run 10/15, lowest in field.

Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5

claude-sonnet-5 scored 13/15 on provenance, within repeat variation vs prior run 12/15.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

claude-sonnet-5 scored 14/15 on instruction authority.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

kimi-k3

kimi-k3 scored 11/15 on provenance, within repeat variation vs prior run 13/15.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gemini-3.6-flash

gemini-3.6-flash scored 12/15 on provenance.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gemini-3.6-flash scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

claude-opus-5

claude-opus-5 scored 10/15 on provenance, lowest in field.

Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-opus-5 scored 13/15 on instruction authority, within repeat variation vs prior run 15/15.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).