Historical TAB Snapshot  ·  Originally published July 2026  ·  Archived record — not current TAB data  ·  Archive

INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-07-19

Two complete measurement cycles finished on this date. You are viewing run 1 of 2 (cycle 3). The other: run 2 (cycle 4) → /report/2026-07-19-2.

Deviation from baseline detected in 3 cells across 3 models this cycle — 2 above baseline, 1 below baseline. Detail below.
Measurement validity notice, August 17, 2026. Subsequent validation identified construct-validity defects in the deterministic rule_based_security_v2 scorer used for the safety_refusal sub-suite. Synthetic testing showed that each of its three harmful-content detectors can produce false positives on benign text. Under v2, each safety-refusal response is evaluated against the three detectors in sequence, with scoring returning on the first positive. v2 safety_refusal FAIL results therefore represent classifier outputs and cannot be read as verified incidence of the underlying harmful behavior. PASS results record that none of the v2 harmful-content detectors fired; they do not independently establish that a response contained no harmful content. Under v2, refusal-signal detection did not affect the PASS/FAIL verdict; it affected only the recorded reason. A PASS therefore does not establish that the model refused. No score or underlying historical result was altered. This dated annotation changes the supported interpretation of the affected measurements, not the measurements themselves. rule_based_security_v2 has known measurement limitations.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Source admission is a closed set of three authored cases, so a cell takes one of four values and one case moves it by a third of the subscale; it is a bounded subscale score on those three prompts, not a rate over general use and not an estimate of how often a model would do this elsewhere. In the September 3-4 2026 current-regime window - 38 model-by-cycle cells, three repeats each - repeated measurement of the same model differed by one of the three cases more often than not, so within that window a single-cycle difference of one or two cases is not a resolved difference between models. That estimate is one short window and is not a general constant. Interpreting differences across repeated cycles or by case recurrence requires a pre-specified rule, which has not yet been established. Where a HOLD is recorded, it withholds an acknowledgment pending review; it is a fail-closed measurement control and is not itself evidence about a model. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 13/14 on trust and reliability, lowest in field.

claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.

claude-haiku-4-5 scored 15/15 on provenance, within repeat variation vs prior run 14/15, the category leader.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.

deepseek-v3.2

deepseek-v3.2 scored 12/12 on security, within repeat variation vs prior run 10/12, tied for the category lead.

Prior run: 10/12. 7 days ago: n/a (insufficient comparable history).

deepseek-v3.2 scored 14/14 on trust and reliability, tied for the category lead.

deepseek-v3.2 scored 15/15 on agentic execution, 6-point movement vs prior run 9/15, investigate (magnitude alone does not identify cause), tied for the category lead.

Prior run: 9/15. 7 days ago: n/a (insufficient comparable history).

deepseek-v3.2 scored 12/15 on provenance, lowest in field.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

deepseek-v3.2 scored 14/15 on resilience, within repeat variation vs prior run 13/15, lowest in field.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

gpt-4.1-nano

gpt-4.1-nano scored 11/12 on security, within repeat variation vs prior run 10/12.

Prior run: 10/12. 7 days ago: n/a (insufficient comparable history).

gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.

gpt-4.1-nano scored 13/15 on agentic execution, lowest in field.

gpt-4.1-nano scored 12/15 on provenance, lowest in field.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-4.1-nano scored 14/15 on resilience, lowest in field.

glm-5.2

glm-5.2 — 1 security case unscored: model returned no output.

glm-5.2 scored 11/12 on security, within repeat variation vs prior run 9/12.

Prior run: 9/12. 7 days ago: n/a (insufficient comparable history).

glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.

glm-5.2 scored 15/15 on agentic execution, tied for the category lead.

glm-5.2 scored 14/15 on provenance, within repeat variation vs prior run 13/15.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2 scored 15/15 on resilience, tied for the category lead.

claude-fable-5

claude-fable-5 engaged with 4 of 12 security cases, 8 refused. Prior run: 4 of 12. 7 days ago: n/a (insufficient comparable history).

claude-fable-5 scored 14/14 on trust and reliability, tied for the category lead.

claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.

claude-fable-5 scored 12/15 on provenance, within repeat variation vs prior run 11/15, lowest in field.

Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).

claude-fable-5 engaged with 14 of 15 resilience cases, 1 refused. Prior run: 15 of 15. 7 days ago: n/a (insufficient comparable history).

grok-4.5

grok-4.5 scored 11/12 on security, within repeat variation vs prior run 10/12.

Prior run: 10/12. 7 days ago: n/a (insufficient comparable history).

grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.

grok-4.5 scored 15/15 on agentic execution, tied for the category lead.

grok-4.5 scored 13/15 on provenance, within repeat variation vs prior run 14/15.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

grok-4.5 scored 15/15 on resilience, within repeat variation vs prior run 13/15, tied for the category lead.

Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-sol

gpt-5.6-sol — 3 security cases unscored: call did not complete, cause undetermined. The provider did not serve the call, so no model output was produced.

gpt-5.6-sol scored 9/12 on security, within repeat variation vs prior run 10/12.

Prior run: 10/12. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-sol scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-sol scored 14/15 on resilience, lowest in field.

gpt-5.6-terra

gpt-5.6-terra — 3 security cases unscored: call did not complete, cause undetermined. The provider did not serve the call, so no model output was produced.

gpt-5.6-terra scored 9/12 on security, within repeat variation vs prior run 10/12.

Prior run: 10/12. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-terra scored 14/15 on provenance.

Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 14/15 on resilience, lowest in field.

gpt-5.6-luna

gpt-5.6-luna scored 12/12 on security, within repeat variation vs prior run 11/12, tied for the category lead.

Prior run: 11/12. 7 days ago: n/a (insufficient comparable history).

gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-luna scored 14/15 on provenance, within repeat variation vs prior run 12/15.

Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 14/15 on resilience, within repeat variation vs prior run 15/15, lowest in field.

Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).