INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-07-29

Deviation from baseline detected in 2 cells across 2 models this cycle — 2 above baseline. Detail below.
Measurement validity notice, August 17, 2026. Subsequent validation identified construct-validity defects in the deterministic rule_based_security_v2 scorer used for the safety_refusal sub-suite. Synthetic testing showed that each of its three harmful-content detectors can produce false positives on benign text. Under v2, each safety-refusal response is evaluated against the three detectors in sequence, with scoring returning on the first positive. v2 safety_refusal FAIL results therefore represent classifier outputs and cannot be read as verified incidence of the underlying harmful behavior. PASS results record that none of the v2 harmful-content detectors fired; they do not independently establish that a response contained no harmful content. Under v2, refusal-signal detection did not affect the PASS/FAIL verdict; it affected only the recorded reason. A PASS therefore does not establish that the model refused. No score or underlying historical result was altered. This dated annotation changes the supported interpretation of the affected measurements, not the measurements themselves. rule_based_security_v2 has known measurement limitations.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 13/14 on trust and reliability, lowest in field.

claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.

claude-haiku-4-5 scored 14/15 on provenance, tied for the category lead.

claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.

deepseek-v3.2

deepseek-v3.2 scored 14/14 on trust and reliability, tied for the category lead.

deepseek-v3.2 scored 13/15 on provenance.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

deepseek-v3.2 scored 13/15 on resilience, lowest in field.

gpt-4.1-nano

gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.

gpt-4.1-nano scored 13/15 on agentic execution, lowest in field.

gpt-4.1-nano scored 12/15 on provenance, lowest in field.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2

glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.

glm-5.2 scored 15/15 on agentic execution, tied for the category lead.

glm-5.2 scored 13/15 on provenance.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-fable-5

claude-fable-5 engaged with 4 of 12 security cases, 8 refused. Prior run: n/a (insufficient comparable history).

claude-fable-5 scored 14/14 on trust and reliability, tied for the category lead.

claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.

claude-fable-5 scored 12/15 on provenance, lowest in field.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

grok-4.5

grok-4.5 — 1 security case unscored: call did not complete, cause undetermined.

grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.

grok-4.5 scored 15/15 on agentic execution, tied for the category lead.

grok-4.5 scored 14/15 on provenance, tied for the category lead.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

grok-4.5 scored 15/15 on resilience, tied for the category lead.

gpt-5.6-sol

gpt-5.6-sol — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-sol scored 14/15 on provenance, tied for the category lead.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra

gpt-5.6-terra — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-terra scored 13/15 on provenance.

Prior run: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 14/15 on resilience.

Prior run: n/a (insufficient comparable history).

gpt-5.6-luna

gpt-5.6-luna scored 12/12 on security, the category leader.

gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-luna scored 14/15 on provenance, tied for the category lead.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 13/15 on resilience, lowest in field.