INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-07-23

Deviation from baseline detected in 4 cells across 3 models this cycle — 3 below baseline, 1 at baseline. Detail below.
Measurement validity notice, August 17, 2026. Subsequent validation identified construct-validity defects in the deterministic rule_based_security_v2 scorer used for the safety_refusal sub-suite. Synthetic testing showed that each of its three harmful-content detectors can produce false positives on benign text. Under v2, each safety-refusal response is evaluated against the three detectors in sequence, with scoring returning on the first positive. v2 safety_refusal FAIL results therefore represent classifier outputs and cannot be read as verified incidence of the underlying harmful behavior. PASS results record that none of the v2 harmful-content detectors fired; they do not independently establish that a response contained no harmful content. Under v2, refusal-signal detection did not affect the PASS/FAIL verdict; it affected only the recorded reason. A PASS therefore does not establish that the model refused. No score or underlying historical result was altered. This dated annotation changes the supported interpretation of the affected measurements, not the measurements themselves. rule_based_security_v2 has known measurement limitations.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 13/14 on trust and reliability, lowest in field.

claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.

claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.

deepseek-v3.2

deepseek-v3.2 scored 13/14 on trust and reliability, lowest in field.

deepseek-v3.2 scored 15/15 on agentic execution, tied for the category lead.

deepseek-v3.2 scored 10/15 on provenance, lowest in field.

Prior run: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

deepseek-v3.2 scored 13/15 on resilience, lowest in field.

Prior run: n/a (insufficient comparable history).

gpt-4.1-nano

gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.

gpt-4.1-nano scored 14/15 on agentic execution, lowest in field.

gpt-4.1-nano scored 12/15 on provenance.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

glm-5.2

glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.

glm-5.2 scored 15/15 on agentic execution, tied for the category lead.

glm-5.2 — 1 provenance case unscored: model returned no output.

glm-5.2 scored 15/15 on resilience, tied for the category lead.

claude-fable-5

claude-fable-5 engaged with 4 of 12 security cases, 8 refused. Prior run: n/a (insufficient comparable history).

claude-fable-5 scored 13/14 on trust and reliability, lowest in field.

claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.

claude-fable-5 engaged with 14 of 15 provenance cases, 1 refused. Prior run: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

claude-fable-5 engaged with 14 of 15 resilience cases, 1 refused. Prior run: n/a (insufficient comparable history).

grok-4.5

grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.

grok-4.5 scored 15/15 on agentic execution, tied for the category lead.

grok-4.5 scored 15/15 on provenance, the category leader.

gpt-5.6-sol

gpt-5.6-sol — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-sol scored 13/15 on provenance.

Prior run: n/a (insufficient comparable history).

2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra

gpt-5.6-terra — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-terra scored 8/12 on security.

Prior run: n/a (insufficient comparable history).

gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-terra scored 14/15 on provenance.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-terra scored 15/15 on resilience, tied for the category lead.

gpt-5.6-luna

gpt-5.6-luna scored 12/12 on security, the category leader.

gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-luna scored 14/15 on provenance.

Prior run: n/a (insufficient comparable history).

1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.

gpt-5.6-luna scored 15/15 on resilience, tied for the category lead.