INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY
TAB Daily Results
2026-07-29
By model
What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.
claude-haiku-4-5
claude-haiku-4-5 scored 13/14 on trust and reliability, lowest in field.
claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.
claude-haiku-4-5 scored 14/15 on provenance, tied for the category lead.
claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.
deepseek-v3.2
deepseek-v3.2 scored 14/14 on trust and reliability, tied for the category lead.
deepseek-v3.2 scored 13/15 on provenance.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
deepseek-v3.2 scored 13/15 on resilience, lowest in field.
gpt-4.1-nano
gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.
gpt-4.1-nano scored 13/15 on agentic execution, lowest in field.
gpt-4.1-nano scored 12/15 on provenance, lowest in field.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
glm-5.2
glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.
glm-5.2 scored 15/15 on agentic execution, tied for the category lead.
glm-5.2 scored 13/15 on provenance.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
claude-fable-5
claude-fable-5 engaged with 4 of 12 security cases, 8 refused. Prior run: n/a (insufficient comparable history).
claude-fable-5 scored 14/14 on trust and reliability, tied for the category lead.
claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.
claude-fable-5 scored 12/15 on provenance, lowest in field.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
grok-4.5
grok-4.5 — 1 security case unscored: call did not complete, cause undetermined.
grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.
grok-4.5 scored 15/15 on agentic execution, tied for the category lead.
grok-4.5 scored 14/15 on provenance, tied for the category lead.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
grok-4.5 scored 15/15 on resilience, tied for the category lead.
gpt-5.6-sol
gpt-5.6-sol — 3 security cases unscored: provider policy block, not a model failure.
gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-sol scored 14/15 on provenance, tied for the category lead.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-terra
gpt-5.6-terra — 3 security cases unscored: provider policy block, not a model failure.
gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-terra scored 13/15 on provenance.
Prior run: n/a (insufficient comparable history).
2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-terra scored 14/15 on resilience.
Prior run: n/a (insufficient comparable history).
gpt-5.6-luna
gpt-5.6-luna scored 12/12 on security, the category leader.
gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-luna scored 14/15 on provenance, tied for the category lead.
Prior run: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-luna scored 13/15 on resilience, lowest in field.