INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY
TAB Daily Results
2026-08-02
By model
What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.
claude-haiku-4-5
claude-haiku-4-5 scored 13/14 on trust and reliability, lowest in field.
claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.
claude-haiku-4-5 scored 14/15 on provenance, tied for the category lead.
claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.
Consistency this cycle: 4 of 5 categories produced the same result across all runs.
deepseek-v3.2
deepseek-v3.2 scored 10/12 on security, down from 11/12 the prior cycle.
Prior cycle: 11/12. 7 days ago: 11/12. 30 days ago: n/a (insufficient history).
deepseek-v3.2 scored 14/14 on trust and reliability, tied for the category lead.
deepseek-v3.2 scored 15/15 on agentic execution, tied for the category lead.
deepseek-v3.2 scored 13/15 on provenance, up from 12/15 the prior cycle.
Prior cycle: 12/15. 7 days ago: 13/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
deepseek-v3.2 scored 14/15 on resilience.
Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 3 of 5 categories produced the same result across all runs.
gpt-4.1-nano
gpt-4.1-nano scored 10/12 on security, down from 11/12 the prior cycle.
Prior cycle: 11/12. 7 days ago: 10/12. 30 days ago: n/a (insufficient history).
gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.
gpt-4.1-nano scored 14/15 on agentic execution, lowest in field.
gpt-4.1-nano scored 12/15 on provenance, lowest in field.
Prior cycle: 12/15. 7 days ago: 12/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
Consistency this cycle: 5 of 5 categories produced the same result across all runs.
glm-5.2
glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.
glm-5.2 scored 15/15 on agentic execution, up from 14/15 the prior cycle, tied for the category lead.
Prior cycle: 14/15. 7 days ago: 15/15. 30 days ago: n/a (insufficient history).
glm-5.2 scored 13/15 on provenance, down from 14/15 the prior cycle.
Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
glm-5.2 scored 14/15 on resilience, down from 15/15 the prior cycle.
Prior cycle: 15/15. 7 days ago: 15/15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 2 of 5 categories produced the same result across all runs.
claude-fable-5
claude-fable-5 engaged with 3 of 12 security cases, 9 refused. Prior cycle: 4 of 12. 7 days ago: 4 of 12. 30 days ago: n/a (insufficient history).
claude-fable-5 scored 14/14 on trust and reliability, tied for the category lead.
claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.
claude-fable-5 scored 12/15 on provenance, lowest in field.
claude-fable-5 engaged with 14 of 15 resilience cases, 1 refused. Prior cycle: 14 of 15. 7 days ago: 15 of 15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 3 of 5 categories produced the same result across all runs.
grok-4.5
grok-4.5 — 1 security case unscored: model returned no output.
grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.
grok-4.5 scored 15/15 on agentic execution, tied for the category lead.
grok-4.5 scored 14/15 on provenance, down from 15/15 the prior cycle, tied for the category lead.
Prior cycle: 15/15. 7 days ago: 15/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
grok-4.5 scored 15/15 on resilience, tied for the category lead.
Consistency this cycle: 5 of 5 categories produced the same result across all runs.
gpt-5.6-sol
gpt-5.6-sol — 3 security cases unscored: provider policy block, not a model failure.
gpt-5.6-sol scored 9/12 on security, down from 10/12 the prior cycle.
Prior cycle: 10/12. 7 days ago: 9/12. 30 days ago: n/a (insufficient history).
gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-sol scored 14/15 on provenance, up from 13/15 the prior cycle, tied for the category lead.
Prior cycle: 13/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
Consistency this cycle: 2 of 5 categories produced the same result across all runs.
gpt-5.6-terra
gpt-5.6-terra — 3 security cases unscored: provider policy block, not a model failure.
gpt-5.6-terra scored 9/12 on security, down from 10/12 the prior cycle.
Prior cycle: 10/12. 7 days ago: 9/12. 30 days ago: n/a (insufficient history).
gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-terra scored 14/15 on provenance, up from 13/15 the prior cycle, tied for the category lead.
Prior cycle: 13/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
gpt-5.6-terra scored 14/15 on resilience.
Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 3 of 5 categories produced the same result across all runs.
gpt-5.6-luna
gpt-5.6-luna scored 12/12 on security, tied for the category lead.
gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.
gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.
gpt-5.6-luna scored 14/15 on provenance, tied for the category lead.
Prior cycle: 14/15. 7 days ago: 13/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
gpt-5.6-luna scored 13/15 on resilience, down from 15/15 the prior cycle, lowest in field.
Prior cycle: 15/15. 7 days ago: 13/15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 4 of 5 categories produced the same result across all runs.
gpt-4o
gpt-4o scored 14/14 on trust and reliability, tied for the category lead.
gpt-4o scored 15/15 on agentic execution, tied for the category lead.
gpt-4o scored 13/15 on provenance.
Prior cycle: 13/15. 7 days ago: n/a (insufficient history).
1 fabricated citation detected.
gpt-4o scored 15/15 on resilience, tied for the category lead.
Consistency this cycle: 5 of 5 categories produced the same result across all runs.
gpt-4.1-mini
gpt-4.1-mini scored 14/14 on trust and reliability, tied for the category lead.
gpt-4.1-mini scored 15/15 on agentic execution, tied for the category lead.
gpt-4.1-mini scored 14/15 on provenance, tied for the category lead.
Prior cycle: 14/15. 7 days ago: n/a (insufficient history).
1 fabricated citation detected.
gpt-4.1-mini scored 15/15 on resilience, tied for the category lead.
Consistency this cycle: 5 of 5 categories produced the same result across all runs.
claude-opus-4-8
claude-opus-4-8 engaged with 10 of 12 security cases, 2 refused. Prior cycle: 10 of 12. 7 days ago: n/a (insufficient history).
claude-opus-4-8 scored 14/14 on trust and reliability, tied for the category lead.
claude-opus-4-8 scored 15/15 on agentic execution, tied for the category lead.
claude-opus-4-8 scored 14/15 on provenance, down from 15/15 the prior cycle, tied for the category lead.
Prior cycle: 15/15. 7 days ago: n/a (insufficient history).
claude-opus-4-8 scored 15/15 on resilience, tied for the category lead.
Consistency this cycle: 2 of 5 categories produced the same result across all runs.
claude-sonnet-5
claude-sonnet-5 engaged with 10 of 12 security cases, 2 refused. Prior cycle: 10 of 12. 7 days ago: n/a (insufficient history).
claude-sonnet-5 scored 14/14 on trust and reliability, tied for the category lead.
claude-sonnet-5 scored 15/15 on agentic execution, tied for the category lead.
claude-sonnet-5 scored 14/15 on provenance, up from 12/15 the prior cycle, tied for the category lead.
Prior cycle: 12/15. 7 days ago: n/a (insufficient history).
Consistency this cycle: 4 of 5 categories produced the same result across all runs.
kimi-k3
kimi-k3 scored 12/12 on security, up from 11/12 the prior cycle, tied for the category lead.
Prior cycle: 11/12. 7 days ago: 0/12. 30 days ago: n/a (insufficient history).
kimi-k3 scored 14/14 on trust and reliability, tied for the category lead.
kimi-k3 scored 15/15 on agentic execution, tied for the category lead.
kimi-k3 scored 13/15 on provenance, up from 12/15 the prior cycle.
Prior cycle: 12/15. 7 days ago: 0/15. 30 days ago: n/a (insufficient history).
1 fabricated citation detected.
Consistency this cycle: 2 of 5 categories produced the same result across all runs.
gemini-3.6-flash
gemini-3.6-flash scored 8/12 on security, down from 10/12 the prior cycle.
Prior cycle: 10/12. 7 days ago: 9/12. 30 days ago: n/a (insufficient history).
gemini-3.6-flash scored 14/14 on trust and reliability, tied for the category lead.
gemini-3.6-flash scored 15/15 on agentic execution, tied for the category lead.
gemini-3.6-flash scored 12/15 on provenance, lowest in field.
Prior cycle: 12/15. 7 days ago: 12/15. 30 days ago: n/a (insufficient history).
2 fabricated citation results detected.
gemini-3.6-flash scored 14/15 on resilience, up from 13/15 the prior cycle.
Prior cycle: 13/15. 7 days ago: 15/15. 30 days ago: n/a (insufficient history).
Consistency this cycle: 2 of 5 categories produced the same result across all runs.