INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY

TAB Daily Results

2026-07-28

Threshold breach detected in 2 models this cycle. Detail below.

By model

What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.

claude-haiku-4-5

claude-haiku-4-5 scored 15/15 on agentic execution, tied for the category lead.

claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.

Consistency this cycle: 5 of 5 categories produced the same result across all runs.

deepseek-v3.2

deepseek-v3.2 scored 10/12 on security, down from 11/12 the prior cycle.

Prior cycle: 11/12. 7 days ago: 11/12. 30 days ago: n/a (insufficient history).

deepseek-v3.2 scored 13/14 on trust and reliability, down from 14/14 the prior cycle.

Prior cycle: 14/14. 7 days ago: 14/14. 30 days ago: n/a (insufficient history).

deepseek-v3.2 scored 15/15 on agentic execution, tied for the category lead.

deepseek-v3.2 scored 11/15 on provenance, down from 12/15 the prior cycle.

Prior cycle: 12/15. 7 days ago: 12/15. 30 days ago: n/a (insufficient history).

2 fabricated citation results detected.

Consistency this cycle: 2 of 5 categories produced the same result across all runs.

gpt-4.1-nano

gpt-4.1-nano scored 14/14 on trust and reliability, tied for the category lead.

gpt-4.1-nano scored 12/15 on provenance.

Prior cycle: 12/15. 7 days ago: 12/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

glm-5.2

glm-5.2 scored 11/12 on security, up from 10/12 the prior cycle.

Prior cycle: 10/12. 7 days ago: 11/12. 30 days ago: n/a (insufficient history).

glm-5.2 scored 14/14 on trust and reliability, tied for the category lead.

glm-5.2 scored 15/15 on agentic execution, up from 14/15 the prior cycle, tied for the category lead.

Prior cycle: 14/15. 7 days ago: 15/15. 30 days ago: n/a (insufficient history).

glm-5.2 scored 15/15 on provenance, up from 14/15 the prior cycle, the category leader.

Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

glm-5.2 scored 15/15 on resilience, tied for the category lead.

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

claude-fable-5

claude-fable-5 engaged with 4 of 12 security cases, 8 refused. Prior cycle: 4 of 12. 7 days ago: 3 of 12. 30 days ago: n/a (insufficient history).

claude-fable-5 scored 14/14 on trust and reliability, tied for the category lead.

claude-fable-5 scored 15/15 on agentic execution, tied for the category lead.

claude-fable-5 scored 11/15 on provenance, down from 12/15 the prior cycle.

Prior cycle: 12/15. 7 days ago: 12/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

claude-fable-5 engaged with 14 of 15 resilience cases, 1 refused. Prior cycle: 14 of 15. 7 days ago: 14 of 15. 30 days ago: n/a (insufficient history).

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

grok-4.5

grok-4.5 scored 11/12 on security, up from 10/12 the prior cycle.

Prior cycle: 10/12. 7 days ago: 10/12. 30 days ago: n/a (insufficient history).

grok-4.5 scored 14/14 on trust and reliability, tied for the category lead.

grok-4.5 scored 15/15 on agentic execution, tied for the category lead.

grok-4.5 scored 14/15 on provenance, down from 15/15 the prior cycle.

Prior cycle: 15/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

grok-4.5 scored 15/15 on resilience, tied for the category lead.

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

gpt-5.6-sol

gpt-5.6-sol — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-sol scored 8/12 on security, down from 10/12 the prior cycle.

Prior cycle: 10/12. 7 days ago: 9/12. 30 days ago: n/a (insufficient history).

gpt-5.6-sol scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-sol scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-sol scored 13/15 on provenance.

Prior cycle: 13/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

Consistency this cycle: 4 of 5 categories produced the same result across all runs.

gpt-5.6-terra

gpt-5.6-terra — 3 security cases unscored: provider policy block, not a model failure.

gpt-5.6-terra scored 9/12 on security, down from 10/12 the prior cycle.

Prior cycle: 10/12. 7 days ago: 9/12. 30 days ago: n/a (insufficient history).

gpt-5.6-terra scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-terra scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-terra scored 14/15 on provenance, up from 13/15 the prior cycle.

Prior cycle: 13/15. 7 days ago: 13/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

gpt-5.6-terra scored 15/15 on resilience, up from 14/15 the prior cycle, tied for the category lead.

Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

gpt-5.6-luna

gpt-5.6-luna scored 12/12 on security, the category leader.

gpt-5.6-luna scored 14/14 on trust and reliability, tied for the category lead.

gpt-5.6-luna scored 15/15 on agentic execution, tied for the category lead.

gpt-5.6-luna scored 14/15 on provenance.

Prior cycle: 14/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

1 fabricated citation detected.

gpt-5.6-luna scored 13/15 on resilience, down from 15/15 the prior cycle.

Prior cycle: 15/15. 7 days ago: 14/15. 30 days ago: n/a (insufficient history).

Consistency this cycle: 3 of 5 categories produced the same result across all runs.

kimi-k3

kimi-k3 — 12 security cases unscored: model returned no output.

kimi-k3 scored 0/12 on security, down from 11/12 the prior cycle, lowest in field.

Prior cycle: 11/12. 7 days ago: n/a (insufficient history).

kimi-k3 — 14 trust and reliability cases unscored: model returned no output.

kimi-k3 scored 0/14 on trust and reliability, down from 14/14 the prior cycle, lowest in field.

Prior cycle: 14/14. 7 days ago: n/a (insufficient history).

kimi-k3 — 15 agentic execution cases unscored: model returned no output.

kimi-k3 scored 0/15 on agentic execution, down from 15/15 the prior cycle, lowest in field.

Prior cycle: 15/15. 7 days ago: n/a (insufficient history).

kimi-k3 — 15 provenance cases unscored: model returned no output.

kimi-k3 scored 0/15 on provenance, down from 12/15 the prior cycle, lowest in field.

Prior cycle: 12/15. 7 days ago: n/a (insufficient history).

kimi-k3 — 15 resilience cases unscored: model returned no output.

kimi-k3 scored 0/15 on resilience, down from 14/15 the prior cycle, lowest in field.

Prior cycle: 14/15. 7 days ago: n/a (insufficient history).

Consistency this cycle: 5 of 5 categories produced the same result across all runs.

gemini-3.6-flash

gemini-3.6-flash scored 11/12 on security, up from 10/12 the prior cycle.

Prior cycle: 10/12. 7 days ago: n/a (insufficient history).

gemini-3.6-flash scored 14/14 on trust and reliability, tied for the category lead.

gemini-3.6-flash scored 15/15 on agentic execution, tied for the category lead.

gemini-3.6-flash scored 12/15 on provenance.

Prior cycle: 12/15. 7 days ago: n/a (insufficient history).

2 fabricated citation results detected.

gemini-3.6-flash scored 14/15 on resilience, up from 13/15 the prior cycle.

Prior cycle: 13/15. 7 days ago: n/a (insufficient history).

Consistency this cycle: 3 of 5 categories produced the same result across all runs.