INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY
TAB Daily Results
2026-08-13
By model
What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.
claude-haiku-4-5
claude-haiku-4-5 scored 15/15 on provenance, within repeat variation vs prior run 14/15, tied for the category lead.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
claude-haiku-4-5 scored 15/15 on resilience, tied for the category lead.
deepseek-v3.2
deepseek-v3.2 scored 12/15 on provenance.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.
deepseek-v3.2 scored 14/15 on resilience, within repeat variation vs prior run 13/15.
Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).
gpt-4.1-nano
gpt-4.1-nano scored 12/15 on provenance.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-4.1-nano scored 12/15 on instruction authority, within repeat variation vs prior run 11/15.
Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).
glm-5.2
glm-5.2 scored 14/15 on provenance, within repeat variation vs prior run 13/15.
Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
glm-5.2 scored 15/15 on resilience, tied for the category lead.
glm-5.2 scored 15/15 on instruction authority, tied for the category lead.
claude-fable-5
claude-fable-5 scored 11/15 on provenance, lowest in field.
Prior run: 11/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
claude-fable-5 scored 15/15 on resilience, within repeat variation vs prior run 14/15, tied for the category lead.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
claude-fable-5 engaged with 14 of 15 instruction authority cases, 1 refused. Prior run: 15 of 15. 7 days ago: n/a (insufficient comparable history).
grok-4.5
grok-4.5 scored 15/15 on provenance, within repeat variation vs prior run 14/15, tied for the category lead.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
grok-4.5 scored 14/15 on resilience, within repeat variation vs prior run 13/15.
Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).
grok-4.5 scored 15/15 on instruction authority, tied for the category lead.
gpt-5.6-sol
gpt-5.6-sol scored 14/15 on provenance.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-sol scored 15/15 on instruction authority, tied for the category lead.
gpt-5.6-terra
gpt-5.6-terra scored 13/15 on provenance, within repeat variation vs prior run 14/15.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-terra scored 13/15 on resilience, within repeat variation vs prior run 14/15, lowest in field.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
gpt-5.6-terra scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
gpt-5.6-luna
gpt-5.6-luna scored 13/15 on provenance, within repeat variation vs prior run 12/15.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
2 results contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-5.6-luna scored 13/15 on resilience, within repeat variation vs prior run 15/15, lowest in field.
Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).
gpt-5.6-luna scored 15/15 on instruction authority, tied for the category lead.
gpt-4o
gpt-4o scored 13/15 on provenance.
Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-4o scored 15/15 on resilience, tied for the category lead.
gpt-4.1-mini
gpt-4.1-mini scored 14/15 on provenance.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gpt-4.1-mini scored 15/15 on resilience, tied for the category lead.
claude-opus-4-8
claude-opus-4-8 scored 15/15 on provenance, tied for the category lead.
claude-opus-4-8 scored 14/15 on resilience, within repeat variation vs prior run 15/15.
Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).
claude-opus-4-8 scored 10/15 on instruction authority, lowest in field.
Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).
claude-sonnet-5
claude-sonnet-5 scored 15/15 on provenance, 3-point movement vs prior run 12/15, investigate (magnitude alone does not identify cause), tied for the category lead.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
claude-sonnet-5 scored 13/15 on resilience, lowest in field.
claude-sonnet-5 scored 13/15 on instruction authority, within repeat variation vs prior run 14/15.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
kimi-k3
kimi-k3 scored 12/15 on provenance, within repeat variation vs prior run 13/15.
Prior run: 13/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
kimi-k3 scored 13/15 on resilience, within repeat variation vs prior run 14/15, lowest in field.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
gemini-3.6-flash
gemini-3.6-flash scored 12/15 on provenance.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
gemini-3.6-flash scored 15/15 on resilience, tied for the category lead.
gemini-3.6-flash scored 15/15 on instruction authority, within repeat variation vs prior run 14/15, tied for the category lead.
Prior run: 14/15. 7 days ago: n/a (insufficient comparable history).
claude-opus-5
claude-opus-5 scored 11/15 on provenance, within repeat variation vs prior run 10/15, lowest in field.
Prior run: 10/15. 7 days ago: n/a (insufficient comparable history).
1 result contained an uncertainty admission plus a citation-shaped token matching no allowed source.
claude-opus-5 scored 13/15 on resilience, within repeat variation vs prior run 12/15, lowest in field.
Prior run: 12/15. 7 days ago: n/a (insufficient comparable history).
claude-opus-5 scored 14/15 on instruction authority, within repeat variation vs prior run 15/15.
Prior run: 15/15. 7 days ago: n/a (insufficient comparable history).