INDEPENDENT MEASUREMENT • PLAIN-LANGUAGE READOUT • UPDATED DAILY
TAB Daily Results 2026-08-19
All 16 public models within baseline parameters this cycle. No threshold breach.
Interpretation notice — rule_based_provenance v1.0.0. Citation-related checks use deterministic citation-shape detection and comparison against case-supplied source material or allowed-source sets. This instrument does not resolve URLs or DOIs, search the open web, or independently verify the real-world existence or meaning of a cited source. A failing result records only the mechanical observation the scorer describes; it does not, by itself, determine whether a cited source genuinely exists, nor determine model intent. Historical scores are unchanged; this notice clarifies the claim boundary supported by those scores.
Measurement validity notice, September 1, 2026. The current flattened Instruction Authority instrument is frozen for aggregate directional interpretation. Its 15 scenarios are delivered through one user-role message while authored authority tiers are described in-band, and subsequent raw-response review identified construct/scorer-direction contradictions. Historical observations remain unchanged. The aggregate combines multiple constructs and cannot support a single better/worse interpretation pending decomposition and successor-instrument validation.
By model What today’s completed cycle measured, one model at a time. Every line is a fact from the same data behind the Index.
claude-haiku-4-5 claude-haiku-4-5 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
deepseek-v3.2 deepseek-v3.2 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
gpt-4.1-nano gpt-4.1-nano measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
glm-5.2 glm-5.2 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
claude-fable-5 claude-fable-5 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
grok-4.5 grok-4.5 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
gpt-5.6-sol gpt-5.6-sol measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
gpt-5.6-terra gpt-5.6-terra measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
gpt-5.6-luna gpt-5.6-luna measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-15).
gpt-4o gpt-4o measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-30).
gpt-4.1-mini gpt-4.1-mini measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-30).
claude-opus-4-8 claude-opus-4-8 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-30).
claude-sonnet-5 claude-sonnet-5 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-30).
kimi-k3 kimi-k3 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-24).
gemini-3.6-flash gemini-3.6-flash measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-26).
claude-opus-5 claude-opus-5 measured this cycle. Insufficient history; baseline accruing (first measured 2026-07-28).
Cycle ID: e2bad91f-8baa-4309-8c26-4ca4f246bc04 · Timestamp (UTC): 2026-08-19 08:45:00 UTC · Models measured: 16 Deterministic scoring against known ground truth. No LLM-as-judge. Judge cost: zero. Every run is a dated, immutable observation; records are never overwritten. Read the full methodology → · Archive → Methodology note, August 10, 2026. "Lowest in field" and "category leader" compare the number of cases a model passed, not the fraction of the cases it was scored on. A model that declines cases in a category is scored on fewer cases than the others, so its passed-count is not directly comparable to the count of a model scored on every case. Where a model's unscored count is large, read the score alongside it. Dated pages are unchanged and show what was published on each date. Methodology note, August 10, 2026. From this date, a category cell scored on fewer than half its cases shows its coverage instead of a score, and is excluded from the comparison that names the lowest and the leading model in a category. The floor is half the category's cases, rounded up, minimum two. Cells published before this date are unchanged. Methodology note, August 11, 2026. From this date, a category cell scored on fewer than half its cases is also excluded from deviation testing. Such a cell is not compared against its own baseline and does not contribute to it, because a score computed on a small and changing number of cases carries variation from the count of cases scored, not only from the model's behaviour. Published scores and the dated pages are unchanged; this affects the deviation baseline only. Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB previously recognised only one of them. From this date, a block signalled by xAI is recorded as a provider policy block rather than as a call that did not complete. The distinction matters because one describes the provider and the other describes TAB's own measurement failing. Cells published before this date are unchanged and show what was recorded then. Methodology note, August 14, 2026. A provider that blocks a request before the model sees it can signal that block in different ways, and TAB has recognised those signals as it identified them. One such signal was not recognised before July 20, 2026, so blocks carrying it were recorded in that period as calls that did not complete rather than as provider policy blocks. The measurement did not change and the scores were not affected, only the description of why a case went unscored. Cells published in that period are unchanged and show what was recorded then.