The comparison
All models across the five categories, today’s completed cycle. Green marks the category leader; a refusal is recorded as unscored.
| Model | Class | Security /12 | Trust /14 | Agentic /15 | Provenance /15 | Resilience /15 | Cost/pass |
|---|---|---|---|---|---|---|---|
| claude-haiku-4-5 | closed · low-cost | 11/12 | 13/14 | 15/15 | 14/15 | 15/15 | $0.0140 |
| deepseek-v3.2 | open-weight · lowest-cost | 11/12 | 14/14 | 15/15 | 12/15 | 13/15 | $0.0024 |
| gpt-4.1-nano | closed · low-cost | 11/12 | 14/14 | 14/15 | 12/15 | 14/15 | $0.0009 |
| glm-5.2 | open-weight · low-cost | 11/12 refused 1/12 (unscored) | 14/14 | 15/15 | 14/15 | 15/15 | $0.0512 |
| claude-fable-5 | frontier · Mythos-class | 3/12 refused 9/12 (unscored) | 14/14 | 15/15 | 12/15 refused 1/15 (unscored) | 13/15 refused 1/15 (unscored) | $0.2653 |
| grok-4.5 | closed · mid-cost | 10/12 refused 1/12 (unscored) | 14/14 | 15/15 | 14/15 | 15/15 | $0.0365 |
| gpt-5.6-sol | frontier · reasoning | 9/12 refused 3/12 (unscored) | 14/14 | 15/15 | 14/15 | 14/15 | $0.0782 |
| gpt-5.6-terra | frontier · reasoning | 9/12 refused 3/12 (unscored) | 14/14 | 15/15 | 13/15 | 14/15 | $0.0421 |
| gpt-5.6-luna | frontier · reasoning | 11/12 refused 1/12 (unscored) | 14/14 | 15/15 | 14/15 | 14/15 | $0.0262 |
Measured 2026-07-21
◆ category leader. A refusal is recorded as unscored, never as a wrong answer.
Category shape by model
Each model’s profile across the five categories. The outer edge is a full pass on every category.
Cost against trust
Cost of one full pass against the Provenance pass rate. Position is the fact; read it as you will.
The five categories
What each category tests. These describe the measurement, not any one day’s result.
- Security Screening: does the model resist attempts to turn it toward harmful use, leak data, or drop its safety refusals under pressure.
- Trust and Reliability: does the model give consistent, correctly formatted, correct answers, and does it stay consistent when the same thing is asked more than once.
- Agentic Execution: given a multi-step task with tools, does the model complete it correctly inside a sealed, controlled environment without taking actions it should not.
- Integrity and Provenance: does the model cite real sources that actually support its claims, and does it admit when it cannot source something instead of fabricating a citation.
- Resilience: does the model hold up under noisy, ambiguous, or adversarial input, rather than falling apart.
How it is scored
Every verdict is computed by fixed, deterministic rules against known ground truth. There is no LLM judging the output, so the judge cost is zero and no model grades another model. A model that declines a case is recorded as unscored, never as a false zero. Every run is kept as a dated, immutable observation that is never overwritten.