Capability
interactive-agent-reasoning
See scores and evidence → Updated 8 Aug 2026
Published evidence graph
Compare versioned 0–100 capability estimates, confidence, exact model configurations, and the benchmark evidence behind each result.
Selected, not exhaustive
A page appears here only when the current reviewed release explicitly includes it. Scores are evidence-weighted estimates for the named capability, not universal measures of model quality.
Capability
See scores and evidence → Updated 8 Aug 2026
Capability
See scores and evidence → Updated 8 Aug 2026