Confirm Action

Are you sure you want to proceed?

Evidence-backed task

Reviewed release

Structured Output

This task measures whether a configured model can return faithful values in the exact structure an application requested.

Published model results
119
Scale
0–100 estimate
Evidence records
0
Release
a4d49d5f8c54

Why it matters

Reliable structured output reduces parsing failures and silent corruption in applications that consume model responses programmatically.

Confidence

Confidence stays at zero for the task claim; the 65% single-lineage cap applies only to calculated capability estimates.

Evidence coverage

The task pins its reviewed requirement-edge revisions and preserves missing coverage rather than converting it into an estimate.

Versioned estimate

Model scores and confidence

Updated 8 Aug 2026

The score is a precise output of this release’s declared method. Confidence, directness, coverage, and limitations are shown separately so a precise number is not mistaken for certainty or universal ability.

  1. Published model

    GPT-5.6 Sol

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  2. Published model

    Claude Fable 5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  3. Published model

    Gemini 3.6 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  4. Published model

    Grok 4.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  5. Published model

    DeepSeek V4 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  6. Published model

    Qwen3.7 Max

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  7. Published model

    GLM 5.2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  8. Published model

    Kimi K3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  9. Published model

    GPT-3.5 Turbo

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  10. Published model

    GPT-4

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  11. Published model

    GPT-4 Turbo

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  12. Published model

    GPT-4o

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  13. Published model

    GPT-4o mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  14. Published model

    GPT-4.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  15. Published model

    GPT-4.1 mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  16. Published model

    GPT-4.1 nano

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  17. Published model

    GPT-5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  18. Published model

    GPT-5 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  19. Published model

    GPT-5 mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  20. Published model

    GPT-5 nano

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  21. Published model

    GPT-5 Codex

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  22. Published model

    GPT-5.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  23. Published model

    GPT-5.1 Codex

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  24. Published model

    GPT-5.1 Codex Max

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  25. Published model

    GPT-5.1 Codex mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  26. Published model

    GPT-5.2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  27. Published model

    GPT-5.2 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  28. Published model

    GPT-5.2 Codex

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  29. Published model

    GPT-5.3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  30. Published model

    GPT-5.3 Codex

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  31. Published model

    GPT-5.4

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  32. Published model

    GPT-5.4 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  33. Published model

    GPT-5.4 mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  34. Published model

    GPT-5.4 nano

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  35. Published model

    GPT-5.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  36. Published model

    GPT-5.5 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  37. Published model

    GPT-5.6 Luna

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  38. Published model

    GPT-5.6 Luna Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  39. Published model

    GPT-5.6 Terra

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  40. Published model

    GPT-5.6 Terra Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  41. Published model

    GPT-5.6 Sol Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  42. Published model

    GPT-OSS 20B

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  43. Published model

    GPT-OSS 120B

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  44. Published model

    Claude 3 Haiku

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  45. Published model

    Claude 3 Opus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  46. Published model

    Claude 3.5 Haiku

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  47. Published model

    Claude 3.5 Sonnet

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  48. Published model

    Claude 3.7 Sonnet

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  49. Published model

    Claude Haiku 4.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  50. Published model

    Claude Sonnet 4.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  51. Published model

    Claude Sonnet 4.6

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  52. Published model

    Claude Sonnet 5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  53. Published model

    Claude Opus 4.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  54. Published model

    Claude Opus 4.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  55. Published model

    Claude Opus 4.6

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  56. Published model

    Claude Opus 4.7

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  57. Published model

    Claude Opus 4.8

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  58. Published model

    Claude Opus 5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  59. Published model

    Gemini 2.0 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  60. Published model

    Gemini 2.0 Flash Lite

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  61. Published model

    Gemini 2.5 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  62. Published model

    Gemini 2.5 Flash Lite

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  63. Published model

    Gemini 2.5 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  64. Published model

    Gemini 3 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  65. Published model

    Gemini 3 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  66. Published model

    Gemini 3.1 Flash Lite

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  67. Published model

    Gemini 3.1 Pro

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  68. Published model

    Gemini 3.5 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  69. Published model

    Gemini 3.5 Flash Lite

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  70. Published model

    Qwen3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  71. Published model

    Qwen3 Max

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  72. Published model

    Qwen3 Next

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  73. Published model

    Qwen3.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  74. Published model

    Qwen3.5 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  75. Published model

    Qwen3.5 Plus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  76. Published model

    Qwen3.6

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  77. Published model

    Qwen3.6 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  78. Published model

    Qwen3.6 Max

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  79. Published model

    Qwen3.6 Plus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  80. Published model

    Qwen3.7 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  81. Published model

    Qwen3.7 Plus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  82. Published model

    Qwen3.8 Max

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  83. Published model

    Qwen3 Coder

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  84. Published model

    Qwen3 Coder Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  85. Published model

    Qwen3 Coder Next

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  86. Published model

    Qwen3 Coder Plus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  87. Published model

    DeepSeek R1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  88. Published model

    DeepSeek R1 Distill Qwen

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  89. Published model

    DeepSeek R1 Distill Llama

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  90. Published model

    DeepSeek Prover V2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  91. Published model

    DeepSeek V3.1 Terminus

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  92. Published model

    DeepSeek V3.2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  93. Published model

    DeepSeek V4 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  94. Published model

    GLM 4.7

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  95. Published model

    GLM 4.7 Flash

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  96. Published model

    GLM 5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  97. Published model

    GLM 5.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  98. Published model

    Grok 3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  99. Published model

    Grok 3 Mini

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  100. Published model

    Grok 4

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  101. Published model

    Grok 4.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  102. Published model

    Grok 4.20

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  103. Published model

    Grok 4.3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  104. Published model

    Grok Build 0.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  105. Published model

    Grok Code Fast 1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  106. Published model

    Kimi K2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  107. Published model

    Kimi K2.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  108. Published model

    Kimi K2.6

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  109. Published model

    Kimi K2.7 Code

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  110. Published model

    MiniMax M3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  111. Published model

    Codestral

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  112. Published model

    Devstral 2

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  113. Published model

    Devstral Medium

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  114. Published model

    Devstral Small

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  115. Published model

    Mistral Medium 3.1

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  116. Published model

    Mistral Medium 3.5

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  117. Published model

    Mistral Large 3

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  118. Published model

    Llama 4 Scout

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  119. Published model

    Llama 4 Maverick

    The reviewed task requirement set is not fully covered.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.

What this result does not prove

  • All current task rows are insufficient-evidence states, not zero scores.
  • The page remains noindex while its requirement set is incomplete.
  • Structured extraction does not establish general function-calling or agent performance.

How the score is calculated

A future task score must bind compatible capability estimates and the reviewed weights in a new immutable release.