Confirm Action

Are you sure you want to proceed?

Evidence-backed capability

Reviewed release

Interactive Agent Reasoning

Interactive agent reasoning describes how well a configured system can explore a novel environment, infer what works and adapt its plan.

Published model results
119
Scale
0–100 estimate
Evidence records
12
Release
a4d49d5f8c54

Why it matters

This is useful when a model must learn through interaction instead of receiving a complete task specification up front.

Confidence

Confidence remains moderate and is capped at 65% because all evidence belongs to the same direct ARC-AGI-3 lineage.

Evidence coverage

Calculated rows bind one published SQLite observation, exact agent configuration, contract and source snapshot; unsupported families are not inferred.

Versioned estimate

Model scores and confidence

Updated 8 Aug 2026

The score is a precise output of this release’s declared method. Confidence, directness, coverage, and limitations are shown separately so a precise number is not mistaken for certainty or universal ability.

  1. Published model

    Claude Opus 5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    30.16/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    claude-opus-5-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  2. Published model

    GPT-5.6 Sol

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    7.78/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gpt-5-6-sol-max

    Max

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  3. Published model

    Claude Opus 4.8

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    1.5/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    claude-opus-4-8-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  4. Published model

    GPT-5.6 Terra

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.8/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gpt-5-6-terra-max

    Max

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  5. Published model

    Claude Opus 4.6

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.5/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    claude-opus-4-6-max

    Max

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  6. Published model

    GPT-5.5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.43/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gpt-5-5-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  7. Published model

    Gemini 3.1 Pro

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.4/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gemini-3-1-pro-preview-preview

    Preview

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  8. Published model

    Grok 4.5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.3/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    grok-4-5-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  9. Published model

    GPT-5.4

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.2/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gpt-5-4-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  10. Published model

    GPT-5.6 Luna

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.18/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    gpt-5-6-luna-max

    Max

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  11. Published model

    Claude Opus 4.7

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.18/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    claude-opus-4-7-high

    High

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  12. Published model

    Grok 4.20

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    0.1/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    grok-4-20-beta-reasoning

    Beta Reasoning

    Source and benchmark

    Relative Human Action Efficiency

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  13. Published model

    Claude Fable 5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  14. Published model

    Gemini 3.6 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  15. Published model

    DeepSeek V4 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  16. Published model

    Qwen3.7 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  17. Published model

    GLM 5.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  18. Published model

    Kimi K3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  19. Published model

    GPT-3.5 Turbo

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  20. Published model

    GPT-4

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  21. Published model

    GPT-4 Turbo

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  22. Published model

    GPT-4o

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  23. Published model

    GPT-4o mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  24. Published model

    GPT-4.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  25. Published model

    GPT-4.1 mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  26. Published model

    GPT-4.1 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  27. Published model

    GPT-5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  28. Published model

    GPT-5 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  29. Published model

    GPT-5 mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  30. Published model

    GPT-5 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  31. Published model

    GPT-5 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  32. Published model

    GPT-5.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  33. Published model

    GPT-5.1 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  34. Published model

    GPT-5.1 Codex Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  35. Published model

    GPT-5.1 Codex mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  36. Published model

    GPT-5.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  37. Published model

    GPT-5.2 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  38. Published model

    GPT-5.2 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  39. Published model

    GPT-5.3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  40. Published model

    GPT-5.3 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  41. Published model

    GPT-5.4 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  42. Published model

    GPT-5.4 mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  43. Published model

    GPT-5.4 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  44. Published model

    GPT-5.5 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  45. Published model

    GPT-5.6 Luna Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  46. Published model

    GPT-5.6 Terra Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  47. Published model

    GPT-5.6 Sol Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  48. Published model

    GPT-OSS 20B

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  49. Published model

    GPT-OSS 120B

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  50. Published model

    Claude 3 Haiku

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  51. Published model

    Claude 3 Opus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  52. Published model

    Claude 3.5 Haiku

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  53. Published model

    Claude 3.5 Sonnet

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  54. Published model

    Claude 3.7 Sonnet

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  55. Published model

    Claude Haiku 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  56. Published model

    Claude Sonnet 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  57. Published model

    Claude Sonnet 4.6

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  58. Published model

    Claude Sonnet 5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  59. Published model

    Claude Opus 4.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  60. Published model

    Claude Opus 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  61. Published model

    Gemini 2.0 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  62. Published model

    Gemini 2.0 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  63. Published model

    Gemini 2.5 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  64. Published model

    Gemini 2.5 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  65. Published model

    Gemini 2.5 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  66. Published model

    Gemini 3 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  67. Published model

    Gemini 3 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  68. Published model

    Gemini 3.1 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  69. Published model

    Gemini 3.5 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  70. Published model

    Gemini 3.5 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  71. Published model

    Qwen3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  72. Published model

    Qwen3 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  73. Published model

    Qwen3 Next

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  74. Published model

    Qwen3.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  75. Published model

    Qwen3.5 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  76. Published model

    Qwen3.5 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  77. Published model

    Qwen3.6

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  78. Published model

    Qwen3.6 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  79. Published model

    Qwen3.6 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  80. Published model

    Qwen3.6 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  81. Published model

    Qwen3.7 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  82. Published model

    Qwen3.7 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  83. Published model

    Qwen3.8 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  84. Published model

    Qwen3 Coder

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  85. Published model

    Qwen3 Coder Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  86. Published model

    Qwen3 Coder Next

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  87. Published model

    Qwen3 Coder Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  88. Published model

    DeepSeek R1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  89. Published model

    DeepSeek R1 Distill Qwen

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  90. Published model

    DeepSeek R1 Distill Llama

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  91. Published model

    DeepSeek Prover V2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  92. Published model

    DeepSeek V3.1 Terminus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  93. Published model

    DeepSeek V3.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  94. Published model

    DeepSeek V4 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  95. Published model

    GLM 4.7

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  96. Published model

    GLM 4.7 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  97. Published model

    GLM 5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  98. Published model

    GLM 5.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  99. Published model

    Grok 3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  100. Published model

    Grok 3 Mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  101. Published model

    Grok 4

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  102. Published model

    Grok 4.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  103. Published model

    Grok 4.3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  104. Published model

    Grok Build 0.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  105. Published model

    Grok Code Fast 1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  106. Published model

    Kimi K2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  107. Published model

    Kimi K2.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  108. Published model

    Kimi K2.6

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  109. Published model

    Kimi K2.7 Code

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  110. Published model

    MiniMax M3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  111. Published model

    Codestral

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  112. Published model

    Devstral 2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  113. Published model

    Devstral Medium

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  114. Published model

    Devstral Small

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  115. Published model

    Mistral Medium 3.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  116. Published model

    Mistral Medium 3.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  117. Published model

    Mistral Large 3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  118. Published model

    Llama 4 Scout

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  119. Published model

    Llama 4 Maverick

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.

Traceable inputs

Supporting evidence

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
24 Jul 2026
Configuration
claude-opus-5-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
9 Jul 2026
Configuration
gpt-5-6-sol-max

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
1 Jun 2026
Configuration
claude-opus-4-8-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
9 Jul 2026
Configuration
gpt-5-6-terra-max

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
17 Dec 2025
Configuration
claude-opus-4-6-max

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
23 Apr 2026
Configuration
gpt-5-5-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
5 Mar 2026
Configuration
gemini-3-1-pro-preview-preview

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
16 Jul 2026
Configuration
grok-4-5-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
5 Mar 2026
Configuration
gpt-5-4-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
9 Jul 2026
Configuration
gpt-5-6-luna-max

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
16 Apr 2026
Configuration
claude-opus-4-7-high

View source ↗

arc-agi-3-official-verified

Relative Human Action Efficiency

Version
ARC-AGI-3 (2026)
Directness
Direct evidence
Release weight
1.0
Observed
5 Mar 2026
Configuration
grok-4-20-beta-reasoning

View source ↗

What this result does not prove

  • Only 12 priority families currently have a reviewed ARC-AGI-3 observation.
  • The score describes the cited agent configuration, not a context-free model.
  • One benchmark lineage caps evidence confidence at 65%.

How the score is calculated

Scores preserve ARC-AGI-3's native 0-100 RHAE scale. The release chooses the highest reviewed result per canonical family and breaks ties by configuration slug.