Confirm Action

Are you sure you want to proceed?

Evidence-backed capability

Reviewed release

Structured Output Reliability

Structured-output reliability describes whether a configured model can preserve source values while obeying an exact output structure.

Published model results
119
Scale
0–100 estimate
Evidence records
22
Release
a4d49d5f8c54

Why it matters

Applications that parse model output need both valid structure and values that remain faithful to the source input.

Confidence

Confidence remains moderate and is capped at 65%: the evidence is direct, but it all belongs to one benchmark lineage.

Evidence coverage

Every displayed estimate traces to the selected configuration, source observation, contract, immutable snapshot and mapping revision.

Versioned estimate

Model scores and confidence

Updated 8 Aug 2026

The score is a precise output of this release’s declared method. Confidence, directness, coverage, and limitations are shown separately so a precise number is not mistaken for certainty or universal ability.

  1. Published model

    GPT-5.4

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    87.017213/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-1e0e3486b1f238ce35871544

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  2. Published model

    Gemini 3.1 Pro

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.944535/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-cf45c94b7cd67d79d8a525df

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  3. Published model

    GLM 5.1

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.610235/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-984c789c85fec606ba5f94ba

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  4. Published model

    Claude Opus 4.7

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.373753/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-dd2a368fe2be7708c1f08b42

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  5. Published model

    Claude Sonnet 5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.242454/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-13a2e54d392a29815d1fea83

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  6. Published model

    GLM 4.7

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.086045/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-b35b1a8be4df85b84c094ec1

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  7. Published model

    GPT-5.5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.031279/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-19838225470376f155eba33a

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  8. Published model

    Gemini 2.5 Flash

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    86.029227/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-fe20b4d57ff9f86b5ee889a2

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  9. Published model

    Gemini 3.5 Flash

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.641663/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-6c859177aeea19eef1fbf998

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  10. Published model

    Gemini 2.5 Pro

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.621643/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-e681e85aae438b4aa66fac6c

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  11. Published model

    Claude Sonnet 4.6

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.438657/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-f924bcaa6afa6dd2848f2c07

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  12. Published model

    Claude Opus 4.6

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.308764/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-505f934f52b18ce6add3f7af

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  13. Published model

    Kimi K2.6

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.275872/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-41182d44bb0774d5483674ad

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  14. Published model

    DeepSeek V4 Pro

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.275213/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-5ba4264f08c80f97f9500c1e

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  15. Published model

    Claude Fable 5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    85.090575/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-c67f1f90b2c4462e611dcbd7

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  16. Published model

    GPT-4.1

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    84.991182/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-63d62997aa3c5868af2424ba

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  17. Published model

    GPT-5

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    84.892585/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-13f74cc33a8cc91c1425b46d

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  18. Published model

    GPT-5.4 mini

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    84.696202/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-84cd0a6f4b3d4ff1c0324d18

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  19. Published model

    Grok 4.3

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    84.098756/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-9ae9a199cbcee774e69cef2f

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  20. Published model

    GPT-5 mini

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    83.524065/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-1ee85cf12fb37d0b9c0deb46

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  21. Published model

    Gemini 3 Flash

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    83.27062/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-a35fba91c706667f5f25b8d6

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  22. Published model

    GPT-OSS 20B

    The 0-100 estimate preserves one reviewed source-native primary metric.

    Score

    73.173539/100

    Confidence
    65%
    Evidence coverage
    1
    Directness
    1/1 direct
    Status
    Calculated

    Representative configuration

    sob-configuration-8beb522273f97432467c03e4

    Source and benchmark

    Structured Output Benchmark Overall

    Limitations

    • Not a calibrated probability of task success.
    • No incompatible benchmark scale is blended into the point estimate.
  23. Published model

    GPT-5.6 Sol

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  24. Published model

    Gemini 3.6 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  25. Published model

    Grok 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  26. Published model

    Qwen3.7 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  27. Published model

    GLM 5.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  28. Published model

    Kimi K3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  29. Published model

    GPT-3.5 Turbo

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  30. Published model

    GPT-4

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  31. Published model

    GPT-4 Turbo

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  32. Published model

    GPT-4o

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  33. Published model

    GPT-4o mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  34. Published model

    GPT-4.1 mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  35. Published model

    GPT-4.1 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  36. Published model

    GPT-5 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  37. Published model

    GPT-5 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  38. Published model

    GPT-5 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  39. Published model

    GPT-5.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  40. Published model

    GPT-5.1 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  41. Published model

    GPT-5.1 Codex Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  42. Published model

    GPT-5.1 Codex mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  43. Published model

    GPT-5.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  44. Published model

    GPT-5.2 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  45. Published model

    GPT-5.2 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  46. Published model

    GPT-5.3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  47. Published model

    GPT-5.3 Codex

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  48. Published model

    GPT-5.4 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  49. Published model

    GPT-5.4 nano

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  50. Published model

    GPT-5.5 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  51. Published model

    GPT-5.6 Luna

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  52. Published model

    GPT-5.6 Luna Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  53. Published model

    GPT-5.6 Terra

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  54. Published model

    GPT-5.6 Terra Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  55. Published model

    GPT-5.6 Sol Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  56. Published model

    GPT-OSS 120B

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  57. Published model

    Claude 3 Haiku

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  58. Published model

    Claude 3 Opus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  59. Published model

    Claude 3.5 Haiku

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  60. Published model

    Claude 3.5 Sonnet

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  61. Published model

    Claude 3.7 Sonnet

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  62. Published model

    Claude Haiku 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  63. Published model

    Claude Sonnet 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  64. Published model

    Claude Opus 4.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  65. Published model

    Claude Opus 4.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  66. Published model

    Claude Opus 4.8

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  67. Published model

    Claude Opus 5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  68. Published model

    Gemini 2.0 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  69. Published model

    Gemini 2.0 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  70. Published model

    Gemini 2.5 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  71. Published model

    Gemini 3 Pro

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  72. Published model

    Gemini 3.1 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  73. Published model

    Gemini 3.5 Flash Lite

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  74. Published model

    Qwen3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  75. Published model

    Qwen3 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  76. Published model

    Qwen3 Next

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  77. Published model

    Qwen3.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  78. Published model

    Qwen3.5 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  79. Published model

    Qwen3.5 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  80. Published model

    Qwen3.6

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  81. Published model

    Qwen3.6 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  82. Published model

    Qwen3.6 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  83. Published model

    Qwen3.6 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  84. Published model

    Qwen3.7 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  85. Published model

    Qwen3.7 Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  86. Published model

    Qwen3.8 Max

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  87. Published model

    Qwen3 Coder

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  88. Published model

    Qwen3 Coder Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  89. Published model

    Qwen3 Coder Next

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  90. Published model

    Qwen3 Coder Plus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  91. Published model

    DeepSeek R1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  92. Published model

    DeepSeek R1 Distill Qwen

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  93. Published model

    DeepSeek R1 Distill Llama

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  94. Published model

    DeepSeek Prover V2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  95. Published model

    DeepSeek V3.1 Terminus

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  96. Published model

    DeepSeek V3.2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  97. Published model

    DeepSeek V4 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  98. Published model

    GLM 4.7 Flash

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  99. Published model

    GLM 5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  100. Published model

    Grok 3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  101. Published model

    Grok 3 Mini

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  102. Published model

    Grok 4

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  103. Published model

    Grok 4.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  104. Published model

    Grok 4.20

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  105. Published model

    Grok Build 0.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  106. Published model

    Grok Code Fast 1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  107. Published model

    Kimi K2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  108. Published model

    Kimi K2.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  109. Published model

    Kimi K2.7 Code

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  110. Published model

    MiniMax M3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  111. Published model

    Codestral

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  112. Published model

    Devstral 2

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  113. Published model

    Devstral Medium

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  114. Published model

    Devstral Small

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  115. Published model

    Mistral Medium 3.1

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  116. Published model

    Mistral Medium 3.5

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  117. Published model

    Mistral Large 3

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  118. Published model

    Llama 4 Scout

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.
  119. Published model

    Llama 4 Maverick

    No reviewed primary evidence lane covers this family and scope.

    Score

    Insufficient evidence

    Confidence
    0%
    Evidence coverage
    See evidence below
    Directness
    Not stated
    Status
    Insufficient evidence

    Limitations

    • Missing evidence is not reweighted or inferred.

Traceable inputs

Supporting evidence

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-1e0e3486b1f238ce35871544

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-cf45c94b7cd67d79d8a525df

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-984c789c85fec606ba5f94ba

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-dd2a368fe2be7708c1f08b42

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-13a2e54d392a29815d1fea83

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-b35b1a8be4df85b84c094ec1

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-19838225470376f155eba33a

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-fe20b4d57ff9f86b5ee889a2

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-6c859177aeea19eef1fbf998

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-e681e85aae438b4aa66fac6c

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-f924bcaa6afa6dd2848f2c07

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-505f934f52b18ce6add3f7af

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-41182d44bb0774d5483674ad

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-5ba4264f08c80f97f9500c1e

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-c67f1f90b2c4462e611dcbd7

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-63d62997aa3c5868af2424ba

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-13f74cc33a8cc91c1425b46d

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-84cd0a6f4b3d4ff1c0324d18

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-9ae9a199cbcee774e69cef2f

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-1ee85cf12fb37d0b9c0deb46

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-a35fba91c706667f5f25b8d6

View source ↗

sob-v1

Structured Output Benchmark Overall

Version
da785a8521c8954283b2989d01e54d80c4e023c6
Directness
Direct evidence
Release weight
1.0
Observed
17 Jul 2026
Configuration
sob-configuration-8beb522273f97432467c03e4

View source ↗

What this result does not prove

  • The score describes the pinned SOB product-family configuration.
  • Thirteen source-exact identities remain evidence-only pending reviewed family mapping.
  • Seven component values share one correlated source lineage and do not receive independent weight.
  • A source-native score is not automatically comparable with unrelated benchmark scales.

How the score is calculated

Scores retain SOB's native 0-100 Overall value. Correlated components are shown as supporting context and are not averaged into the score again.