Confirm Action

Are you sure you want to proceed?

Planned benchmark facet · ranking not published

Which models are best at structured schemas and values?

Produce schema-valid output with correct grounded values.

This is the evaluation plan, not a model leaderboard.

No model order or score is shown. Candidate external benchmarks and planned first-party comparisons become evidence only after their exact release gates pass; retired benchmark grades are not reused.

First-party pairwise plan

Six distinct tasks for this facet

5

exploration tasks

Used for comparison coverage and calibrated refinement.

1

sealed holdout task

Kept out of model and judge tuning until release validation.

6

tasks in total

One of four equally weighted facets in this collection.

First-party, task-local pairwise comparisons with deterministic validators acting as hard gates. The final task remains sealed from model and judge tuning until release validation.

Collection-level external evidence map

Candidate benchmarks that can inform this skill

These sources are mapped to Tool Calling & Structured Output, then identity-partitioned and lineage-deduplicated. They are supporting inputs, not a substitute for direct task-local pairwise results, and no score is reproduced on this page.

Berkeley Function Calling Leaderboard (BFCL) V4

Berkeley Function Calling Leaderboard V4

Mapped and runnable
Evidence role
primary
Maximum directness
90%
Declared subject
harnessed model
Usable lineages mapped
1

Capabilities: function calling, multi turn tools, call decision.

Mapped fields: BFCL V4 single-turn non-live AST accuracy, BFCL V4 single-turn live AST accuracy, BFCL V4 multi-turn function-calling accuracy, BFCL V4 relevance detection accuracy, BFCL V4 irrelevance/no-call detection accuracy.

Ifstruct

IFStruct

Score feed not yet integrated
Evidence role
primary
Maximum directness
75%
Declared subject
foundation model
Usable lineages mapped
0

Capabilities: structural instruction following.

  • • source is absent from the checked-in source registry

Skillsbench

SkillsBench

Score feed not yet integrated
Evidence role
supporting
Maximum directness
55%
Declared subject
agent
Usable lineages mapped
0

Capabilities: tool workflow, reusable skills.

  • • source is absent from the checked-in source registry

Structured Output Benchmark (SOB)

Structured Output Benchmark

Mapped and runnable
Evidence role
primary
Maximum directness
85%
Declared subject
foundation model
Usable lineages mapped
1

Capabilities: value accuracy, schema compliance, exactness.

Mapped fields: SOB unified value accuracy, SOB unified faithfulness, SOB unified schema compliance, SOB unified perfect response rate.

Current blockers

Why no ranking is live

  • No approved combined collection artifact has been released.
  • The parent collection has not passed its publication gate.
  • No approved facet-level result artifact has been released.

Publication gate

What must pass first

  1. 1Combine at least one eligible external result with at least one first-party pairwise result.
  2. 2Retain at least 2 independent evidence lineages and 8 eligible model configurations.
  3. 3Complete at least 3 facet tasks with at least 80% required-task coverage.
  4. 4Show that this facet is materially distinct, then pass reliability and sealed-release stability checks.

Tool Calling & Structured Output

Explore all four planned skills

Back to Tool Calling & Structured Output Sources and safeguards