Confirm Action

Are you sure you want to proceed?

Planned benchmark facet · ranking not published

Which models are best at memos and reports?

Structure evidence and recommendations for busy business readers.

This is the evaluation plan, not a model leaderboard.

No model order or score is shown. Candidate external benchmarks and planned first-party comparisons become evidence only after their exact release gates pass; retired benchmark grades are not reused.

First-party pairwise plan

Six distinct tasks for this facet

5

exploration tasks

Used for comparison coverage and calibrated refinement.

1

sealed holdout task

Kept out of model and judge tuning until release validation.

6

tasks in total

One of four equally weighted facets in this collection.

First-party, task-local pairwise comparisons with deterministic validators acting as hard gates. The final task remains sealed from model and judge tuning until release validation.

Collection-level external evidence map

Candidate benchmarks that can inform this skill

These sources are mapped to Business Writing & Email, then identity-partitioned and lineage-deduplicated. They are supporting inputs, not a substitute for direct task-local pairwise results, and no score is reproduced on this page.

EQ-Bench 3

EQ-Bench message tailoring

Reuse rights review required
Evidence role
supporting
Maximum directness
40%
Declared subject
foundation model
Usable lineages mapped
1

Capabilities: tone control, message tailoring.

Mapped fields: EQ-Bench 3 — Message Tailoring, EQ-Bench 3 — Pragmatic EI.

  • • result-feed reuse rights require review

LMArena Leaderboard Dataset

Style-controlled writing and business

Mapped and runnable
Evidence role
primary
Maximum directness
55%
Declared subject
foundation model
Usable lineages mapped
1

Capabilities: human preference, instruction following.

Mapped fields: LMArena Text — Writing, literature and language, LMArena Text — Business, management and financial operations, LMArena Text (style controlled) — Instruction following.

UGI Leaderboard

UGI writing components

Mapped and runnable
Evidence role
primary
Maximum directness
65%
Declared subject
foundation model
Usable lineages mapped
1

Capabilities: style control, length control, originality, repetition.

Mapped fields: UGI style adherence, UGI requested-length error, UGI writing originality, UGI semantic redundancy.

Writingbench

WritingBench business and finance

Reuse rights review required
Evidence role
primary
Maximum directness
75%
Declared subject
foundation model
Usable lineages mapped
0

Capabilities: business writing, finance writing.

  • • result-feed reuse rights require review
  • • source is absent from the checked-in source registry

Current blockers

Why no ranking is live

  • No approved combined collection artifact has been released.
  • The parent collection has not passed its publication gate.
  • No approved facet-level result artifact has been released.

Publication gate

What must pass first

  1. 1Combine at least one eligible external result with at least one first-party pairwise result.
  2. 2Retain at least 2 independent evidence lineages and 8 eligible model configurations.
  3. 3Complete at least 3 facet tasks with at least 80% required-task coverage.
  4. 4Show that this facet is materially distinct, then pass reliability and sealed-release stability checks.

Business Writing & Email

Explore all four planned skills

Back to Business Writing & Email Sources and safeguards