Benchmarks / Remote Labor Index
Reported by Remote Labor Index
Remote Labor Index
Share of 240 real paid freelance projects (design, software, data, architecture and more) an agent delivers at a standard a client would accept.
- Results dated
- 1 Oct 2026
- Models
- 14
- Unit
- % of projects
Full results
| # | Model | Projects done to client standard % of projects, higher is better |
|---|---|---|
| 1 | GPT-6 AstraOpenAI |
20.8%
|
| 2 | Claude Fable 5.1Anthropic |
17.9%
|
| 3 | Claude Fable 5Anthropic |
16.1%
|
| 4 | Claude Opus 4.8Anthropic |
8.3%
|
| 5 | GPT-5.5OpenAI |
6.2%
|
| 6 | Gemini 3.7 FlashGoogle |
5.0%
|
| 7 | Claude Opus 4.6Anthropic |
4.2%
|
| 8 | Claude Opus 4.5Anthropic |
3.8%
|
| 9 | GPT-5.2 (medium reasoning)OpenAI |
2.5%
|
| 10 | Claude Sonnet 4.5Anthropic |
2.1%
|
| 10 | GPT-5.2OpenAI |
2.1%
|
| 12 | GPT-5OpenAI |
1.7%
|
| 13 | gemini-3-pro-previewGoogle |
1.2%
|
| 14 | Gemini 2.5 Pro Preview 06-05Google |
0.8%
|
Results as published by Remote Labor Index; we do not re-run them.
What it measures
Share of 240 real paid freelance projects (design, software, data, architecture and more) an agent delivers at a standard a client would accept.
What it does not measure
Not speed or cost; a small set of projects, so a few points either way are noise.
Source
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
Not ranked
- claude-fable-5_unknown: listed twice
Scale AI and the Center for AI Safety; collected by Epoch AI. Licence: Published with permission (Scale AI and the Center for AI Safety results via Epoch AI).