Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot · Hermes 4 - Llama-3.1 70B (Non-reasoning)

Results through 2026-07-30

Hermes 4 - Llama-3.1 70B (Non-reasoning) on GPQA Diamond: scores and setup

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

Best published configuration
Hermes 4 - Llama-3.1 70B (Non-reasoning)
Best score
49%5-way tie
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

Hermes 4 - Llama-3.1 70B (Non-reasoning) score

1 configuration · 1 model entry

This model view includes every scored Hermes 4 - Llama-3.1 70B (Non-reasoning) configuration in Artificial Analysis snapshot · 30 July 2026. Return to the benchmark page to compare all published rows. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Hermes 4 - Llama-3.1 70B (Non-reasoning) published configurations

Each mark is one scored reasoning configuration published for this model on the same benchmark source and evaluation contract. Score labels reproduce Artificial Analysis's rounded display value.

Compare all models →
  1. Non-reasoning 49%
0.0 24.5 49.1

Accuracy · higher is better

Current configurations with a reported GPQA Diamond value; missing values are omitted.

How to interpret the result

What do the Hermes 4 - Llama-3.1 70B (Non-reasoning) results mean?

This view isolates one model family without changing the evaluation track or pretending its configurations are equal-compute runs.

1. Read the best configuration

The strongest published Hermes 4 - Llama-3.1 70B (Non-reasoning) setup scored 49% on GPQA Diamond at Non-reasoning reasoning. That is a system result under the cited evaluation protocol.

2. Reasoning settings matter

Changing a stated reasoning variant or other configuration detail changes the evaluated system. The spread between rows is evidence about the whole setup, not noise to average into one bare score.

3. Compare on the same track

Use the main benchmark page for comparison with other rows from the same source field and evaluation contract. Results from a different harness remain a separate methodology.

Best takeaway: Hermes 4 - Llama-3.1 70B (Non-reasoning)'s published GPQA Diamond score depends on its stated configuration and the evaluation contract.

Treat the Hermes 4 - Llama-3.1 70B (Non-reasoning) rows as configuration results under the cited protocol, not as a context-free model rating.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 Hermes 4 - Llama-3.1 70B (Non-reasoning)Best Non-reasoning 49%
What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same gpqa field.

This filtered table contains only the published Hermes 4 - Llama-3.1 70B (Non-reasoning) configurations from the complete official comparison. It does not hide other model families from the benchmark page or merge results from a different evaluation track.

Official GPQA Diamond resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗