Benchmarks / Vectara Hallucination Leaderboard

Reported by Vectara Hallucination Leaderboard

Vectara Hallucination Leaderboard

Share of documents the model agreed to summarise rather than refusing.

Results dated
22 Sep 2026
Models
108
Unit
% of documents
Licence
Apache License 2.0
Vectara: answer rate, % of documents, higher is better
#ModelAnswer rate
% of documents, higher is better
1 Nova Micro 1.0Amazon
100.0%
1 Gemma 4 31BGoogle
100.0%
1 granite-3.3-8b-instructIBM
100.0%
1 granite-4.0-h-smallIBM
100.0%
1 Mercury 2Inception
100.0%
1 llama-4-maverick-17b-128e-instruct-fp8Meta
100.0%
1 GPT-5.1OpenAI · GPT-5.1 (high)
100.0%
1 GPT-5.1OpenAI · gpt-5.1
100.0%
1 GPT-5.2OpenAI · GPT-5.2 (high)
100.0%
1 GPT-5.2OpenAI · gpt-5.2
100.0%
1 GPT-5.4 MiniOpenAI
100.0%
1 GPT-5.4 NanoOpenAI
100.0%
1 GPT-5.4 ProOpenAI
100.0%
1 GPT-5.5OpenAI
100.0%
1 GPT-5 NanoOpenAI
100.0%
1 GPT-6 AstraOpenAI
100.0%
1 GPT-6 SolOpenAI
100.0%
1 o3 ProOpenAI
100.0%
19 Qwen3 14BAlibaba
99.9%
19 Qwen3 32BAlibaba
99.9%
19 qwen3-4bAlibaba
99.9%
19 Qwen3 8BAlibaba
99.9%
19 Nova Lite 1.0Amazon
99.9%
19 Claude Sonnet 4.6Anthropic
99.9%
19 ministral-3b-2410Mistral
99.9%
19 ministral-8b-2410Mistral
99.9%
19 mistral-large-2411Mistral
99.9%
19 GPT-4.1OpenAI
99.9%
19 GPT-5.4OpenAI
99.9%
19 GPT-5 MiniOpenAI
99.9%
19 GPT-5OpenAI · GPT-5 (high)
99.9%
19 GPT-5OpenAI · gpt-5
99.9%
19 gpt-oss-120bOpenAI
99.9%
34 Qwen3.5-122B-A10BAlibaba
99.8%
34 Qwen3.5-27BAlibaba
99.8%
34 Qwen3.5-35B-A3BAlibaba
99.8%
34 Qwen3.5-FlashAlibaba
99.8%
34 Qwen3.5 Plus 2026-02-15Alibaba
99.8%
34 Claude Opus 4.6Anthropic
99.8%
34 c4ai-aya-expanse-32bCohere
99.8%
34 Gemini 3 Flash PreviewGoogle
99.8%
34 Gemma 4 26B A4BGoogle
99.8%
34 GLM 4.7Z.ai
99.8%
44 Mistral Medium 3.1Mistral
99.7%
44 Kimi K2.6Moonshot AI
99.7%
44 grok-4-1-fast-reasoningxAI
99.7%
44 GLM 5Z.ai
99.7%
48 jamba-mini-2AI21 Labs
99.6%
48 Nova 2 LiteAmazon
99.6%
48 Gemini 3.1 Flash Lite PreviewGoogle
99.6%
48 Ministral 3 14B 2512Mistral
99.6%
48 Nemotron 3 Nano 30B A3BNVIDIA
99.6%
53 finix_s1_32bAnt Group
99.5%
53 Claude Haiku 4.5Anthropic
99.5%
53 Gemini 2.5 Flash LiteGoogle
99.5%
53 llama-3.3-70b-instruct-turboMeta
99.5%
53 grok-4-fast-reasoningxAI
99.5%
58 Gemini 3.1 Pro PreviewGoogle
99.4%
58 gemini-3-pro-previewGoogle
99.4%
58 MiniMax M2.7MiniMax
99.4%
61 Nova Pro 1.0Amazon
99.3%
62 o4 Mini HighOpenAI
99.2%
62 grok-4-fast-non-reasoningxAI
99.2%
64 ai21-jamba-mini-1.7AI21 Labs
99.1%
64 Gemini 2.5 ProGoogle
99.1%
64 Ministral 3 8B 2512Mistral
99.1%
64 GPT-5.6 SolOpenAI
99.1%
68 trinity-large-previewArcee Ai
99.0%
68 Gemini 2.5 FlashGoogle
99.0%
68 Llama 4 ScoutMeta
99.0%
71 ai21-jamba-large-1.7AI21 Labs
98.9%
72 Gemma 3 27BGoogle
98.8%
72 Mistral Large 3 2512Mistral
98.8%
74 Claude Opus 4.5Anthropic
98.7%
74 o4 MiniOpenAI · o4-mini
98.7%
76 Claude Sonnet 4Anthropic
98.6%
76 Kimi K2 0905Moonshot AI
98.6%
78 MiniMax M2.1MiniMax
98.5%
78 grok-4-1-fast-non-reasoningxAI
98.5%
80 MiniMax M2.5MiniMax
98.2%
81 glm-4.5-air-fp8Z.ai
98.1%
82 Claude Opus 4.7Anthropic
98.0%
83 Mistral Small 3Mistral
97.9%
84 Command ACohere
97.6%
85 DeepSeek V3DeepSeek
97.5%
86 Gemma 3 12BGoogle
97.4%
87 DeepSeek V4 Pro 0423DeepSeek
97.2%
88 R1DeepSeek
97.0%
89 DeepSeek V3.2 ExpDeepSeek
96.6%
90 Claude Sonnet 4.5Anthropic
95.6%
91 Command R+ (08-2024)Cohere
95.0%
92 Qwen3 235B A22BAlibaba
94.9%
93 DeepSeek V3.1DeepSeek
94.5%
93 GLM 4.6Z.ai
94.5%
95 Qwen3 Next 80B A3B ThinkingAlibaba
94.4%
96 GPT-4o (2024-08-06)OpenAI
93.8%
97 grok-3xAI
93.0%
98 DeepSeek V3.2DeepSeek
92.6%
99 phi-4-mini-instructMicrosoft
92.5%
100 Claude Opus 4.1Anthropic
92.4%
101 Kimi K2.5Moonshot AI
92.2%
102 GLM 4.7 FlashZ.ai
91.6%
103 claude-opus-4Anthropic
91.0%
104 Phi 4Microsoft
80.7%
105 c4ai-aya-expanse-8bCohere
77.5%
106 Ministral 3 3B 2512Mistral
74.3%
107 Gemma 3 4BGoogle
67.3%
108 snowflake-arctic-instructSnowflake
62.7%

Results as published by Vectara Hallucination Leaderboard; we do not re-run them.

What it measures

Share of documents the model agreed to summarise rather than refusing.

What it does not measure

Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.

Vectara Hallucination Leaderboard by Vectara, Apache License 2.0.