Skip to main content
Qwen models outperform baseline models of similar sizes across a series of benchmark datasets, evaluating capabilities in natural language understanding, mathematical problem solving, coding, and more.

Overall Performance

Qwen-72B achieves better performance than LLaMA2-70B on all tasks and outperforms GPT-3.5 on 7 out of 10 tasks. Qwen-72B Performance Radar

Base Model Benchmarks

Performance comparison across major benchmarks for all Qwen base models:
For all compared models, we report the best scores between their official reported results and OpenCompass.

World Knowledge

C-Eval Performance

C-Eval is a comprehensive benchmark testing common-sense capability in Chinese, covering 52 subjects across humanities, social sciences, STEM, and other specialties. Qwen-7B on C-Eval Validation Set: Qwen-7B on C-Eval Test Set:

MMLU Performance

MMLU evaluates English comprehension abilities across 57 subtasks spanning different academic fields and difficulty levels. 5-shot MMLU Accuracy:

Coding Capabilities

HumanEval

Zero-shot Pass@1 performance on HumanEval benchmark:

Mathematical Reasoning

GSM8K

8-shot accuracy on GSM8K benchmark:

Translation Capabilities

WMT22

5-shot BLEU scores on WMT22 translation tasks:

Chat Model Performance

World Knowledge (Chat)

Zero-shot C-Eval Validation Set: Zero-shot MMLU:

Coding (Chat)

Zero-shot Pass@1 on HumanEval:

Math (Chat)

GSM8K performance:

Quantized Model Performance

Quantized models maintain near-lossless performance while improving memory efficiency:

Tool Usage Capabilities

Qwen-7B-Chat performance on tool selection and usage: ReAct Prompting Evaluation: HuggingFace Agent Benchmark:
The plugins in the evaluation set do not appear in Qwen’s training set, demonstrating genuine generalization capability.

Long Context Performance

Perplexity (PPL) on arXiv dataset with extended context lengths:
Qwen supports training-free long-context inference from 2048 to over 8192 tokens using NTK-aware interpolation, LogN attention scaling, and local window attention.

Additional Resources

For detailed model performance on additional benchmark datasets, please refer to the technical report.