Overall Performance
Qwen-72B achieves better performance than LLaMA2-70B on all tasks and outperforms GPT-3.5 on 7 out of 10 tasks.
Base Model Benchmarks
Performance comparison across major benchmarks for all Qwen base models:For all compared models, we report the best scores between their official reported results and OpenCompass.
World Knowledge
C-Eval Performance
C-Eval is a comprehensive benchmark testing common-sense capability in Chinese, covering 52 subjects across humanities, social sciences, STEM, and other specialties. Qwen-7B on C-Eval Validation Set:
Qwen-7B on C-Eval Test Set:
MMLU Performance
MMLU evaluates English comprehension abilities across 57 subtasks spanning different academic fields and difficulty levels. 5-shot MMLU Accuracy:Coding Capabilities
HumanEval
Zero-shot Pass@1 performance on HumanEval benchmark:Mathematical Reasoning
GSM8K
8-shot accuracy on GSM8K benchmark:Translation Capabilities
WMT22
5-shot BLEU scores on WMT22 translation tasks:Chat Model Performance
World Knowledge (Chat)
Zero-shot C-Eval Validation Set:
Zero-shot MMLU:
Coding (Chat)
Zero-shot Pass@1 on HumanEval:Math (Chat)
GSM8K performance:Quantized Model Performance
Quantized models maintain near-lossless performance while improving memory efficiency:Tool Usage Capabilities
Qwen-7B-Chat performance on tool selection and usage: ReAct Prompting Evaluation:
HuggingFace Agent Benchmark:
The plugins in the evaluation set do not appear in Qwen’s training set, demonstrating genuine generalization capability.
Long Context Performance
Perplexity (PPL) on arXiv dataset with extended context lengths:Qwen supports training-free long-context inference from 2048 to over 8192 tokens using NTK-aware interpolation, LogN attention scaling, and local window attention.