☰
VLLM/sglang evalscope/lm_eval MMLU等精度 准确度评测
2026/10/9 4:53:21 网站建设 项目流程

Ref

欢迎来到 EvalScope 中文教程! | EvalScope

支持的数据集 | EvalScope

https://github.com/EleutherAI/lm-evaluation-harness

如何使用lm-evaluation-harness零代码评估大模型

【GPT】中文大语言模型梳理与测评(C-Eval 、AGIEval、MMLU、SuperCLUE)_superclue c-eval哪个更权威-CSDN博客

Note: 本文主要针对评测本地部署的vllm/sglang部署。

evalscope

使用lmeval评测sglang本地部署依赖vllm导致有一些问题,推荐使用evalscope。

pip install evalscope --upgrade -i https://pypi.tuna.tsinghua.edu.cn/simple pip install 'evalscope[perf]' -i https://pypi.tuna.tsinghua.edu.cn/simple

评测sglang本地部署

evalscope eval \ --model DeepSeek-V3.1-Terminus \ --api-url http://localhost:30000/v1 \ --api-key EMPTY \ --eval-type openai_api \ --datasets mmlu \ --dataset-args '{"mmlu": {"subset_list": ["high_school_physics", "high_school_psychology"], "few_shot_num": 5}}' \ --eval-batch-size 64

只需要评测子任务的话,加上--dataset-args。

evalscope eval \ --model DeepSeek-V3.2 \ --api-url http://localhost:30000/v1 \ --api-key EMPTY \ --eval-type openai_api \ --datasets aime25 \ --dataset-hub huggingface \ --eval-batch-size 128 evalscope eval \ --model DeepSeek-V3.2 \ --api-url http://localhost:30000/v1 \ --api-key EMPTY \ --eval-type openai_api \ --datasets bbh \ --dataset-args '{"bbh": {"subset_list": ["boolean_expressions"], "few_shot_num": 3}}' \ --eval-batch-size 64

设置数据集本地路径:

--dataset-args '{"bbh": {"local_path": "/data1/dataset/bbh"}}' \

generation-config

--generation-config '{"do_sample":true,"temperature":0.5,"chat_template_kwargs":{"enable_thinking":false}}'
# DeepSeek-V3.2 --generation-config '{"chat_template_kwargs":{"thinking": true}}'
# subset_list: chinese, english evalscope eval \ --model DeepSeek-V3.2 \ --api-url http://xxx_ip:30000/v1 \ --api-key EMPTY \ --eval-type openai_api \ --datasets needle_haystack \ --dataset-args '{"needle_haystack": {"subset_list": ["chinese"], "local_path": "/data1/dataset/Needle-in-a-Haystack-Corpus"}}' \ --judge-strategy auto \ --judge-worker-num 80 \ --judge-model-args '{"api_url": "http://judge_model_ip:30000/v1", "model_id": "DeepSeek-V3.2"}' \ --eval-batch-size 80

LongBench-v2

LongBench-v2 | EvalScope

LongBench-v2

各子集统计数据:

子集

样本数

提示词平均长度

提示词最小长度

提示词最大长度

short

180

124200.42

49433

841252

medium

215

501002.72

172108

2233351

long

108

2861217.94

720823

16184015

evalscope eval \ --model GLM-5.3-Flash \ --api-url http://localhost:12121/v1 \ --api-key EMPTY \ --eval-type openai_api \ --generation-config timeout=1800 \ --datasets longbench_v2 \ --dataset-args '{"longbench_v2": {"subset_list": ["short"], "local_path": "/data1/datasets/llm_dataset/llm_accuracy_bench/LongBench-v2/"}}' \ --generation-config '{"max_tokens":24576}' \ --eval-batch-size 20 \ --limit 100

sglang自带评测

gsm8k

python benchmark/gsm8k/bench_sglang.py --port 30000 --num-shots 5 --num-questions 500 python benchmark/gsm8k/bench_sglang.py --host http://127.0.0.1 --port 30000 --num-shots 5 --num-questions 500

使用lm_eval

lm_eval安装

# pip install lm-eval pip install lm-eval[api]

源码安装

git clone https://github.com/EleutherAI/lm-evaluation-harness cd lm-evaluation-harness pip install -e .

查看支持的评测任务

lm-eval --tasks list

VLLM/SGLang serving API评测

lm_eval \ --model local-completions \ --tasks mmlu \ --batch_size=8 \ --model_args '{"model": "Qwen/Qwen3-8B-FP8", "base_url": "http://localhost:8000/v1/completions", "num_concurrent": 8}'

这个可以评测VLLM/SGLang启动的serving服务提供的api接口,从而评测VLLM/sglang不同部署方案的效果。

MMLU

除了mmlu整体评测,还可以分为4个子任务单独评测

mmlu_stem
mmlu_other
mmlu_social_sciences
mmlu_humanities

或者更精细的子任务评测。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询