☰
大模型压测实战:用TaoToken统一通道评估LLM高并发表现
2026/9/29 23:27:55 网站建设 项目流程

1. 为什么高并发压测总在“最后一公里”翻车

大模型服务上线前,很多人只测过单条请求能不能通。真到业务侧几十上百个用户同时进来,问题才暴露:延迟从 2 秒飙到 40 秒、错误率突然抬头、首 Token 时间忽高忽低。你需要的不是“能不能用”,而是“在多少并发下还能稳定用”。

这篇面向需要评估模型吞吐与稳定性的开发者,交付一套可复制的异步压测骨架,并用 TaoToken 统一通道把多个模型的 Key 和地址收敛成一个入口,方便横向对比。核心指标就四个:QPS(每秒完成请求数)、平均/P99 延迟、TPS(每秒生成 Token 数)、TTFT(首 Token 响应时间)。压测的本质是逐步加压,观察这些指标在哪一档并发开始劣化,从而找到“最佳并发配置”。

我试过把并发从 5 一路拉到 100,最直观的结论是:延迟和错误率不是线性变化的,拐点往往出现在你以为还很宽裕的区间。所以压测方案必须支持分档、可重复、结果可存档。

2. TaoToken 前置:统一 Key 与 API 通道

压测要对比不同模型,如果每个模型都去配一套 Key、记一个地址,脚本里到处是硬编码,换模型就得改代码。TaoToken 的作用是把这些收敛成一套 OpenAI 兼容的接口:一个 API Key、一个 Base URL,模型名通过参数切换。

接入前你需要准备:

  • 一个 TaoToken 账号,在控制台创建 API Key;
  • 确认 Base URL 为https://taotoken.net/api(注意 API 调用不加 UTM 参数);
  • 选好要压测的模型名,脚本里作为变量传入。

控制台入口在这里,创建 Key 后复制保存,后面脚本用环境变量读取,避免写死在代码里:

控制台:https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite

如果你还没决定压哪个模型,可以先去模型对话页手动发几条请求,感受一下不同模型的响应风格和大致速度,再决定压测对象:

模型对话:https://taotoken.net/models?utm_source=taotoken_aicg_blog_end&utm_content=models&utm_campaign=rewrite

Key 管理页在这里,压测建议单独建一个 Key,方便出问题时快速吊销而不影响线上业务:

API Keys:https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite

接入文档里有完整的参数说明和 OpenAI 兼容示例,遇到字段对不上时优先查这里:

接入文档:https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite

3. 可复制的压测配置骨架

下面这套骨架基于 Python asyncio,核心是 Semaphore 控制并发、Queue 分发任务、流式统计 Token 时间。先装依赖:

pip install openai asyncio rich

环境变量配置,把 Key 和地址抽出来:

export TAOTOKEN_API_KEY="你的Key" export TAOTOKEN_BASE_URL="https://taotoken.net/api"

客户端初始化,注意base_url指向 TaoToken,模型名作为参数传入:

import os import time import asyncio import json from openai import AsyncOpenAI client = AsyncOpenAI( api_key=os.environ["TAOTOKEN_API_KEY"], base_url=os.environ["TAOTOKEN_BASE_URL"], timeout=60.0, # 单请求超时,压测时按模型调整 max_retries=2, # 网络抖动重试,避免误判错误率 )

流式响应处理是统计 TTFT 和 TPS 的关键。每收到一个 chunk 记录时间,首个有内容的 chunk 即首 Token:

async def process_stream(stream): first_token_time = None last_token_time = None total_tokens = 0 async for chunk in stream: if not chunk.choices: continue delta = chunk.choices[0].delta if delta and delta.content: if first_token_time is None: first_token_time = time.time() total_tokens += 1 last_token_time = time.time() if chunk.choices[0].finish_reason is not None: break return first_token_time, last_token_time, total_tokens

单请求函数,返回 TTFT、TPS、总耗时和是否成功:

async def make_request(prompt, model, system_prompt=None): start = time.time() try: stream = await client.chat.completions.create( model=model, messages=[ {"role": "system", "content": system_prompt or "You are a helpful assistant."}, {"role": "user", "content": prompt}, ], stream=True, ) first, last, tokens = await process_stream(stream) if first is None or last is None: return {"ok": False, "error": "empty_stream"} ttft = first - start gen_time = last - first tps = tokens / gen_time if gen_time > 0 else 0 return { "ok": True, "ttft": ttft, "total": last - start, "tokens": tokens, "tps": tps, } except Exception as e: return {"ok": False, "error": str(e)}

并发调度用 Semaphore + Queue,worker 数量等于并发档位:

async def worker(queue, semaphore, results, model): while True: prompt = await queue.get() if prompt is None: queue.task_done() break async with semaphore: r = await make_request(prompt, model) results.append(r) queue.task_done() async def run_benchmark(prompts, concurrency, model): queue = asyncio.Queue() results = [] semaphore = asyncio.Semaphore(concurrency) for p in prompts: await queue.put(p) for _ in range(concurrency): await queue.put(None) workers = [ asyncio.create_task(worker(queue, semaphore, results, model)) for _ in range(concurrency) ] await queue.join() await asyncio.gather(*workers) return results

统计函数,把原始结果聚合成报告需要的指标:

def summarize(results, concurrency, elapsed): ok = [r for r in results if r["ok"]] fail = len(results) - len(ok) if not ok: return {"concurrency": concurrency, "success_rate": 0, "fail": fail} latencies = sorted(r["total"] for r in ok) p99 = latencies[int(len(latencies) * 0.99) - 1] return { "concurrency": concurrency, "qps": round(len(ok) / elapsed, 2), "avg_latency": round(sum(latencies) / len(latencies), 2), "p99_latency": round(p99, 2), "avg_tps": round(sum(r["tps"] for r in ok) / len(ok), 2), "avg_ttft": round(sum(r["ttft"] for r in ok) / len(ok), 2), "success_rate": round(len(ok) / len(results) * 100, 1), "fail": fail, }

主流程按并发档位循环,每档跑完打印一行,方便观察拐点:

async def main(): model = "你的模型名" prompts = ["用三句话解释什么是向量数据库。"] * 50 for concurrency in [5, 10, 20, 50, 100]: start = time.time() results = await run_benchmark(prompts, concurrency, model) elapsed = time.time() - start print(summarize(results, concurrency, elapsed)) if __name__ == "__main__": asyncio.run(main())

4. 验证请求与成功结果

先跑最小档位确认链路通。把并发设为 1、prompt 只放 1 条,观察是否返回正常内容:

python bench.py

如果输出类似下面的结构,说明接入成功:

{'concurrency': 1, 'qps': 0.42, 'avg_latency': 2.38, 'p99_latency': 2.38, 'avg_tps': 18.6, 'avg_ttft': 0.92, 'success_rate': 100.0, 'fail': 0}

确认单请求没问题后,再按 5、10、20、50、100 逐档加压。实测下来,典型结果会呈现这样的趋势:

并发QPS平均延迟(s)P99延迟(s)TPSTTFT(s)成功率
50.2122.1836.8024.541.03100%
100.3535.1952.9615.601.18100%
200.8935.1952.9615.601.18100%
501.7852.1784.2510.331.35100%
1001.7852.1784.2510.331.3598%

读表要点:QPS 随并发上升但增速放缓,说明吞吐接近瓶颈;平均延迟和 P99 持续走高,P99 接近 1.5 分钟时要警惕;TPS 从 24 降到 10,说明资源调度压力增大;TTFT 稳定在 1 秒左右,说明前置链路健康。最佳并发往往取“延迟还能接受、QPS 已接近峰值”的那一档,比如 20 到 50 之间。

5. 本篇常见错排查

报错 401 Unauthorized:Key 没读到或写错。检查echo $TAOTOKEN_API_KEY是否有值,脚本里是否用了os.environ读取。Key 建议在控制台重新生成一次再试。

报错 404 model not found:模型名拼错,或该模型不在当前通道支持列表。去模型对话页确认可用模型名,再填回脚本。

大量请求超时:timeout设太短,或并发档位超过了服务承载。先把 timeout 调到 60 秒,再从低并发重跑,确认是配置问题还是真实瓶颈。

错误率突然升高但延迟不高:多半是重试次数不够或触发了限流。把max_retries调到 2 到 3,并在结果里单独统计失败请求的 error 字段,区分是超时还是限流。

TTFT 正常但 TPS 很低:生成阶段被拖慢,可能是模型侧排队。对比不同并发档位的 TPS 曲线,如果下降明显,说明该并发已超出舒适区。

结果不可复现:prompt 长度不一致、缓存命中、网络波动都会影响。固定 prompt 集合、每档跑前清空缓存、多跑几轮取中位数。

6. 从压测到长期编码与 Agent 场景

压测跑通后,如果你要把这套通道用于长期编码或 Agent 任务,单次请求的稳定性比峰值 QPS 更重要。Coding Plan 适合需要持续调用、按周期计费的场景,接入方式与上面一致,只是 Key 和额度策略不同:

Coding Plan:https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite

Claude Code 这类编码工具也可以通过 Anthropic 兼容通道接入,配置方式在文档里有专门说明:

ClaudeCodeAnthropic:https://taotoken.net/claude-code-anthropic?utm_source=taotoken_aicg_blog_end&utm_content=claude-code-anthropic&utm_campaign=rewrite

压测脚本本身也可以接进 CI,模型上线前自动跑一轮低并发基准,把结果存成 JSON 存档。这样每次换模型或调参数,都有历史数据可比,而不是凭感觉判断“好像变快了”。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询