vLLM 是一款专为大语言模型推理加速而设计的框架,实现了 KV 缓存内存几乎零浪费,解决了内存管理瓶颈问题。
更多 vLLM 中文文档及教程可访问 →go.hyper.ai/Wa62f
本目录包含用于部署 vLLM 应用程序的 Helm 图表。该图表包含部署配置、自动扩缩容、资源管理及其他相关配置项。
*在线运行 vLLM 入门教程:零基础分步指南
源码 examples/online_serving/cohere_rerank_client.py
# SPDX-License-Identifier: Apache-2.0""" 使用 OpenAI 入口点的 Rerank API 的示例, 该 API 兼容Cohere SDK:https://github.com/cohere-ai/cohere-python run: vllm serve BAAI/bge-reranker-base """ import cohere# cohere v1 client# cohere v1 客户端co = cohere.Client(base_url="http://localhost:8000", api_key="sk-fake-key") rerank_v1_result = co.rerank( model="BAAI/bge-reranker-base", query="What is the capital of France?", documents=[ "The capital of France is Paris", "Reranking is fun!", "vLLM is an open-source framework for fast AI serving" ]) print(rerank_v1_result)# or the v2# 或 V2co2 = cohere.ClientV2("sk-fake-key", base_url="http://localhost:8000") v2_rerank_result = co2.rerank( model="BAAI/bge-reranker-base", query="What is the capital of France?", documents=[ "The capital of France is Paris", "Reranking is fun!", "vLLM is an open-source framework for fast AI serving" ]) print(v2_rerank_result)