vllm-project
★ 91,926vllm
适用于法学硕士的高吞吐量和内存高效的推理和服务引擎。
HYSEN LABS 项目目录
来自 vllm-project 组织的可信仓库、分类与深度解析。
适用于法学硕士的高吞吐量和内存高效的推理和服务引擎。
使用全模态模型进行高效模型推理的框架。
A programmable Mixture-of-Models router for heterogeneous LLM inference
用于 GenAI 推理的经济高效且可插拔的基础设施组件。
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling