vllm-project
★ 91,844vllm
適用於法學碩士的高吞吐量和記憶體高效的推理和服務引擎。
HYSEN LABS 專案目錄
來自 vllm-project 組織的可信儲存庫、分類與深度解析。
適用於法學碩士的高吞吐量和記憶體高效的推理和服務引擎。
使用全模態模型進行高效模型推理的架構。
A programmable Mixture-of-Models router for heterogeneous LLM inference
用於 GenAI 推理的經濟高效且可插拔的基礎設施組件。
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling