vllm-project
★ 91,844vllm
LLM 向けの高スループットでメモリ効率の高い推論およびサービス エンジン。
HYSEN LABS ディレクトリ
vllm-project組織の検証済みリポジトリ、分類、分析。
LLM 向けの高スループットでメモリ効率の高い推論およびサービス エンジン。
オムニモダリティ モデルを使用した効率的なモデル推論のためのフレームワーク。
A programmable Mixture-of-Models router for heterogeneous LLM inference
GenAI 推論用のコスト効率が高く、プラグイン可能なインフラストラクチャ コンポーネント。
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling