vllm
A high-throughput and memory-efficient inference and serving engine for LLMs.
HYSEN LABS DIRECTORY
Verified repositories, classifications and analysis from the vllm-project organization.
A high-throughput and memory-efficient inference and serving engine for LLMs.
A framework for efficient model inference with omni-modality models.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Cost-efficient and pluggable Infrastructure components for GenAI inference.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling