vllm
A high-throughput and memory-efficient inference and serving engine for LLMs.
HYSEN LABS DIRECTORY
Verified repositories, classifications and analysis from the vllm-project organization.
A high-throughput and memory-efficient inference and serving engine for LLMs.
A framework for efficient model inference with omni-modality models.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Cost-efficient and pluggable Infrastructure components for GenAI inference.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling