vllm-project
★ 86,907vllm
A high-throughput and memory-efficient inference and serving engine for LLMs.
HYSEN LABS DIRECTORY
Verified repositories, classifications and analysis from the vllm-project organization.
A high-throughput and memory-efficient inference and serving engine for LLMs.
A framework for efficient model inference with omni-modality models.
Cost-efficient and pluggable Infrastructure components for GenAI inference.