vllm
Eine speichereffiziente Inferenz- und Serving-Engine mit hohem Durchsatz für LLMs.
HYSEN LABS VERZEICHNIS
Geprüfte Repositories, Klassifikationen und Analysen der Organisation vllm-project.
Eine speichereffiziente Inferenz- und Serving-Engine mit hohem Durchsatz für LLMs.
Ein Framework für effiziente Modellinferenz mit Omnimodalitätsmodellen.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Kosteneffiziente und steckbare Infrastrukturkomponenten für GenAI-Inferenz.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling