vllm
Un moteur d'inférence et de service à haut débit et économe en mémoire pour les LLM.
RÉPERTOIRE HYSEN LABS
Dépôts vérifiés, classifications et analyses de l’organisation vllm-project.
Un moteur d'inférence et de service à haut débit et économe en mémoire pour les LLM.
Un cadre pour une inférence de modèle efficace avec des modèles omnimodaux.
A programmable Mixture-of-Models router for heterogeneous LLM inference
Composants d'infrastructure rentables et enfichables pour l'inférence GenAI.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
An LLM post-training framework with vLLM for RL Scaling