Hysen Labs
Open-Source-Projekt
kekzl/imp avatar
kekzl

imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

36 Sterne2 ForksCudaMIT
GitHub
Datenaktualität

Klassifizierung

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

Der Projektkontext auf dieser Seite ist kostenlos lesbar. Das ursprüngliche GitHub-Repository bleibt maßgeblich; nur zum Speichern oder Diskutieren ist eine Anmeldung nötig.

Community-Notizen

Community-Notizen