Hysen Labs
開源專案
kekzl/imp avatar
kekzl

imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

36 個 Star2 個 ForkCudaMIT
GitHub
資料新鮮度

專案分類

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

本頁的專案背景與編輯內容均可免費閱讀。原始 GitHub 儲存庫仍是最終依據;只有收藏或參與討論時才需要登入。

社群筆記

社群筆記