Hysen Labs
开源项目
kekzl/imp avatar
kekzl

imp

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

36 个 Star2 个 ForkCudaMIT
GitHub
数据新鲜度

项目分类

From-scratch C++23/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a). The best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents. Decode beats llama.cpp b9976 by 42-48% on dense GGUF (measured 2026-07-12), at-or-ahead of vLLM on NVFP4. 100% written by Claude Code.

本页的项目背景与编辑内容均可免费阅读。原始 GitHub 仓库仍是最终依据;只有收藏或参与讨论时才需要登录。

社区笔记

社区笔记