函式庫 / SDK
microsoft/MInference avatar
microsoft/MInference

MInference

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

1,229 個 Star82 個 ForkPythonMIT

秒懂

它是什麼?
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
可以商用嗎?
可以。MIT 是寬鬆授權:你可以使用、修改並販售以它為基礎的軟體,只需保留著作權與授權聲明。
還在維護嗎?
有在維護。儲存庫最近一次提交在 5 天前。
用什麼語言寫的?
主要是 Python(依據 GitHub 的語言統計)。

以上回答依據專案的 GitHub 資料(最近同步於 2026年9月15日)與我們的分析,不構成法律意見。

資料新鮮度

編輯狀態

本專案的完整編輯分析尚未發布。上方的資訊來自專案的公開 GitHub 中繼資料。在正式環境採用前,請先查看倉庫、授權條款與 issue 追蹤。

社群筆記

社群筆記