intel
★ 2,707neural-compressor
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
HYSEN LABS 项目目录
来自 intel 组织的可信仓库、分类与深度解析。
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
英特尔 AutoRound 是一款用于量化工作流程的模型优化工具包,可降低推理成本,同时保持人工智能部署的准确性。