radixark/miles:README 来源编辑指南
基于 README、仓库元数据和许可证整理 radixark/miles 的安装与核验路径。
项目定位
radixark/miles 的 README 将项目描述为"Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime."。本文只整理仓库能直接核验的内容,不把星标、Fork 或宣传语当成质量证明。README 在"What is Miles?"下的说明是:Miles is a high-performance, enterprise-ready reinforcement learning (RL) framework specifically optimized for Large-Scale model Post-Training.。这给出的首先是项目边界,而不是已经完成的生产验证。
适用场景
从 README 的"Latest Updates"和相关条目看,读者可以先判断它是否解决自己的具体问题:[2026/01] 💎 INT4 Quantization-Aware Training (QAT): Inspired by the Kimi K2-Thinking report, Miles now features a full-stack INT4 W4A16 QAT pipeline.。如果你的目标与这段说明不一致,就不应仅凭项目热度采用它。这里保留原项目名、命令和组件名,方便回到一手来源核对。 README 还列出了另一条可核对的信息:[2026/02] 💡 Miles Detailed Arguments: We've added a detailed command-line argument guide used to configure Miles for RL training and inference.。这类原文条目可以帮助读者设计试运行步骤,但不能代替自己的环境测试。
工作方式
README 把工作方式分散写在"🏗️ Supported Models"等段落中。可确认的线索包括:Miles supports a wide range of state-of-the-art architectures, with a special emphasis on DeepSeek, Qwen, Llama and mainstream models.。这篇整理没有把未写出的架构、性能或安全边界补成结论;真正的运行链仍应结合仓库目录、配置文件和版本标签检查。
安装与第一次运行
第一次安装应从 README 给出的入口开始。当前可复核的命令是: # Pull the latest image docker pull radixark/miles:latest # Or install from source pip install -r requirements.txt pip install -e . 如果仓库没有提供命令,本文不会替它编造安装步骤,而是建议先打开 README 的"High-Performance Rollout • Low Precision Training • Production Stability"部分,确认系统依赖、默认端口和首次初始化动作。
配置与日常使用
日常使用的细节取决于项目实际文档。README 的"🏗️ Supported Models"段落提到:| Family | Supported Models | | :--- | :--- | | DeepSeek | R1, V3, V3.2 | | Qwen | Qwen 2, 2.5, 3 | | Llama | Llama 3, 3.1, 3.3, 4 | | Gemma | Gemma 2, 3, 3N | | GLM | GLM-4.5, GLM-4.6, GLM-4.7 | | MiniMax | M2, M2.1 | | Others | Mistral,。对于配置文件、环境变量、权限和数据目录,当前稿只记录来源明确的部分;未写明的默认值必须在测试环境中验证,并保留可回滚的配置副本。 同一部分还提到:[2026/01] 💎 Unified VLM/LLM Multi-Turn Training: We provided an implementation for the VLM multi-turn sampling paradigm. Developers only need to write a customized rollout function to easily start multi-turn RL for VLM, just like training。