radixark/miles:README に基づく導入ガイド
README、メタデータ、ライセンスに基づく radixark/miles の導入と確認ガイドです。
プロジェクトの範囲
radixark/miles の README はプロジェクトを「Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.」と説明しています。ここではリポジトリで確認できる事実だけを整理します。star 数やバッジは注目度の手掛かりであり、品質の証明ではありません。「What is Miles?」には次の説明があります。Miles is a high-performance, enterprise-ready reinforcement learning (RL) framework specifically optimized for Large-Scale model Post-Training.。これは範囲の説明であり、本番検証の結果ではありません。
向いている用途
README の「Latest Updates」にある内容から、用途が合うかを先に判断できます。[2026/01] 💎 INT4 Quantization-Aware Training (QAT): Inspired by the Kimi K2-Thinking report, Miles now features a full-stack INT4 W4A16 QAT pipeline.。目的が違うなら、人気だけで採用する理由にはなりません。プロジェクト名やコマンドは原文のまま残し、一次資料へ戻って用語を確認できるようにしています。 README には次の確認可能な項目もあります。[2026/02] 💡 Miles Detailed Arguments: We've added a detailed command-line argument guide used to configure Miles for RL training and inference.。初回テストの材料にはなりますが、実際の環境での確認を省略する理由にはなりません。
動作の考え方
動作の説明は「🏗️ Supported Models」など複数の箇所に分かれています。確認できる情報は次の通りです。Miles supports a wide range of state-of-the-art architectures, with a special emphasis on DeepSeek, Qwen, Llama and mainstream models.。書かれていない構成、性能、セキュリティを推測で補いません。導入時はディレクトリ、設定ファイル、release 履歴を確認してください。
インストールと初回起動
初回導入は README の入口から始めます。確認できるコマンドは次の通りです。 # Pull the latest image docker pull radixark/miles:latest # Or install from source pip install -r requirements.txt pip install -e . 実行可能なコマンドがない場合は手順を作らず、「High-Performance Rollout • Low Precision Training • Production Stability」で依存関係、待受ポート、初回設定を確認します。
設定と日常運用
日常運用は公式文書の範囲に限ります。「🏗️ Supported Models」には| Family | Supported Models | | :--- | :--- | | DeepSeek | R1, V3, V3.2 | | Qwen | Qwen 2, 2.5, 3 | | Llama | Llama 3, 3.1, 3.3, 4 | | Gemma | Gemma 2, 3, 3N | | GLM | GLM-4.5, GLM-4.6, GLM-4.7 | | MiniMax | M2, M2.1 | | Others | Mistral,とあります。設定、環境変数、権限、データ保存先は明記されたものだけを扱います。未記載の既定値は隔離環境で確認し、戻せる設定を保存してください。 同じ資料には[2026/01] 💎 Unified VLM/LLM Multi-Turn Training: We provided an implementation for the VLM multi-turn sampling paradigm. Developers only need to write a customized rollout function to easily start multi-turn RL for VLM, just like trainingともあります。