SkillOpt: Training Agent Skills as Trainable Parameters Without Touching Model Weights
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
At a glance
- What is it?
- SkillOpt is a Python framework from Microsoft that trains natural-language skill documents for frozen LLM agents using a loop of scored rollouts, bounded edits, and validation gates, producing a compact `best_skill.md` that runs against the unchanged target model at deployment.
- Who is it for?
- SkillOpt is worth evaluating for any team running LLM agents on benchmarks or structured tasks where prompt engineering has plateaued and fine-tuning the model weights is not an option. The trained `best_skill.md` is small, portable, and adds zero inference-time overhead at deployment.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 24 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem SkillOpt Addresses
Most LLM agent frameworks treat the skill document, the system prompt or instruction file that tells the agent how to behave, as a static artifact. It is either hand-crafted, generated once by a strong model, or revised through informal self-reflection. None of these approaches apply the feedback discipline used in gradient-based optimization: systematic measurement of performance, bounded updates, and a held-out validation gate that rejects changes that do not improve things.
SkillOpt treats the skill document as the trainable state of a frozen agent and applies an optimization loop to it. A separate optimizer model turns scored rollouts into bounded add, delete, or replace edits to a single skill document. A candidate edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, a rejected-edit buffer, and epoch-wise update scheduling make training stable. The deployed artifact is a compact `best_skill.md`, typically 300 to 2,000 tokens, that runs against the unchanged target model.
The README describes results across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex CLI, Claude Code CLI). The README states that on GPT-5.5, SkillOpt lifts the average no-skill accuracy by 23.5 points in direct chat, 24.8 inside the Codex agentic loop, and 19.1 inside Claude Code. Optimized skill artifacts are described as transferring across model scales, between Codex and Claude Code harnesses, and to nearby benchmarks without further optimization.
How the Training Loop Works
The core training loop has five stages: rollout, reflect, aggregate, select, update, and evaluate. The optimizer model generates candidate edits to the skill document based on scored rollouts from the target model. Edits are bounded, meaning they can only add, delete, or replace a limited amount of text per step, controlled by the textual learning-rate setting. A held-out validation set gates each update: only edits that improve the validation score are accepted into the next version of the skill document.
The framework uses two model roles: the optimizer, which proposes edits, and the target, which is the frozen agent whose skill document is being trained. These can be different models, or the same model playing both roles. Multi-backend support means the optimizer can be an Azure OpenAI deployment while the target runs through Claude Code CLI.
SkillOpt-Sleep, added in v0.2.0, is a nightly offline variant. The `skillopt-sleep` CLI reviews past agent sessions, replays recurring tasks, and consolidates validated skills behind a held-out gate. It is described in `docs/sleep/README.md`.
Installing SkillOpt
SkillOpt is available on PyPI. The v0.1.0 release announcement in the README states the install command:
pip install skilloptThe base install supports OpenAI and Azure backends. The `pyproject.toml` defines optional extras for additional backends and tools: `claude` for the Claude Agent SDK backend, `qwen` for Qwen via vLLM, `searchqa` for SearchQA data materialization, `alfworld` for the ALFWorld benchmark, `docs` for the MkDocs documentation site, and a WebUI extra for the Gradio dashboard. The extras are installed using pip's bracket syntax with the package name.
Runtime configuration uses environment variables. Copy `.env.example` to `.env` and fill in the credentials for your chosen backend before running training. The `.env.example` shows settings for Azure OpenAI, generic OpenAI-compatible endpoints (including DeepSeek and Novita AI), and per-role overrides: `OPTIMIZER_OPENAI_COMPATIBLE_*` and `TARGET_OPENAI_COMPATIBLE_*` variables allow the optimizer model and target model to use different providers.
The package requires Python 3.10 or newer. Core dependencies include `openai>=1.30.0`, `pyyaml>=6.0`, `numpy>=1.24.0`, `httpx>=0.27.0`, and Azure SDK packages (`azure-identity>=1.15.0`, `azure-core>=1.30.0`).
Supported Backends and Harnesses
SkillOpt distinguishes between chat backends (models used for rollout and optimization via a chat API) and exec harnesses (environments in which the target agent runs commands). Chat backends include `openai_chat`, `claude_chat`, `qwen_chat`, `minimax_chat`, `copilot_chat`, and `openai_compatible`. The `openai_compatible` backend supports any provider that implements the OpenAI Chat Completions protocol, including DeepSeek and Novita AI, without requiring custom code.
Exec harnesses include `codex_exec`, `claude_code_exec`, `cursor_exec`, and `copilot_exec`. These are the environments in which the optimized skill is actually exercised during training rollouts. The `skillopt/model/__init__.py` file is the registration point for all backends.
Plugin and MCP integration files for Claude Code, Codex, Copilot, and Devin are included in the repository but not in the PyPI wheel. They are accessible by cloning the repository. The `plugins/` directory and `.cursor-plugin/` directory contain these integration files.
Configuration and Built-in Benchmarks
Training runs are controlled by YAML configuration files stored in the `configs/` directory. The README points to the documentation at `docs/` for the full configuration reference. The `.env.example` shows the environment variable naming conventions for all supported backends, including per-role overrides for scenarios where the optimizer and target should use different models or endpoints.
SkillOpt ships with six built-in benchmarks. The `skillopt/envs/` directory contains packages for each benchmark. Each package provides an adapter, a data loader, a scored rollout helper, and a YAML config. The SearchQA benchmark requires the optional `searchqa` extra for data materialization. ALFWorld requires `alfworld>=0.4.0` and `gymnasium>=0.29.0`.
The WebUI dashboard provides a visual interface for monitoring training progress. It is separate from the main CLI and requires the `webui` extra.
Limitations and Where SkillOpt Is the Wrong Tool
SkillOpt optimizes a natural-language skill document. It assumes the performance gap between a skilled and unskilled agent on the target task is large enough that text-space edits can meaningfully close it. For tasks where performance is already near ceiling, or where the limiting factor is model capability rather than prompt quality, SkillOpt adds complexity without benefit.
The training loop makes optimizer model calls at training time, not at inference time. The README notes that the deployed `best_skill.md` adds zero inference-time model calls. However, the training run itself is not free: it requires multiple rollouts per epoch, which means API costs or compute time proportional to the number of training epochs times the rollout budget.
The `claude` and `qwen` backend extras pull in `json_repair>=0.61.0` for tolerant JSON parsing of free-form model output. Without this package, malformed JSON from non-OpenAI backends silently drops edits rather than recovering them. The `requirements.txt` lists this dependency as commented out, so it must be installed explicitly or through the appropriate extra.
The six built-in benchmarks cover specific domains. Teams working on tasks outside these domains need to implement a custom benchmark package following the structure described in `docs/` for new benchmark adapters.
SkillOpt vs. GEPA and Other Prompt Optimization Approaches
GEPA (Generative Prompt Advisor) is another approach to automatic prompt optimization, using LLM feedback to suggest prompt improvements. The core distinction SkillOpt draws is the optimization discipline: it applies epoch structure, learning-rate bounding, and a validation gate that rejects non-improving edits. The README describes this as making skill training behave like weight-space optimization in terms of reproducibility, rather than loosely controlled self-revision.
DSPy is another framework that treats prompts as optimizable objects, using a compilation step to tune prompt text or few-shot examples for a given task. DSPy operates over the full pipeline and can tune components individually. SkillOpt focuses on a single skill document for an agent running through a specific harness, rather than a multi-stage pipeline. These tools address overlapping but distinct problem framings.
Maintenance and Licensing
SkillOpt is under active development. The last push was on 2026-09-05. Version 0.2.0 was released on 2026-07-02 and version 0.1.0 on 2026-06-02. The project is licensed under the MIT license.
The MIT license places no restrictions on commercial use, modification, or redistribution. The `best_skill.md` artifact produced by training is a plain Markdown file with no embedded license terms. Teams can deploy it in commercial products without any licensing friction from SkillOpt itself.
The repository includes a SECURITY.md and CONTRIBUTING.md. The paper describing the method is available at arXiv (arxiv.org/abs/2605.23904), and a Microsoft Research blog post and a project page are linked from the README for context on the research background.
Editorial conclusion
SkillOpt is worth evaluating for any team running LLM agents on benchmarks or structured tasks where prompt engineering has plateaued and fine-tuning the model weights is not an option. The trained `best_skill.md` is small, portable, and adds zero inference-time overhead at deployment. Teams with no agentic workload or teams using models that already perform adequately have less to gain. Before starting, check that your target model and harness are among the supported backends listed in `skillopt/model/__init__.py`, and confirm your YAML config paths match the structure in the `configs/` directory.
Frequently asked questions
What is SkillOpt?
SkillOpt is a Microsoft open-source framework that trains natural-language skill documents for frozen LLM agents using a loop of scored rollouts, bounded text edits, and held-out validation gates. The output is a compact `best_skill.md` file that improves agent performance without modifying model weights.
How do I install SkillOpt?
Run `pip install skillopt` for the base package, which supports OpenAI and Azure backends. Install `skillopt[claude]` for Claude support, `skillopt[qwen]` for Qwen via vLLM, or `skillopt[webui]` for the dashboard. Python 3.10 or newer is required.
How do I use SkillOpt?
Copy `.env.example` to `.env` and fill in credentials for your chosen backend, then create a YAML config file referencing one of the built-in benchmarks or a custom benchmark adapter in `skillopt/envs/`. The README points to the `docs/` directory for the full configuration reference and training commands.
How does SkillOpt compare to GEPA?
The README describes SkillOpt as applying optimization discipline (epoch structure, learning-rate bounding, and a validation gate that rejects non-improving edits) that makes skill training reproducible, distinguishing it from loosely controlled self-revision approaches. The README does not include a direct quantitative comparison to GEPA.
What are self-evolving skills in SkillOpt?
Self-evolving skills in SkillOpt are natural-language skill documents that improve automatically through a training loop of rollouts, scored edits, and validation gates. SkillOpt-Sleep extends this with a nightly offline engine that reviews past sessions, replays recurring tasks, and consolidates validated improvements.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-skillopt)