Model or dataset
Optim-Agent/optim-agent avatar
Optim-Agent/optim-agent

optim-agent: the agent proposes, the declared space decides

LLM agents as your hyperparameter optimizer.

800 stars43 forksPythonMIT

At a glance

What is it?
optim-agent wraps a coding agent as a sampler for parameter search and keeps objective evaluations authoritative over it. Its packaged version, its only release, and one row of its own benchmark table do not line up.
Who is it for?
optim-agent is worth a look when you have a system with configurable parameters, an objective you can measure honestly, and evaluations expensive enough that a ten-trial budget matters. It is the wrong tool when the space is cheap to explore exhaustively, and its own table says as much: at ten trials the classical TPE baseline lands worse than random search on Branin.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 52 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

TPE lands worse than random search on Branin at ten trials

The hard-function benchmark is the most informative table in the project, and one row of it does not behave. Agents receive no task context at all here: only generic `x1...x5` parameter names, numeric bounds and the trial history, run over 10 trials and five seeds, with Random and TPE left as unchanged baselines.

On Branin 2D, Random reaches a mean best of 5.008 and TPE reaches 11.395. On Ackley 5D the ordering flips, with TPE at 18.843 against Random at 19.639. So on one of the two functions the classical surrogate is worse than random sampling under the same budget, and no explanation is offered for that row. The agent rows sit far below both: Opus-4.8 at 0.398 and 0.061, GPT-5.5 at 1.326 and 3.960, Sonnet-5 at 3.850 and 0.143, Kimi-K3 at 2.082 and 0.907, Minimax-M3 at 0.970 and 0.574, GLM-5.2 at 3.609 and 15.023.

A reader deciding whether to pay for agent calls on a cheap objective should weigh that comparison rather than the agent rows alone.

The trajectory figure is one seed, the tables are five

The benchmark section opens by describing a seed-0 Branin trace that compares TPE and GPT-5.5 under the same 10-trial budget, with incumbent values plotted after each trial. It labels itself a trajectory illustration and says aggregate results and reproduction commands follow. The tables that follow report means over five seeds, so the opening figure and the tables are not the same measurement.

Model identity is pinned in prose right after the first table, as `gpt-5.5`, `claude-opus-4-8`, `claude-sonnet-5`, `kimi-k3`, `MiniMax-M3` and `glm-5.2`. Note that the row label and the pinned identifier differ: the table row reads Opus-4.8 while the identifier is `claude-opus-4-8`. Any reproduction has to use the identifiers, not the row labels.

A second table for OpenCode agents on free tiers starts beneath it and gives only its first data row before the document ends. Those rows exist and are meant to be compared against the paid ones, but the visible page does not carry them.

The package says 0.2.0 while the only release is 0.1.1

The packaging metadata and the release history disagree by a minor version. `pyproject.toml` declares `version = "0.2.0"`, while the project's single published release is 0.1.1, dated 2026-07-14. There is no 0.2.0 release to match, and the repository was pushed to again on 2026-08-14, after that tag.

The classifier is more conservative than either number: `Development Status :: 3 - Alpha`. Combined with one release and an empty-looking upgrade path, that tells you what kind of dependency this is. `requires-python` is `>=3.9`, the ruff target is `py39`, and the build system asks for `setuptools>=77.0.3` with a PEP 639 style `license = "MIT"` and `license-files`.

For an installed copy this matters in a specific way: pinning by version number gives you 0.1.1, while tracking the branch gives you metadata claiming 0.2.0. Neither tells you which of the two corresponds to the code you are reading.

The agent proposes, the declared space disposes

The design decision that separates this from an agent that edits your code is stated once and then repeated as a property: objective evaluations remain authoritative. The agent proposes values, optim-agent validates them against the declared space, records the outcome, and falls back to safe sampling when an agent reply is invalid. An unusable reply costs a trial, not correctness.

That boundary is enforced through the two suggest calls your objective function makes. `trial.suggest_float("threshold", 0.05, 0.95, context=...)` declares a continuous dimension with its own bounds, and `trial.suggest_int("budget", 10, 200, log=True, context=...)` declares an integer dimension sampled on a log scale. Anything outside those bounds is not a candidate.

The optional `context` argument is where the semantic part lives, and it has three separate entry points: study-wide on `AgentSampler(context=...)`, per parameter on each suggest call, or both. Combined with the trial history, that context is what the agent reasons over, which is also why the context-free benchmark above is the harder of the two setups.

Three install routes produce three different artifacts

Installation is documented three times, and each route installs something different. The Codex route adds a skill:

text
$skill-installer install https://github.com/Optim-Agent/optim-agent

That line is published as plain text rather than as a shell command, so it is the one route where you have to know which tool consumes it. The Claude Code route adds a marketplace and then a plugin:

bash
claude plugin marketplace add Optim-Agent/optim-agent && claude plugin install optim-agent@optim-agent

The Python route is a normal package install, from PyPI or straight from the repository:

bash
python -m pip install optim-agent
python -m pip install "optim-agent @ git+https://github.com/Optim-Agent/optim-agent.git"

Whichever you pick, one authenticated agent CLI has to be on `PATH`: `claude`, `codex` or `opencode`. That dependency is not expressed in the package metadata, so a clean environment can satisfy `pip` and still fail to run.

Five extras groups and one packaged module

Optional dependencies are grouped by use case rather than listed flat, and the groups tell you what the examples need. `examples` pulls numpy, matplotlib, optuna and scikit-optimize, which is the tell that the comparison baselines in the benchmarks are Optuna-class tools rather than something hand-rolled. `rl` adds gymnasium with box2d and imageio, `vision` adds torch and torchvision, `ml` adds scikit-learn, pandas and xlrd, and `dev` adds build, pre-commit, pytest, pytest-cov and ruff.

Two details are worth noting for anyone vendoring this. The optional list contains no agent SDK at all, because the agent is reached by running a CLI on PATH rather than by calling a library. And the packaging table declares `packages = ["optim_agent"]` as a literal list with no discovery directive, so what ships is exactly that one top-level package.

The repository root carries the rest of the surface: `SKILL.md`, a `plugins/` directory, a `benchmarks/` directory, a `paper/`, a `ROADMAP.md`, and eight translations of the README under `docs/i18n/`.

The quickstart stops inside the effort argument

The quickstart is the shortest path to a working study and the visible page ends in the middle of it. What is there sets up an objective function that declares its own dimensions, then begins the study:

python
import optim_agent as oa

def objective(trial):
    threshold = trial.suggest_float(
        "threshold", 0.05, 0.95,
        context="decision threshold; higher values trade recall for precision",
    )
    budget = trial.suggest_int(
        "budget", 10, 200, log=True,
        context="compute or operating budget; larger values may improve quality",
    )
    return evaluate_system(threshold=threshold, budget=budget)  # domain code

The study construction that follows names `direction="maximize"` and an `AgentSampler` whose `backend="claude"` is commented as interchangeable with `codex` or `opencode`. The next argument begins `effort=` and the page stops there, so the accepted values for that argument are not written down anywhere in the visible text.

Eight runnable examples sit under `examples/`, named for the job rather than the API: `cifar10.py`, `credit_card.py`, `hard_functions.py`, `inference_tuning.py`, `mnist.py`, `quickstart.py`, `rl_control.py` and `sklearn_tuning.py`.

Editorial conclusion

optim-agent is worth a look when you have a system with configurable parameters, an objective you can measure honestly, and evaluations expensive enough that a ten-trial budget matters. It is the wrong tool when the space is cheap to explore exhaustively, and its own table says as much: at ten trials the classical TPE baseline lands worse than random search on Branin. Before adopting it, check which version you actually install, because the package metadata reads 0.2.0 while the only release tag is 0.1.1 dated 2026-07-14 and the classifier still reads Alpha, confirm that one authenticated agent CLI is on PATH, and read the row for your own objective rather than the seed-0 trajectory figure. The last push to main was on 2026-08-14.

Frequently asked questions

What does optim-agent actually change in my system?

It changes values you declare inside your objective function through `trial.suggest_float` and `trial.suggest_int`, each with its own bounds. The agent proposes those values, optim-agent validates them against the declared space, records the measured outcome and falls back to safe sampling when a reply is invalid, so objective evaluations stay authoritative.

Which coding agents can optim-agent drive?

One authenticated agent CLI has to be on `PATH`: `claude`, `codex` or `opencode`. The sampler takes the choice as `backend="claude"`, with `codex` and `opencode` named as the alternatives, and no agent SDK appears in the package dependencies because the agent is invoked as an external process.

What does optim-agent need installed alongside it?

Python 3.9 or newer, since `requires-python` is `>=3.9`. Optional groups cover the examples, reinforcement learning, vision, machine learning and development work. The groups are separate installs, so an RL example needs the `rl` group with gymnasium and imageio rather than the base package alone.

What does an optim-agent study keep on disk?

Studies are retained as JSON or SQLite and hold configurations, outcomes, states, the context you supplied and, optionally, the agent's own rationale. That record is the point of the audit trail: the proposed value, the declared space it was checked against and the measured result all stay inspectable after the run.

Official sources

  1. License: MIT
  2. Optim-Agent/optim-agent on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/optim-agent-optim-agent.svg)](https://hysenlabs.com/projects/optim-agent-optim-agent)