MLE-Agent: a CLI pairing agent for ML baselines and Kaggle runs
🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 OpenAI, Anthropic, Gemini, Ollama, etc supported. :fireworks: Code RAG
At a glance
- What is it?
- MLE-Agent is a Python CLI that plans and writes ML code from a project directory, pulling in arXiv and Papers with Code references. It installs from PyPI as mle-agent and exposes the mle command, but the README leaves rollback and cost control undocumented.
- Who is it for?
- Adopt MLE-Agent if you want a terminal agent that scaffolds a baseline project with mle new and mle start, and you are comfortable reading generated code before running it. Skip it if you need a documented rollback path, cost controls, or a stable API for automation; the README does not describe any of those.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 81 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap MLE-Agent targets: requirements in, runnable ML code out
Most ML work starts with a sentence and ends with a repository. MLE-Agent is aimed at that gap. The README describes it as a pairing LLM agent for machine learning engineers and researchers, and the use cases it lists are concrete: prototype an ML baseline from a vague requirement such as predicting stock prices from historical data, run a Kaggle competition end to end, and generate a weekly work report from a Git repository. The intended user is someone who already knows what a training loop looks like and wants the scaffolding, the file layout and the debugging loop handled by an agent rather than typed by hand. It is not a notebook replacement and it is not a hosted service. Everything runs from a project directory you create on your own machine, which means the agent writes files into your working tree. That single design decision shapes most of the trade-offs below.
How the agent is put together: CLI, project directory, tools
The architecture visible in the repository is a Python package named mle with a Click-based entry point. The pyproject.toml declares two console scripts, mle-agent and mle, both pointing at mle.cli:cli, so the two names are interchangeable. The dependency list is the clearest statement of what the agent actually does: openai for model calls, google-genai for Gemini, tavily-python for web search, GitPython for reading commit history, kaggle for competition data, lancedb and tantivy for local retrieval, tree-sitter for parsing code, mem0ai for memory, and fastapi with uvicorn for the report web application. A tool layer wraps those libraries, and the agent calls into it while working inside a project directory created by mle new. The retrieval side is local: lancedb and tantivy suggest an on-disk index rather than a hosted vector store, which is why the README can advertise a personal coding assistant without a server. The arXiv and Papers with Code integration sits on top of the same tool layer, feeding references into planning. There is also a bench extra in pyproject.toml that pulls mlebench from a Git URL, restricted to Python 3.11 and 3.12, which is the only place the repository points at an evaluation harness.
Install MLE-Agent and run your first baseline
The README gives two install paths. The short one is PyPI, and it is the one to start with. The -U flag upgrades an existing install, so re-running it is how you move between versions.
pip install -U mle-agentIf you use uv, the README lists the equivalent command as uv pip install -U mle-agent. The source path is documented as well: clone the repository, create a virtual environment, and install in editable mode. The README shows uv venv .venv followed by source .venv/bin/activate on Linux or macOS, then pip install -e . from the repository root.
Once installed, the first real use is creating a project. Run this from the directory where you want the project to live, not from inside an existing one.
mle new <project name>The README states that a project directory is created under the current path and that you must start the project from inside that directory. So the next two commands are a cd and a start.
cd <project name>
mle startmle start is the command that drives baseline prototyping. The README says the requirements you give it can be vague. What you should see is the agent planning, writing files into the project directory, and attempting to run and debug the code it produced. If you prefer a conversational loop instead of a one-shot run, the README documents mle chat as an interactive terminal chat under the project directory. For report generation there are two modes: mle report starts a local web application that the README says you reach at http://localhost:3000/, and mle report-local works from a local Git repository using --email, --start-date, --end-date and a repository path, with the date flags optional and defaulting to the last seven days.
The Kaggle mode and what --auto actually requires
The most opinionated feature is the Kaggle integration. The interactive form is a single command from inside the project directory.
mle kaggleThe automated form is where the constraints become visible. The README shows mle kaggle --auto with --datasets, --description, --submission, --sub_example and --comp_id, and it states plainly that you must have joined the competition before running the command. That is a real precondition, not a formality: the kaggle dependency is in the install list, but credentials and competition membership are on you. The --datasets flag takes a comma-separated list, and --description accepts either a file path or literal text, which is a small but useful detail because it means you can pass a paragraph without creating a file first. The --auto path is described as completing a task without human interaction, which is the strongest claim in the README and also the one to treat most carefully. An unattended agent writing and executing code against a competition dataset is exactly the scenario where you want the run to be reproducible and bounded, and the README does not describe either property.
Where MLE-Agent is the wrong tool
Three limitations stand out. First, the README does not document rollback. The agent writes into a project directory and executes code, but there is no described mechanism for reverting a change it made or recovering a file it overwrote. If that matters to you, the answer is version control outside the tool, and the README does not say that either. Second, cost and rate limits are absent from the documentation. The agent calls OpenAI, Anthropic, Gemini or Ollama depending on configuration, and a long mle start session with debugging loops can issue many calls. Nothing in the README describes a budget, a token ceiling or a dry-run mode. Third, the release history and the version metadata disagree in a way worth noting. The most recent release listed is 0.4.2 from 2024-10-12, while pyproject.toml declares version 0.4.3, and the last push to the default branch was on 2026-07-10. So the repository has moved since the last tagged release, and installing from PyPI may give you something older than the source tree. That is a normal state for a project of this kind, but it means the README and the code you run are not guaranteed to match. Finally, if your work is exploratory analysis in a notebook, or a pipeline that must be deterministic and auditable, an agent that rewrites files is the wrong layer to add.
How it differs from a general coding agent
The closest alternative category is a general-purpose coding agent such as Aider or Claude Code, and the difference is in what the tools are wired to. A general agent edits files and runs shell commands; MLE-Agent ships a dependency list aimed at ML specifically. That list includes kaggle for competition data, pandas, tree-sitter for parsing Python source, lancedb and tantivy for a local index, and tavily-python for web search. The README also names arXiv and Papers with Code as sources for state-of-the-art methods, which a general agent would only reach through ad hoc search. The second difference is the project directory convention. mle new creates a structure and mle start operates inside it, so the agent has a known place to write. A general agent works wherever you point it. The trade-off runs the other way too: a general coding agent is not tied to a project layout, so it fits an existing repository, while MLE-Agent expects you to start from mle new. If your codebase already exists and you want targeted edits, the project-directory model is friction rather than help. If you are starting from a blank requirements sentence, it removes a decision.
Licence, maintenance and upgrade cost
The README badge and GitHub metadata identify the licence as MIT, while pyproject.toml declares license = {text = "Apache-2.0"}. Those are both permissive licences, and both allow commercial use, but they are not the same document, and a redistribution question cannot be answered from the repository alone. Check the LICENSE file in the repository root before you rely on either. This is a factual discrepancy, not legal advice. On maintenance, the last push to the default branch was on 2026-07-10, so the repository is not dormant, but the newest tagged release is 0.4.2 from 2024-10-12. Upgrades therefore carry a specific cost: if you install from PyPI you get a released version, and if you install from source you get whatever is on main, and the two are not the same. The dependency list is also wide, spanning openai, google-genai, fastapi, uvicorn, lancedb, tantivy, mem0ai and tree-sitter, with pinned versions on several of them. A wide dependency set with pins means upgrades are not a single-command affair in practice; expect to resolve conflicts when one of the pinned packages moves. Python 3.9 or newer is required, and the bench extra narrows that to 3.11 or 3.12.
Editorial conclusion
Adopt MLE-Agent if you want a terminal agent that scaffolds a baseline project with mle new and mle start, and you are comfortable reading generated code before running it. Skip it if you need a documented rollback path, cost controls, or a stable API for automation; the README does not describe any of those. Before installing, check the current version on PyPI against the 0.4.2 release notes and confirm Python 3.9 or newer, since the project metadata in pyproject.toml requires it.
Frequently asked questions
What does MLE-Agent install as, and what command do I run?
It installs from PyPI as mle-agent, and the README shows pip install -U mle-agent. The package declares two console scripts, mle and mle-agent, both pointing at the same entry point, so either name works.
How do I start a first project with MLE-Agent?
Run mle new with a project name, then cd into the created directory and run mle start. The README states the project directory is created under the current path and that the project must be started from inside it.
Does MLE-Agent need a Kaggle account for the Kaggle mode?
Yes. The README says to make sure you have joined the competition before running mle kaggle, including the --auto variant.
What port does the MLE-Agent report web application use?
The README says to run mle report from the project directory and then visit http://localhost:3000/ to generate the report locally.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mlsysops-mle-agent)