Model or dataset
JackHopkins/factorio-learning-environment avatar
JackHopkins/factorio-learning-environment

The Factorio environment's agent interface is a Python interpreter, with the game's own schemas at the root

A non-saturating, open-ended environment for evaluating LLMs in Factorio

1,189 stars102 forksPythonNOASSERTION

At a glance

What is it?
An evaluation harness where an LLM agent is given no action space at all: it writes Python programs, the environment runs them against a Factorio server, and the output stream of the last program is the next observation. What the packaging adds is where the interest lies, because the base install already contains everything the three advertised extras are for, the manifest sits five patches ahead of the newest release, and the front page links to documentation for version 0.3.0.
Who is it for?
This suits anyone evaluating whether a model can plan over a long horizon in a world that does not end, because a code interface measures reasoning about state rather than about a fixed action list, and the repository ships the game's own interface schemas so the model is not guessing function names.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The manifest is five patches ahead of the newest release

Three version numbers describe this project and none of them agree. The packaging manifest declares 0.4.8. The newest release tag is 0.4.3, cut in April. And the project homepage, which is also the target of the Website link in the readme's own header row, points at documentation for version 0.3.0 specifically. So the page a visitor lands on describes a build two minor versions behind the newest tag and five patch versions behind the manifest. The release titles are more informative than the numbers, because each one names a single change rather than a version: the newest adds a sandbox evaluation through an inspection framework, the one before adds test coverage for lab observations, and the one before that is a hotfix to a directional system. That pattern, one substantive change per tag with a name attached, is easier to track than version numbers alone and is worth reading before deciding which tag to pin.

The base install already contains everything the three extras are for

The readme advertises three extras and a combined one, for running experiments, for agent protocol support and for PostgreSQL. Every one of those capabilities is already a base dependency. The base list includes an evaluation harness and its agent-to-agent package, the MCP client, a PostgreSQL driver, the Docker client, a plotting library, a web framework and its server, two model provider SDKs, a model hub client with its own helper package, a formatter, a test runner, an HTTP client and an async HTTP server. So `pip install factorio-learning-environment` on its own gives you everything the extras select, and the extras exist mainly to name subsets of what you already have. That is a packaging choice rather than a bug, but it has a cost worth naming: the floor install for someone who only wants to drive a Factorio server programmatically is the size of the whole research stack, and pinning that stack is how this project breaks.

Two dependencies are pinned exactly, and one backports a standard library module

Most of the dependency list floats with a lower bound, which is the right default for a research harness. Three entries do not. The game's remote console client is pinned to an exact version, which is defensible because it is the channel everything else depends on and a protocol change there would be invisible. The agent-to-agent SDK is also pinned exactly, at version 1, with no explanation of why it cannot move. And a TOML parser backport is listed unconditionally, even though the project requires Python 3.10 or newer and classifiers run through 3.13, which means every supported interpreter except the oldest already has that module in its standard library. So the dependency set mixes three policies in one list: exact pins where breakage would be silent, lower bounds where the ecosystem should be allowed to move, and a backport shipped to everyone for the benefit of the single oldest supported version.

The agent interface is a Python interpreter, not an action space

The design decision that defines this environment is that there is no action space. The loop is described in three steps and each one is about code. Observation is whatever the agent saw on the output streams of its last program, both standard output and standard error. Action is a Python program the agent generates. Feedback is the environment executing that program, assigning variables, adding classes and functions to the namespace, and handing back an output stream. So the policy is a code generator and the observation is a log, and the words observation, action and feedback are used in the shape of a reinforcement learning loop while describing a read-eval-print loop. The dependency list confirms how the code reaches the game: a Lua runtime embedded in Python, a string-building bridge for talking to it, the remote console client, a binary parsing library, an image library for screenshots and a numerical library. There is no tool schema and no JSON action format anywhere in the loop, which is what makes the result hard to compare against a conventional benchmark and easy to extend.

The game's own interface schemas sit at the repository root

Two files at the root are unusual enough to be the answer to the obvious objection about a code interface. One is a machine-readable index of the game's Lua documentation and the other is its GraphQL schema, both checked in at the top level rather than fetched at runtime. In a loop where the model invents Python that has to call the game's API correctly, those two files are the grounding: the model can be pointed at real function signatures and real query shapes instead of recalling them. It also explains a dependency that otherwise looks out of place, a config parser and a binary format library, since reading and re-emitting those schemas is part of the setup rather than an afterthought. The pattern is worth copying regardless of the game: if your agent writes code against an external API, ship that API's schema. The cost is that the files are a snapshot, so a game update makes them stale, and nothing in the readme says who refreshes them or how.

The quickstart points at a configuration directory the repository does not have

The quickstart is three commands. Activate the environment, start a Factorio cluster, then run an evaluation against a configuration file:

bash
# Activate venv
source .venv/bin/activate

# Start Factorio cluster
fle cluster start

# Run evaluation trajectories (requires [eval] dependencies)
fle eval --config configs/gym_run_config.json

That configuration path does not exist in the repository. The top level holds a library package, tests, documentation, examples, a data directory, two schema files and the usual project files, and there is no configuration directory among them. The example agent directory under examples is the only place example code is grouped, and the named configuration is not in it as far as the listing shows. So the one command a new user is most likely to run after installing points at a file they would have to reconstruct, and the error they would hit is a missing path rather than anything about the environment. Everything else in that block is self-contained, which makes the gap more noticeable rather than less.

A hotfix and a coverage release went out 71 seconds apart

The release history has one detail that says more about how the project is worked on than any feature list. Two consecutive releases were published on the same evening 71 seconds apart, and they are a hotfix to a directional system followed by a change adding test coverage for lab observations. A third release followed ten days later adding a sandbox evaluation through an inspection framework. All three are from late March and early April, while the branch itself was pushed to in mid-September, so roughly five and a half months of commits sit behind the newest tag with no release in front of them. For a research harness that is an understandable rhythm, since the interesting artifact is the recorded trajectories rather than the package. It does mean the pinned version is a moving target in one direction only: nothing tells you which commit the leaderboard entries were produced at, and the leaderboard, the paper and the documentation all live outside this repository.

Docker is a prerequisite and the game itself is only needed to render

The prerequisites are three lines and one of them is conditional in an instructive way. Docker. Python 3.10 or newer. And the game, at version 2.0.73 or later, marked as needed only for optional rendering. That split tells you the architecture: the agent does not need to see the game, it needs a server it can talk to, and the game client is a debugging convenience. It is also why the Docker client is a base dependency and why the first command in the quickstart starts a cluster rather than launching anything locally. The readme is explicit that the framework is for developing and evaluating agents, and the contributing section frames the goal as open-ended evaluations that frontier models cannot saturate, spelled in one place with a word that does not exist. That framing is the actual research claim here, and it is also the hardest to verify from inside the repository, since the tasks, the scoring and the saturation argument are all on the external site.

Editorial conclusion

This suits anyone evaluating whether a model can plan over a long horizon in a world that does not end, because a code interface measures reasoning about state rather than about a fixed action list, and the repository ships the game's own interface schemas so the model is not guessing function names. It is a poor fit if you want a standard reinforcement learning interface, since the observation is a text stream rather than a vector, and if you want a light install, since the base package pulls an evaluation harness, two model providers, a database driver and a web framework. Before you build against it, check three things: the configuration path in the quickstart is not in the repository, the newest release is five patches behind the manifest, and Docker plus a headless server are part of the setup rather than an optional extra.

Frequently asked questions

What does an agent actually do in the Factorio Learning Environment?

It writes Python programs. Observation is the output streams of the agent's previous program, action is a newly generated Python program, and feedback is the environment executing it, binding variables and adding classes and functions to the namespace before returning an output stream.

Do I need Factorio installed to run an evaluation?

Only for optional rendering. The stated prerequisites are Docker and Python 3.10 or newer, with the game at version 2.0.73 or later listed as needed just for rendering. The first command in the quickstart starts a cluster rather than launching the game.

What does the base pip install of factorio-learning-environment pull in?

More than the readme's extras suggest. The base dependency list already includes the evaluation harness, the agent protocol packages, a PostgreSQL driver, the Docker client, a plotting library, a web framework and server, two model provider SDKs and a formatter, so the eval, MCP and PostgreSQL extras mostly select subsets of what is already installed.

How do I start an evaluation run?

Activate the environment, run `fle cluster start` to bring up the Factorio cluster, then run `fle eval --config configs/gym_run_config.json`, which the quickstart notes requires the eval extras. That configuration path is not present in the repository listing.

What license is the Factorio Learning Environment under?

The packaging metadata declares MIT with a matching classifier, and a LICENSE file sits at the repository root, so the project files agree with each other. The repository-level metadata records no licence name at all.

Official sources

  1. Issues
  2. JackHopkins/factorio-learning-environment on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jackhopkins-factorio-learning-environment.svg)](https://hysenlabs.com/projects/jackhopkins-factorio-learning-environment)