# LLM4AD: a Python platform for algorithm design driven by large language models

> LLM4AD wraps LLM-backed evolutionary search around a task interface, so you can describe a problem and let a model write and refine the solver. It installs with pip, requires Python 3.9 to 3.12, and ships a GUI, examples and a BSD-2-Clause licence.

**Optima-CityU/LLM4AD** — LLM4AD: A Platform for Algorithm Design with Large Language Model

- Repository: https://github.com/Optima-CityU/LLM4AD
- Website: http://www.llm4ad.com
- Stars: 782 · Forks: 103
- Language: Python
- License: BSD-2-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/optima-cityu-llm4ad

## The problem LLM4AD addresses: writing the heuristic, not just tuning it

Most optimisation tooling assumes you already know the algorithm. You pick simulated annealing, a genetic algorithm or a greedy rule, then tune its parameters. LLM4AD attacks the step before that. Its stated purpose is automatic algorithm design: the LLM writes the algorithm itself, and a search loop keeps the versions that score better.

The README says the platform was originally developed for optimisation tasks and that the framework is versatile enough for machine learning, science discovery, game theory and engineering design. That is the intended audience: researchers and engineers who have a well-defined objective function and no strong prior about which heuristic shape will win. If you already know the right algorithm and only need to fit three numbers, this is more machinery than you need.

The project is maintained by the Optima group at City University of Hong Kong and is distributed under BSD-2-Clause. A survey paper on LLMs for algorithm design, accepted by ACM Computing Surveys, is cited in the news section, and the platform paper itself is on arXiv. Those references are the place to look for the methodology and benchmark results; the README does not reproduce them.

## How the search loop works: an LLM proposes, a profiler scores, the population evolves

The architecture visible in the README is three layers. Tasks define a problem and an evaluation function. Methods run the search. LLMs are pluggable endpoints behind one interface.

The quick-start example imports OBPEvaluation from llm4ad.task.optimization.online_bin_packing, HttpsApi from llm4ad.tools.llm.llm_api_https, and EoH plus EoHProfiler from llm4ad.method.eoh. EoH stands for Evolution of Heuristics. The method samples candidate programs from the LLM, evaluates each one through the task's evaluation class, and uses the scores to decide what to ask for next. The profiler records the run.

Two features in the README's table matter for anyone running this on a shared machine. Evaluation uses multiprocessing, and there is what the project calls secure evaluation: main process protection and timeout interruption. That second one is the admission that generated code sometimes hangs or crashes. The platform runs candidate programs it did not write, so a timeout and an isolated process are not luxuries. Logging goes to local files, Wandb or Tensorboard, and runs can be resumed.

The README lists support for other programming languages, more search methods and more task examples as coming soon. Treat the Python-only scope as the current boundary, not a temporary one.

## Installing LLM4AD and running a first EoH search

The README is explicit that the Python version must be at least 3.9 and less than 3.13, and it suggests a conda environment. Two install paths exist: from a local clone, or from PyPI.

```bash
cd LLM4AD
pip install .
```

Or, without cloning:

```bash
pip install llm4ad
```

Optional pieces are installed separately. Numba is needed for Numba acceleration, tensorboard for the Tensorboard logger, wandb for the Wandb logger, gym for GUI and machine learning tasks, and pandas for science discovery tasks. The README warns that the gym version may conflict with your environment and points to gym's own documentation for the right version. The GUI additionally wants everything in requirements.txt.

The quick-start script configures an LLM endpoint and runs EoH on online bin packing. The README's note gives the three fields to set: host, key and model, with api.deepseek.com and deepseek-chat as the worked example.

```python
from llm4ad.task.optimization.online_bin_packing import OBPEvaluation
from llm4ad.tools.llm.llm_api_https import HttpsApi
from llm4ad.method.eoh import EoH, EoHProfiler

if __name__ == '__main__':
    llm = HttpsApi(
        host='xxx',   # your host endpoint, e.g., api.openai.com, api.deepseek.com
        key='sk-xxx', # your key, e.g., sk-xxx
    )
```

The README truncates the snippet at the key argument, and it does not show the EoH constructor call or the arguments to EoHProfiler. For those, the documentation site at llm4ad-doc.readthedocs.io and the example directory are the places the project points to. The repository also ships a Colab notebook under example/online_bin_packing/, which is the fastest way to see the full script without installing anything locally.

## Where LLM4AD breaks down, and when it is the wrong tool

The evaluation function is the bottleneck. Every candidate the LLM writes has to be scored, and the search keeps generating candidates. If a single evaluation takes minutes, a run that samples hundreds of programs becomes an overnight job at best. The multiprocessing support helps with throughput, but it does not change the arithmetic.

The second constraint is the LLM endpoint. LLM4AD does not ship a model. You supply a host, a key and a model name, and the quality of what comes back depends entirely on that choice. A model that produces code that does not parse, or that ignores the required function signature, wastes iterations. The README does not document a retry or repair policy for malformed responses, and it does not document rollback for a run that has produced worse candidates than the starting point.

Third, the platform executes generated code. The README's secure evaluation feature addresses this, but timeout interruption is a mitigation, not a sandbox. Running LLM4AD on a machine with credentials you care about is a decision the documentation does not walk you through.

Finally, if your problem has a known exact solver or a mature library implementation, generated heuristics are unlikely to beat it. LLM4AD is for problems where no good hand-written heuristic exists yet.

## LLM4AD compared with LLaMEA and with general evolutionary code search

The requirements file installs llamea directly from its GitHub repository, and the file's own comment marks it as a new method. So LLaMEA is not a competitor bolted on from outside; it is a method that runs inside LLM4AD's task and LLM interfaces. The difference is in the search strategy, not in what you can evaluate. Choosing between them is choosing a method, not a platform.

Against general LLM code-generation tools, the distinction is the evaluation loop. A coding assistant produces one answer and stops. LLM4AD produces candidates, scores them against your objective, and feeds the scores back into the next prompt. That closed loop is the whole point, and it is also why the cost model is different: you pay per iteration, not per answer.

Against AlphaEvolve and OpenEvolve, which appear in the related searches around this project, the README makes no comparison and this article will not invent one. What can be said from the repository alone is that LLM4AD ships a GUI, a task hierarchy spanning optimisation, machine learning and science discovery, and a documented Python API. Those are the concrete differences a reader can check.

## Licence, maintenance and the cost of keeping up

The repository is not archived and the last push was on 2026-06-30. The single release listed is v1.0.0 from 2025-03-20. That gap between a moving main branch and a single tagged release is the practical maintenance fact: if you install from PyPI you get the release, and if you install from a clone you get whatever main looked like on the day you cloned.

The licence file is BSD-2-Clause, but setup.py carries a narrower statement: permission is granted to use the LLM4AD platform for research purposes, and publications or software using it must acknowledge LLM4AD and cite the paper by Liu, Zhang, Xie, Sun, Li, Lin, Wang, Lu and Zhang. The same file directs commercial licensing enquiries to the project's contact page. Read both before you ship anything, and treat the acknowledgement requirement as a real obligation rather than a formality. This is a description of what the files say, not legal advice.

Upgrade cost is dominated by dependencies rather than by LLM4AD's own code. numpy is pinned below 2, tree-sitter-python is pinned at 0.23, and llamea is pulled from a branch rather than a tag. A branch dependency means two installs a week apart can differ without any version number changing.

## Conclusion

Adopt LLM4AD if you have a scoring function you can call cheaply and you want to see whether an LLM can write a better heuristic than the one you would hand-code. Do not adopt it if your evaluation is slow, non-deterministic or expensive, because the search loop calls it many times. Verify three things first: that your Python is between 3.9 and 3.12, that your LLM endpoint returns code in the expected format, and that EoH's default budget fits your API spend. The repository is not archived and the last push was on 2026-06-30, so the code is still moving; pin the release you install.

## FAQ

### What Python version does LLM4AD require?

The README states the version must be at least 3.9 and less than 3.13, and setup.py declares python_requires as >=3.9,<3.13. The project suggests running it in a conda environment.

### How do I install LLM4AD?

Either clone the repository and run pip install . from inside it, or run pip install llm4ad directly from PyPI. Optional extras such as numba, tensorboard, wandb, gym and pandas are installed separately depending on which logger or task family you use.

### Do I need an API key to run LLM4AD?

Yes. The quick-start example configures an HttpsApi object with a host, a key and a model name, using api.deepseek.com and deepseek-chat as the worked example. LLM4AD does not bundle a model.

### What licence is LLM4AD released under?

The repository is BSD-2-Clause, while setup.py states that permission is granted for research purposes and that works using the platform must acknowledge LLM4AD and cite the platform paper. Commercial licensing enquiries are directed to the project's contact page.

## Sources

- [License: BSD-2-Clause](https://github.com/Optima-CityU/LLM4AD/blob/main/LICENSE)
- [Optima-CityU/LLM4AD on GitHub](https://github.com/Optima-CityU/LLM4AD)
- [Project website](http://www.llm4ad.com)
- [README](https://github.com/Optima-CityU/LLM4AD/blob/main/README.md)
- [Releases](https://github.com/Optima-CityU/LLM4AD/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/optima-cityu-llm4ad
