# AMA-CMFAI/LAMBDA: a data analysis agent that runs the code it writes, on your machine

> LAMBDA turns a natural-language question about an uploaded dataset into an executable analysis workflow, with a FastAPI backend, a React frontend and one persistent workspace per conversation. Its own documentation tells you it executes generated code on the host, and its manual backend command binds every interface.

**AMA-CMFAI/LAMBDA** — This is the offical repository of paper "LAMBDA: A large Model Based Data Agent". https://www.polyu.edu.hk/ama/cmfai/lambda.html

- Repository: https://github.com/AMA-CMFAI/LAMBDA
- Website: https://arxiv.org/abs/2407.17535
- Stars: 596 · Forks: 59
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ama-cmfai-lambda

## Three release tags, and only one of them names a version

The release list is the first thing to read, because it is not a version history.

There are three tags. One is a name for the current line and its title says the second version and the web application were released. One is named after what it holds, a beta application build, with a date-stamped version inside the title. And one is named after an environment, with an alpha version and a date inside the title.

So of three releases, one is a code line, one is a pre-release build of the app, and one is a production build. None of the three uses plain semantic versioning as its tag.

The dates are more useful. The current line tag is from 3 July 2026, the last push to the default branch is dated 11 August 2026, and the repository is not archived. So there are about five weeks of commits sitting on the branch with no tag to point at.

The readme resolves the naming question in a Versions section rather than leaving it implicit: the latest code is on the main branch tagged as the current line, and the previous open-source version is preserved on a separate named branch with its own earlier tag.

## The manual command binds every interface, and the note says not to

There is a runtime warning at the end of the readme, and one documented command contradicts it.

The warning is unambiguous. LAMBDA executes analysis code on the machine where the backend is running, and the instruction is to use it with trusted users and trusted data, and not to expose the backend directly to untrusted traffic without adding sandboxing or other isolation.

Now compare the manual setup section. The backend start command there binds the server to all interfaces on port 8000 with reload enabled. That is the command a reader copies when they do not use the start script.

So the two halves of the same document give opposite guidance: one says do not expose this thing, the other shows a command that exposes it to the local network.

It is a plausible gap rather than a contradiction in intent, because the start script is the recommended path and may bind differently. But the readme does not say so, and the manual section is presented as an equally valid route with the same result.

For anyone taking the manual path, that host binding is the line to change first.

## An OpenAI-shaped endpoint with a model list that is not OpenAI's

The model interface is deliberately shaped like one large provider's API and pointed at somebody else's models.

LAMBDA talks to models through an OpenAI-style chat completions interface. Three settings in the backend configuration decide everything: an API key, a base URL that should be the base of any OpenAI-compatible API, and a model list that is a JSON array of model identifiers.

The call it makes is the chat completions path appended to that base URL, nothing more exotic.

The two example model identifiers in the sample configuration are worth pausing on, because neither is an OpenAI model. That is the point of the design rather than an oversight: the interface is compatible so that a gateway, a relay or a self-hosted server can sit behind it, and the identifiers are passed to the provider unchanged.

There is one precedence rule. The first model in the list becomes the default unless a separate default setting is also provided, which means the list order is functional rather than cosmetic, and anyone reordering it changes what a new conversation starts with.

## One workspace per conversation, in a directory git ignores

The persistence model is simple and it is what separates this from a chat window.

Every conversation gets its own workspace directory, named by conversation identifier. Uploaded files, charts, reports, notebooks and other generated artefacts all live there and are all surfaced in a Files panel in the interface.

The feature list puts the consequence in plain words: variables and generated files stay available throughout the analysis. That is the difference between an agent that answers a question and an agent that leaves a state you can come back to, and it is why the artefact tracking and the exports are one design rather than two features.

Underneath, runtime data sits in a single directory containing the local SQLite database, an uploads folder and the workspaces folder. And those files are intentionally ignored by git, which is the right call for uploaded datasets and generated reports.

One consequence is worth naming. Because the database and the workspaces are untracked local files rather than a migration-managed store, copying a deployment means copying that directory, and there is no migration path described for moving conversations between machines.

## Chinese reports need fonts and a TeX engine installed on the host

Two of the advertised output formats depend on system packages rather than Python libraries, which is the most deployment-relevant detail in the readme.

PDF reports and Chinese report generation are described as optional, and the condition given for the Chinese case is specific: XeLaTeX and CJK fonts have to be installed. The feature list states the same requirement as a feature, including Chinese report generation when those are present.

There is a script that installs the system dependencies, and what it installs is named: both pdflatex and xelatex, the Chinese LaTeX packages, and the Noto CJK fonts. The comments in the environment example note that Chinese charts need the fonts too, so this is not only about text in a document.

The start script can install them for you when they are missing, but that path uses sudo with a package manager, and the readme is careful to mark it opt-in rather than automatic.

That is the right boundary. A tool that shells out to a system package manager during startup is a tool you cannot run unattended in a container without thinking about it, and the two-step, opt-in arrangement keeps that decision with the operator.

## Two ports, two overrides, two shell scripts

The quick start is unusually short, and the reason is that two shell scripts absorb the complexity.

The setup is: copy the backend's example environment file to its real name, edit that file to set the API key, the base URL and the model list, then run the start script. The interface is on port 3000. The stop script stops it.

Port conflicts are handled by environment variables passed to the start script, and both a backend and a frontend port can be overridden in one command.

The manual path underneath is the conventional one and is what the scripts automate:

```bash
cd backend
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env
uvicorn main:app --host 0.0.0.0 --port 8000 --reload
```

The frontend is then installed and started separately with npm.

The stated requirements are Python 3.11 or newer, Node.js 20 or newer and npm, with the TeX toolchain marked optional.

So the only non-optional chain is a Python runtime and a Node runtime. Everything else, including the model provider and the document toolchain, is configuration or opt-in installation.

## The paper is in a statistics journal, the tool is a web application

The repository is the official home of a paper, and the venue is worth noticing.

The citation block asks for a journal article in a statistical association's journal, with a volume, an issue number, a page range and a 2026 year, published by a commercial academic publisher. A second cited work is a survey on large language model-based agents for statistics and data science.

The repository description points at a university department page for the same project, and the header links both a project page and a cases page alongside the hosted application.

So the research output is aimed at statisticians, while the software is a general-purpose analysis agent: a web application with a React frontend, a FastAPI backend, a model interface, a file workspace and analysis tools, published to a hosted site and to npm-scale configuration.

The repository itself is small in structure. Seven top-level entries, of which two are the frontend and backend directories, one holds example data, one holds figures for the paper, one holds scripts, and the start and stop scripts sit at the root beside a gitignore and the readme.

There is no licence file in the tree and no licence field recorded, which for a research artefact intended for reuse is the gap worth raising with the authors.

## Conclusion

LAMBDA fits a researcher or analyst who wants an agent to actually run code over their data and leave the artefacts behind, rather than describe what it would do, since variables, charts and reports persist per conversation and everything exports to notebook, PDF or slides. It is a poor fit if you plan to expose it to people you do not trust, because the documentation is explicit that analysis code executes on the backend host with no sandboxing, or if you need Chinese PDF reports without touching the host first. Before you start it, check three things: which port the backend ends up on, since the manual command binds all interfaces while the start script is what the rest of the document assumes, what you have set as the model endpoint, since the shipped example lists models from a different provider, and whether you are on the current tag or the preserved earlier branch.

## FAQ

### What is the LAMBDA data analysis agent?

It turns natural-language questions into reproducible analysis workflows. You upload a dataset, ask a question, and it can inspect the data, write and run code, create visualizations, summarize findings and generate reports or notebooks from the session.

### How do I configure which models LAMBDA uses?

Through three values in the backend environment file: an API key, a base URL that should be any OpenAI-compatible API, and a model list that is a JSON array of identifiers. LAMBDA appends the chat completions path to that base URL and passes the identifiers through unchanged.

### Is it safe to run the LAMBDA backend on a shared network?

The readme says analysis code executes on the machine where the backend runs, and that you should not expose the backend to untrusted traffic without adding sandboxing or isolation. Note that the manual setup command binds the backend to all interfaces, so that binding is the first thing to change.

### What does LAMBDA need to generate Chinese PDF reports?

System packages rather than Python libraries: XeLaTeX and CJK fonts, along with pdflatex for plain PDF. A script installs the TeX engines, the Chinese LaTeX packages and the Noto CJK fonts, and the start script can do it when missing, but that path uses a package manager with elevated privileges and is opt-in.

### Where does LAMBDA keep uploads and generated files?

Each conversation has its own workspace directory named by conversation identifier, holding uploads, charts, reports and notebooks. All of it sits under a backend data directory that also contains the local SQLite database, and those files are intentionally ignored by git.

## Sources

- [AMA-CMFAI/LAMBDA on GitHub](https://github.com/AMA-CMFAI/LAMBDA)
- [Issues](https://github.com/AMA-CMFAI/LAMBDA/issues)
- [Project website](https://arxiv.org/abs/2407.17535)
- [README](https://github.com/AMA-CMFAI/LAMBDA/blob/main/README.md)
- [Releases](https://github.com/AMA-CMFAI/LAMBDA/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ama-cmfai-lambda
