Data Formulator: a Microsoft Research workspace where AI agents build charts from your data
🪄 Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data.
At a glance
- What is it?
- Data Formulator is an MIT-licensed Python and TypeScript research prototype that pairs natural language with a visual Data Thread so analysts can branch through data questions. It installs from PyPI or runs in Docker on port 5567, and it depends on an external LLM provider for the agent steps.
- Who is it for?
- Adopt Data Formulator if you already have an LLM API key, you work with pandas-scale or DuckDB-scale tables, and you want to branch through questions instead of scrolling a chat log. Skip it if you need a governed BI deployment with row-level security, or if you cannot send data to an external model provider.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Data Formulator aims at: chat history is a bad place to keep an analysis
The README states the motivation in two lines: data lives in many places, and questions evolve as you explore. The second half is the interesting one. A conversational analysis tool records every question in a single chronological thread, so by the twentieth turn you have lost the shape of the investigation. You cannot see which branch produced the chart you are looking at, and you cannot go back and try a different aggregation without mentally rewinding.
Data Formulator's answer is the Data Thread, described in the README as a way to "branch into different questions, compare paths, and use visualizations to discover deeper insights without losing context." That is a structural claim, not a chat feature. The intended user is an analyst who already knows what a grouped bar chart should look like and wants the agent to handle the transformation and encoding work, not a person looking for an auto-generated dashboard.
The second piece is the data connector layer. The README says connectors "give agents a common way to connect to different data sources and maintains a data memory to remember the relationships between them." The stated failure mode is agents answering before relationships between sources are clear, which is a real problem when you join a CSV to a warehouse table and the agent guesses the key.
How the agent, the DuckDB layer and the Flint chart engine fit together
The repository layout shows a split application. The Python side lives in py-src/ and pyproject.toml, and it pulls in flask, duckdb, pandas, pyarrow, scikit-learn, and a long list of database drivers. The browser side is a Vite and React application (vite.config.ts, src/, package.json) with MUI, CodeMirror, and chart.js in its dependency list. The Python process serves the UI and runs the data work; the front end renders it.
DuckDB is the execution engine for local data. The changelog entry for v0.2 describes "Large data support with DuckDB integration," so the intent is that tables larger than comfortable pandas frames get pushed into DuckDB rather than held in memory. pyarrow>=13.0.0 appears in both requirements.txt and pyproject.toml, which fits a columnar handoff between the two.
Chart rendering is delegated. The README points at Flint, an open-source visualization language that "compiles compact chart specs into polished visualizations," and package.json lists flint-chart with a minimum version of 0.5.0. The v0.7 release notes describe a "semantic chart engine powered by Flint" with more than 30 chart types and a style-refinement agent. So the agent's job is to emit a chart specification, and Flint turns that specification into the rendered graphic.
Model access is abstracted through LiteLLM. The README lists OpenAI, Azure, Ollama, and Anthropic as supported providers via LiteLLM, and litellm is pinned in both dependency files to >=1.84.0,<1.92. The comment in requirements.txt explains the ceiling: litellm 1.92 and later ship manylinux-only wheels, so installs hang on Windows and macOS without a Rust toolchain. That pin is a deliberate compatibility decision, and it means you will not get whatever the newest LiteLLM release adds.
Installing Data Formulator and producing your first chart
The README offers several routes. The fastest is uvx, which runs the package without a permanent install. The stable release is 0.7, and the 0.8 beta is available behind a pre-release flag.
uvx data_formulatorIf you prefer pip, the README gives the equivalent commands. The 0.8 beta is opt-in, and the version string in the README is 0.8.0b1, which matches the 0.8.0b1 version in pyproject.toml.
pip install data_formulator
pip install --pre data_formulator==0.8.0b1For a container deployment, docker-compose.yml is explicit about the sequence: copy the environment template, fill in API keys, build, and open port 5567. The compose file mounts a named volume at /home/appuser/.data_formulator so uploaded files and sessions survive a container restart.
cp .env.template .env
docker compose up --build
# then open http://localhost:5567Once the interface is up, the workflow the README describes is: load data, ask questions, review results, and branch in the Data Thread. Practically, that means pointing a connector at a local file or a supported database, typing a question in natural language, and letting the agent produce a chart specification that you then refine. The model credentials come from your .env file, so the first thing to check if a request fails is whether the provider key and model name in that file are correct.
Where Data Formulator breaks down
The dependency list is the clearest limitation. Both requirements.txt and pyproject.toml include drivers for MongoDB, Cosmos DB, Azure Blob, Kusto, MySQL, ClickHouse, MSSQL, PostgreSQL, S3, Athena, BigQuery, and Databricks SQL, plus Azure Key Vault and azure-identity. That is a large surface installed by default. The requirements.txt comment notes that all data-loader dependencies are included and that try/except at import time keeps things safe, which tells you the project chose breadth over a slim core. If you want a minimal install, this is not it.
The second constraint is that the agent needs a model endpoint. There is no bundled local model in the dependency list; Ollama appears as a supported provider through LiteLLM, which means you supply the server. If your data cannot leave your network and you have no self-hosted model, the agent features are unavailable to you.
Third, the project describes itself in pyproject.toml as a "research prototype" with the classifier Development Status :: 4 - Beta. The 0.8 line is still in beta as of the 0.8.0b1 release on 2026-08-15, so the version you install from the plain pip command is 0.7, not the version whose features the news section leads with. Anyone reading the 0.8 bullet list and running pip install data_formulator will get different behaviour.
Finally, desktop builds are ephemeral. The README states that CI builds Windows and macOS applications for pull requests and main-branch updates, that artifacts are found under the Artifacts section of the desktop-build workflow, and that workflow artifacts are retained for 30 days. If you want a desktop binary rather than a server, you may find the most recent artifact already gone.
Data Formulator versus Power BI and notebook-based charting
The comparison people search for is Data Formulator versus Power BI, and the difference is architectural rather than cosmetic. Power BI is a governed semantic-layer product: you model data once, publish it, and consumers build reports against that model with row-level security and scheduled refresh. Data Formulator has no semantic layer in the repository's own description. It has connectors and a data memory that tracks relationships between sources, and the analysis lives in a Data Thread you branch through. It is a workspace for the person doing the exploring, not a publishing platform for an organization.
Against notebook charting, the difference runs the other way. In a notebook you write the transformation, and the chart library renders it. Here the agent proposes the transformation and the chart specification, and Flint renders it. That is faster when you are exploring and worse when you need an auditable, version-controlled pipeline, because the interesting logic lives in a model call rather than in a cell you can diff.
The honest framing is that Data Formulator sits between the two. It gives you more structure than a chat window and less governance than a BI platform.
Licence, maintenance and what an upgrade costs you
The licence is MIT, declared in pyproject.toml and shown in the README badge. MIT is permissive, so redistributing or modifying the code carries few obligations beyond keeping the notice. That said, the dependencies are not all MIT, and the data loaders pull in vendor SDKs from Microsoft, Google, Amazon, and Databricks. If you ship a bundled build, check those licences separately; this is not legal advice, just a pointer to where the work is.
The repository is not archived and the last push was on 2026-09-20. Releases are frequent: 0.7-alpha.2 on 2026-05-13, 0.7.0 on 2026-05-28, and 0.8b1 on 2026-08-15. That cadence is a cost as well as a signal. The litellm pin at <1.92 is the kind of constraint that has to be revisited whenever the upstream wheel situation changes, and the comment in requirements.txt suggests the maintainers are tracking it deliberately.
Upgrading from 0.7 to 0.8 is not a drop-in if you are on the default pip command, because 0.8 is still published as a pre-release. You either opt in with --pre and accept beta behaviour, or you stay on 0.7 and miss the unified flow, the extra data sources, and the chart improvements listed in the news section.
Editorial conclusion
Adopt Data Formulator if you already have an LLM API key, you work with pandas-scale or DuckDB-scale tables, and you want to branch through questions instead of scrolling a chat log. Skip it if you need a governed BI deployment with row-level security, or if you cannot send data to an external model provider. Before committing, verify which LiteLLM model string your provider accepts, check that the 0.8 beta is the version you actually want, and confirm the desktop artifact from the workflow has not expired.
Frequently asked questions
What is Microsoft Data Formulator?
It is a Microsoft Research project for data exploration with visualizations powered by AI agents, described in pyproject.toml as a research prototype. It combines UI interactions with natural language so analysts can branch into alternative analyses and share results.
How do I use Data Formulator?
The README describes one flow: load data through a connector, ask questions in natural language, review the results, and branch in the Data Thread. Charts are rendered by Flint from compact chart specifications the agent produces.
What is Data Formulator in Excel?
Data Formulator is not an Excel feature. It is a standalone Python package and web application from Microsoft Research that connects to files, databases, and platforms, and it renders charts through the Flint visualization language.
What is data formulator?
It is an interactive AI-powered data analysis system that connects to data sources and visualizes them, released under the MIT licence with a Python package on PyPI and a browser interface served on port 5567.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-data-formulator)