Model or dataset
antvis/AVA avatar
antvis/AVA

antvis/AVA: natural language queries that become executed code or SQL

🤖 AI-native Visual Analytics framework build for agents.

1,571 stars159 forksTypeScriptMIT

At a glance

What is it?
An analytics framework where a question in plain English is answered by loading data, inferring its shape, deciding between a JavaScript helper and a SQL engine, and then executing whatever the model produced. The default branch is named ai, the only release is an alpha, and the trust boundary is the model provider.
Who is it for?
Adopt AVA if you are building a product surface where a user asks a question of their data and expects an answer with a chart, because the load, suggest, analyse and visualise path is short and the result object is well shaped for that.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The model writes code and your process runs it

The architecture diagram in the README is the most important thing in this repository, and it has one line that deserves attention. The Analysis Module is described as generating and executing code or SQL, and the result object distinguishes the two: code for the in-memory JavaScript path, sql for the SQLite path. So the mechanism is not that the model returns an answer you display. It is that the model produces an artefact, and the framework executes it. Everything else in the design is built around that. Data is loaded, metadata is extracted with type inference and statistics, a size check picks an engine, the model writes against that engine, the output is summarised in natural language, and a chart is optionally generated. The consequence is a trust boundary you should draw explicitly before you ship anything. A natural language query becomes a JavaScript string that runs in your Node process or your browser tab. There is no mention in the documentation of a sandbox, an allowlist, a permission prompt, or a statement about which operations the generated code is permitted to perform. If your data is untrusted, the honest reading is that the model provider is inside your trust boundary, and so is anything that managed to influence the prompt. That is a decision for you to make, not a flaw in the framework, but the framework does not make it for you either.

One npm install, an LLM endpoint you supply, and a threshold you tune

Installation is a single package published as @antv/ava, and the quick start offers the same name for three package managers:

bash
npm install @antv/ava

The manifest declares Node 18 or later as the engine requirement, sets sideEffects to false so a bundler can drop unused imports, and compiles its output twice with tsc, once to CommonJS and once to ESM. What you must bring is the model access. The constructor takes an llm object with a model, an apiKey and a baseURL, and the example populates it with a model name and a placeholder key and URL rather than a vendor default, so the framework is not tied to one provider as long as the endpoint speaks a compatible protocol. Two runtime dependencies reveal how much of the work is yours: better-sqlite3 for the Node path, which is a native module and therefore a compilation step in your build, and idb for the browser path, which is an IndexedDB wrapper. The other three are ordinary. csv-parse does the CSV reading, zod validates what the model returns, and the ai package with its OpenAI provider is the layer the request goes through. If you deploy to an edge runtime rather than Node, that better-sqlite3 dependency is the first thing that will need checking against your platform.

Ten kilobytes decides your engine and your result shape

The data path branches on size, and the branch is configurable. The default for sqlThreshold is 10KB, and anything under that is analysed with JavaScript helpers in memory; at or above it, the framework checks the environment and picks IndexedDB in the browser or SQLite in Node. The quick start overrides the default with a two megabyte threshold:

```typescript const ava = new AVA({ llm: { model: 'ling-1t', apiKey: 'YOUR_API_KEY', baseURL: 'LLM_BASE_URL', }, sqlThreshold: 1024 * 1024 * 2, // Threshold for switching to SQLite });

So the same question against the same shape of data can execute on a different engine depending on how much you loaded, and the return value changes with it, since one path fills result.code and the other fills result.sql. That is a design decision with real consequences for a caller. Your error handling has to cover both engines, your logging has to record which one ran, and a query that worked in development can take a different path in production simply because the dataset grew. Raising the threshold is a performance choice with a correctness edge attached to it, and the documentation does not say whether the JavaScript path can express everything the SQL path can. If you are integrating this, treat the threshold as a setting with tests on both sides of it, not as a performance knob. Two supporting details from the manifest: the browser path uses the idb package for IndexedDB, and the Node path uses better-sqlite3, which is a native module and therefore a compilation step in your build.

Four ways to load data, all of them untrusted

The load methods are the front door, and there are four of them, each with a different risk profile. loadCSV takes a file path in Node or a content string in the browser, and the quick start shows the browser variant reading a file input and passing the text straight through. loadObject takes a plain array of records. loadURL fetches a URL and takes a transform function that maps the response, so you are pointing it at a URL you do not control and giving it code to run on the result. loadText takes unstructured text and has the LLM extract structure from it, which is the case where a model is parsing untrusted content into data. Every one of these ends up in the same place: the data becomes context for a model that then writes code. That is the design, and for a trusted internal dataset it is a reasonable one. For a product where a user uploads a spreadsheet, the uploaded file is now a prompt injection surface, and the mitigation is not in the framework. Put a boundary between the data and the execution, and assume the model will eventually be talked into writing something outside your intent. The one piece of the API that leans the other way is suggest, which returns a list of recommended queries each with a score and a reason, so a UI can offer the user something to click rather than accepting free text directly into the execution path.

visualize can return null, and hands back a syntax and an HTML string

The chart step is the one to design around carefully. The signature returns an object with chartType, syntax and html, or null, and the null is documented for two distinct reasons: no visualisation intent in the answer, or no usable data. So a successful analysis can legitimately produce no chart, and any integration that assumes otherwise will break on ordinary questions. The successful shape is itself interesting. The chartType is a plain string such as column, the syntax is described as GPT-Vis chart syntax, and html is standalone HTML that renders the chart. That gives you two integration routes and the documentation does not choose between them. You can render the chart yourself from the syntax, which means adopting whatever renderer speaks that language, or you can drop the html into a frame, which means shipping model-generated markup to a browser. The first route gives you control and more work. The second is a few lines of code and puts generated markup on a page that also holds your application. Neither is wrong, and the difference is a security and styling decision rather than a technical one, so make it deliberately. The examples directory has a visualization example and a suggest example, and those two files are where the intended usage is.

An API shaped for a product, with a branch named ai

Two repository facts belong in any evaluation. First, the default branch is named ai, and the only release is 4.0.0-alpha.1, published on 2026-06-09, while the last push was 2026-09-20. So the version on npm is an alpha roughly three months behind the branch you would read on GitHub, and the 4 major number signals a rewrite of whatever came before. That is an honest pre-1.0 position for a new architecture, and it means you should expect the API to change. Second, the repository contains an llms.txt file at the root and carries an AI Agent badge alongside the website and documentation links. That is a deliberate choice to make the project legible to a model rather than only to a person, and it is consistent with a framework whose subject is model interaction. It also has a practical effect on how you should read this repository: the README is a summary aimed at both audiences, which is why the architecture diagram is the densest part of it. The topics on the repository, including auto-insight, chart recommendation, insight-gpt and narrative charts, describe a product direction rather than a library category, and the suggest method returning a reason string is the clearest expression of that intent.

Dual tsc output, Node 18, and a linter two majors behind the test runner

The build is deliberately plain. The manifest compiles the source twice with tsc, once with the commonjs module setting into lib and once with es2020 into esm, and ships both directories. Main points at the CommonJS build, module and the types field at the ESM side, engines requires Node 18 or later, and sideEffects is false so a bundler can drop unused imports. Tests are vitest, with watch, ui and coverage variants, and there is a top-level __tests__ directory. The examples directory is the other half of the verification story: eight scripts named after the paths they exercise, covering companies, heart disease and loan payments as datasets, plus object, text, URL and CSV loading, suggest, and visualisation. That is a sensible matrix for a framework whose risk surface is the variety of inputs, and it doubles as usage documentation. One inconsistency is worth naming. The test tooling is current, with vitest at 4.0.18, while the lint tooling sits at eslint 7.32.0 with typescript-eslint 4.29.2, which is old enough that it predates the flat config style the repository does not use. For a project that calls a language model, the missing piece in the visible file list is a recorded-response fixture, since a framework whose output is model text cannot be tested deterministically without one.

Editorial conclusion

Adopt AVA if you are building a product surface where a user asks a question of their data and expects an answer with a chart, because the load, suggest, analyse and visualise path is short and the result object is well shaped for that. Do not adopt it in a pipeline that handles untrusted input without a boundary of your own, since the framework generates JavaScript and executes it in your process, and the four load methods, including a browser file input and a fetched URL, are exactly the paths where data can carry instructions. Two things to check before you build on it. The published artifact is 4.0.0-alpha.1 from 2026-06-09 while the last push was 2026-09-20, so npm is three months behind the default branch, which is named ai. And decide your execution model early, because sqlThreshold defaults to 10KB and crossing it changes both the engine and the result shape from code to sql.

Frequently asked questions

How do I install and initialise AVA?

Install with npm install @antv/ava, then construct an AVA instance passing an llm object with model, apiKey and baseURL. You can also pass sqlThreshold, documented as the size in bytes at which analysis switches from in-memory JavaScript to SQLite or IndexedDB, with a default of 10KB.

How does AVA decide between JavaScript analysis and SQL?

It checks the data size against sqlThreshold, which defaults to 10KB. Under the threshold it uses in-memory JavaScript helpers, and at or above it checks the environment and uses IndexedDB in the browser or SQLite in Node, returning sql instead of code in the result.

What does AVA.visualize return?

An object with chartType, syntax and html, or null. The null case applies when there is no visualisation intent or no usable data. The syntax field is described as GPT-Vis chart syntax and html is standalone HTML that renders the chart.

Which data sources can AVA load?

Four: loadCSV, which takes a file path in Node or a content string in the browser, loadObject for a plain array, loadURL with an optional transform function, and loadText, which has the LLM extract structure from unstructured text.

What version of AVA is published?

4.0.0-alpha.1, published on 2026-06-09, and it is the only release. The last push to the repository was 2026-09-20 and the default branch is named ai, so the published artifact is behind the branch.

Official sources

  1. antvis/AVA on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/antvis-ava.svg)](https://hysenlabs.com/projects/antvis-ava)