Hugging Face AI Sheets: No-Code Dataset Building with Open Models
Build, enrich, and transform datasets using AI models with no code
At a glance
- What is it?
- AI Sheets is an Apache-2.0 TypeScript tool from Hugging Face for building, enriching, and transforming datasets using AI models without writing code. It deploys locally or on the Hub, supports thousands of open models through Inference Providers, and accepts custom LLM endpoints that speak the OpenAI API spec.
- Who is it for?
- AI Sheets is well-suited for data scientists and ML engineers who need to enrich or transform datasets using open models without scripting inference calls. Teams that need strict data governance, audit trails, or complex multi-step transformation pipelines tied to a version-controlled workflow should compare it against tools like dbt for structured transformation or purpose-built annotation platforms.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AI Sheets is and the problem it addresses
AI Sheets addresses the gap between having a dataset and having a useful one. Enriching a raw dataset, adding classifications, generating descriptions, translating text, or inferring labels from content, traditionally requires writing inference code, managing API calls, and handling rate limits and errors. AI Sheets exposes those operations as spreadsheet-style column transformations driven by natural language prompts.
The README describes the tool as allowing users to build, enrich, and transform datasets using AI models with no code. The intended audience is data scientists, ML engineers, and researchers who work with datasets on or for the Hugging Face Hub and want to apply model inference to their data without managing the infrastructure themselves.
AI Sheets is also the front end for Hugging Face Jobs-based bulk generation. The README shows how to launch data generation pipelines that extend an existing dataset using HF Jobs, which runs the processing remotely on Hugging Face infrastructure rather than locally.
Deployment options: Hub Space, Docker, and pnpm
Three deployment paths are documented. The fastest is the public Space at huggingface.co/spaces/aisheets/sheets, which requires no local setup.
For local deployment, Docker is the simplest option:
export HF_TOKEN=your_token_here
docker run -p 3000:3000 \
-e HF_TOKEN=HF_TOKEN \
aisheets/sheetsOpen http://localhost:3000 in your browser after the container starts. The Dockerfile is a two-stage build using node:22-slim. The build stage installs pnpm, runs pnpm install --frozen-lockfile, and then pnpm build. The production stage copies the built server and dist directories and exposes port 3000.
For development, the pnpm path gives a faster feedback cycle:
git clone https://github.com/huggingface/aisheets.git
cd sheets
export HF_TOKEN=your_token_here
pnpm install --frozen-lockfile
pnpm devThe dev server opens at http://localhost:5173. To build for production and serve with the built-in Express server:
pnpm build
export HF_TOKEN=your_token_here
pnpm serveThe package.json engines field lists the supported Node versions as ^18.17.0, ^20.3.0, or >=21.0.0.
Model access: Inference Providers, custom endpoints, and Ollama
By default, AI Sheets uses the Hugging Face Inference Providers API to access open models. The HF_TOKEN variable must carry a token with the necessary scopes for authenticated inference.
For custom or local LLMs, two environment variables redirect inference calls. MODEL_ENDPOINT_URL points to the base URL of any OpenAI-compatible API endpoint, and MODEL_ENDPOINT_NAME specifies the model identifier. The README gives this example for Ollama running locally:
export MODEL_ENDPOINT_URL=http://localhost:11434
export MODEL_ENDPOINT_NAME=llama3To use this with Ollama, first start the Ollama server and pull the model:
export OLLAMA_NOHISTORY=1
ollama serveThen run the model:
ollama run llama3Then set the environment variables and start AI Sheets with pnpm serve. The README notes that the text-to-image generation feature cannot use a custom endpoint and will always call the Hugging Face Inference Providers API regardless of the MODEL_ENDPOINT_URL setting.
All models available through the Hugging Face Inference Providers API remain selectable from the column settings even when a custom endpoint is configured; the custom endpoint is used for text tasks, not as a replacement for the entire model catalogue.
Running large-scale pipelines with HF Jobs
For datasets too large to process locally, AI Sheets provides scripts and configuration for HF Jobs. The README shows two variants. The first uses the standard Inference Client:
hf jobs uv run \
-s HF_TOKEN=$HF_TOKEN \
https://github.com/huggingface/aisheets/raw/refs/heads/main/scripts/extend_dataset/with_inference_client.py \
nvidia/Nemotron-Personas dvilasuero/nemotron-kimi-qa-distilled \
--config https://huggingface.co/datasets/dvilasuero/nemotron-personas-kimi-questions/raw/main/config.yml \
--num-rows 100The second uses vllm inference, which the README says saves on inference costs but requires setting a vllm-compatible flavor when launching the job:
hf jobs uv run --flavor l4x1 \
-s HF_TOKEN=$HF_TOKEN \
https://github.com/huggingface/aisheets/raw/refs/heads/main/scripts/extend_dataset/with_vllm.py \
nvidia/Nemotron-Personas dvilasuero/nemotron-kimi-qa-distilled \
--config https://huggingface.co/datasets/dvilasuero/nemotron-personas-kimi-questions/raw/main/config.yml \
--num-rows 100 \
--vllm-model deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5BThese jobs run the pipeline on the specified source and destination datasets using a config YAML that defines the prompts and transformation logic. The --num-rows flag limits the run to a sample; omitting it runs the full dataset.
Authentication options and the OAuth path
Three authentication modes are available. The HF_TOKEN environment variable performs simple token-based auth and is used for both authenticated inference and user identity. The OAUTH_CLIENT_ID variable enables Hugging Face OAuth2 login for a deployed instance where multiple users need their own accounts. The OAUTH_SCOPES variable controls which permissions the OAuth flow requests; the default is openid profile inference-api manage-repos.
For a local single-user deployment, HF_TOKEN alone is sufficient. For a shared deployment where users should log in with their own Hugging Face accounts, OAUTH_CLIENT_ID must be set; the README links to a setup guide for configuring the OAuth application on the Hugging Face platform.
Because AI Sheets runs inference against Hugging Face-hosted models using the requesting user's token, the token scope must include inference-api. A read-only token will authenticate the user but fail when the model call is made.
Limitations and the wrong-tool cases
AI Sheets is built on the Qwik framework (using @builder.io/qwik and @builder.io/qwik-city from package.json) and uses SQLite via the build Docker image for local persistence. The application is designed as a single-user or small-team interface; it is not a batch processing engine for production ETL workloads.
The text-to-image feature cannot be redirected to a custom model endpoint. This is a documented limitation: image generation always uses Hugging Face Inference Providers regardless of the MODEL_ENDPOINT_URL configuration. Teams that need fully air-gapped or on-premises image generation will find this blocking.
For complex multi-step data transformation, tools like dbt handle SQL-based structured transformations with version control, dependency graphs, and test assertions, which AI Sheets does not provide. AI Sheets is designed for AI model inference on individual rows or columns, not for relational joins, aggregations, or schema migrations.
The repository has no GitHub releases and the last push was on 2026-09-23.
Editorial conclusion
AI Sheets is well-suited for data scientists and ML engineers who need to enrich or transform datasets using open models without scripting inference calls. Teams that need strict data governance, audit trails, or complex multi-step transformation pipelines tied to a version-controlled workflow should compare it against tools like dbt for structured transformation or purpose-built annotation platforms. Before deploying, set HF_TOKEN with a token that has the inference-api scope, and decide whether to use Inference Providers for access to hosted models or to point MODEL_ENDPOINT_URL at a local Ollama instance or private vllm server for data that should not leave the machine.
Frequently asked questions
How to use AI Sheets?
AI Sheets can be used immediately at huggingface.co/spaces/aisheets/sheets with no setup. For local use, run `docker run -p 3000:3000 -e HF_TOKEN=your_token_here aisheets/sheets` and open http://localhost:3000. A Hugging Face token with the inference-api scope is required for model access.
What is AI Sheets?
AI Sheets is an open-source tool from Hugging Face for building, enriching, and transforming datasets using AI models with no code. It provides a spreadsheet-style interface where columns can be filled or transformed using model inference from thousands of open models on the Hugging Face Hub.
Can AI Sheets run models locally without sending data to Hugging Face?
Yes for text tasks. Set MODEL_ENDPOINT_URL to point to a local server such as Ollama or a private vllm instance, and set MODEL_ENDPOINT_NAME to the model identifier. The README notes that text-to-image generation is the one exception: it always calls the Hugging Face Inference Providers API.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-aisheets)