Model or dataset
ConardLi/easy-dataset avatar
ConardLi/easy-dataset

Easy Dataset: A Tool for Building LLM Fine-Tuning and RAG Datasets from Documents

A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval

14,961 stars1,537 forksJavaScriptNOASSERTION

At a glance

What is it?
Easy Dataset is a desktop and server application that converts domain-specific documents into structured datasets for LLM fine-tuning, retrieval-augmented generation, and model evaluation. It supports PDF, Markdown, DOCX, TXT, and EPUB input, and exports in Alpaca, ShareGPT, and Multilingual-Thinking formats.
Who is it for?
Easy Dataset is a practical choice for teams or individuals who need to convert internal documents into structured training or evaluation data for LLMs. It handles the full pipeline: document ingestion, segmentation, question generation, answer generation, quality scoring, and export in standard fine-tuning formats.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 151 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Easy Dataset Does and Who Uses It

Easy Dataset addresses the practical bottleneck in LLM fine-tuning: converting raw domain documents into the question-answer pairs and dialogue examples that training pipelines consume. The README describes it as an application specifically designed for building LLM datasets, with support for model fine-tuning, retrieval-augmented generation (RAG), and model performance evaluation.

The application targets practitioners who have domain-specific PDF, Markdown, DOCX, TXT, or EPUB files and need to transform them into structured datasets without writing data-processing code. It is also relevant for teams building evaluation systems: version 1.7.0 added the ability to convert documents into evaluation datasets and run automated multi-dimensional evaluation tasks.

Document Processing and Question Generation Pipeline

Easy Dataset's core pipeline has four stages. First, it ingests documents using intelligent recognition that handles multiple file formats. Second, it splits the document text using one of several algorithms: Markdown structure-aware splitting, recursive separators, fixed-length chunking, or code-aware chunking. Third, it auto-extracts questions from each text segment using an LLM API, drawing on configurable question templates and supporting batch generation. Fourth, it generates answers using the same LLM API, with support for Chain of Thought (COT) generation and AI-based quality optimization.

A Domain Label Tree feature builds a global taxonomy from the document structure and automatically tags generated questions with relevant labels. This enables balanced export: when exporting the final dataset, the user can configure how many question-answer pairs to include per label, avoiding skewed datasets that over-represent one topic area.

Installing and Running Easy Dataset

Easy Dataset offers three installation paths: desktop client downloads for Windows, macOS (Intel and Apple Silicon), and Linux; npm-based local server; and Docker.

For the npm path, clone and start the server:

bash
git clone https://github.com/ConardLi/easy-dataset.git
cd easy-dataset
npm install

Then build and start:

bash
npm run build
npm run start

The application listens on port 1717. Open `http://localhost:1717` in a browser to access the interface.

For Docker, the official image is available at `ghcr.io/conardli/easy-dataset`. The `docker-compose.yml` provided in the repository mounts two local directories for persistence:

yml
services:
  easy-dataset:
    image: ghcr.io/conardli/easy-dataset
    container_name: easy-dataset
    ports:
      - '1717:1717'
    volumes:
      - ./local-db:/app/local-db
      - ./prisma:/app/prisma
    restart: unless-stopped

The `local-db/` volume holds the SQLite database managed by Prisma. The `prisma/` volume holds the schema and migration files. The Dockerfile uses a multi-stage build: an Alpine-based builder stage compiles the Next.js application with pnpm and a separate runner stage installs only production dependencies.

Dataset Types and Export Formats

Easy Dataset generates four dataset types. Single-turn QA datasets produce standard question-answer pairs for basic fine-tuning. Multi-turn dialogue datasets support customizable roles and scenarios. Image QA datasets generate visual question-answer data from images, accepting input via directory, PDF, or ZIP. Data Distillation generates questions directly from domain topic descriptions without requiring document uploads.

Export formats include Alpaca, ShareGPT, and Multilingual-Thinking, in both JSON and JSONL file types. The LLaMA Factory integration generates a one-click configuration file for the LLaMA Factory training framework. Direct upload to Hugging Face Hub is also supported.

The evaluation system, introduced in v1.7.0, can generate true/false, single-choice, multiple-choice, short-answer, and open-ended evaluation questions. A Judge Model automatically evaluates model answer quality against customizable scoring rules. A Human Blind Test (Arena) system supports double-blind comparison of two models' answers.

Model Compatibility and API Requirements

Easy Dataset calls external LLM APIs for question generation, answer generation, data cleaning, and evaluation. The README states it is compatible with all LLM APIs that follow the OpenAI format. Explicitly named providers include OpenAI, MiniMax, Ollama (for local models), Zhipu AI, Alibaba Bailian, and OpenRouter.

For PDF parsing and image QA, the application uses vision-capable models. The README names Gemini and Claude as supported providers for this purpose. The Model Testing Playground allows simultaneous comparison of up to three models.

All LLM API credentials and provider endpoints are configured within the application. The application does not bundle API keys; users supply their own. Token consumption statistics and API call tracking are available through the Resource Monitoring Dashboard.

Limitations and Cases Where It Falls Short

Easy Dataset's last push was on 2026-05-01, which is about five months before the current date. The most recent release is version 1.7.3, dated 2026-04-09. The repository has seen no new releases since then. Teams evaluating it for a production pipeline should check whether open issues blocking their use case have been addressed.

The application is document-centric: it converts text from uploaded files into training data. It is not designed for generating synthetic datasets from scratch without a source document, though the Data Distillation feature allows generating questions from topic descriptions. Teams working primarily with structured tabular data (CSV, SQL exports, time series) will find no ingestion path for those formats in the documented features.

Custom prompt templates are project-level only. There is no account-level template library, so teams running multiple projects must configure templates in each one separately. The README does not document rollback procedures if a generated dataset contains widespread errors introduced by a misconfigured prompt template. The background task management center allows monitoring and interrupting batch generation tasks, which is the primary recovery mechanism for jobs that go wrong mid-run.

Architecture, License, and the Research Paper

Easy Dataset is built as a Next.js application (version 1.7.3, as seen in `package.json`). The database layer uses Prisma with an SQLite backend stored in `local-db/`. The electron package wraps the Next.js server as a desktop application using Electron. The build process uses pnpm and generates platform-specific binaries for Windows, macOS, and Linux via `electron-builder`.

The repository links to a research paper at `https://arxiv.org/abs/2507.04009v1`. The project is licensed with a custom license (shown as `NOASSERTION` in GitHub metadata); the actual terms are in the `LICENSE` file. The desktop client binaries are available at the GitHub Releases page.

An alternative to Easy Dataset for dataset construction is LabelStudio (github.com/heartexlabs/label-studio), an open-source data labeling platform. LabelStudio provides manual annotation workflows for text, images, and other modalities, while Easy Dataset automates the generation step using LLM APIs. The choice depends on whether the team needs AI-assisted generation or human-verified labeling.

Editorial conclusion

Easy Dataset is a practical choice for teams or individuals who need to convert internal documents into structured training or evaluation data for LLMs. It handles the full pipeline: document ingestion, segmentation, question generation, answer generation, quality scoring, and export in standard fine-tuning formats. The last push to the repository was on 2026-05-01. Teams that need a document-to-dataset pipeline compatible with LLaMA Factory or Hugging Face Hub, and who want a UI-driven workflow rather than a Python scripting approach, will find Easy Dataset covers that ground directly. Teams who need a fully scriptable pipeline, or who are working primarily with structured tabular data rather than text documents, should evaluate whether the document-centric workflow fits their use case before adopting it.

Frequently asked questions

How do I use Easy Dataset?

Install Easy Dataset by cloning the repository and running `npm install`, `npm run build`, and `npm run start`, then open `http://localhost:1717`. Create a project, upload your documents, configure the LLM API provider, and run the question and answer generation pipeline to build your dataset.

What is a good alternative to Easy Dataset?

LabelStudio is an open-source alternative for dataset creation that focuses on manual annotation rather than AI-assisted generation. For teams that want an automated approach similar to Easy Dataset but fully code-driven, the README mentions compatibility with LLaMA Factory and Hugging Face Hub for downstream use.

What export formats does Easy Dataset support?

Easy Dataset exports datasets in Alpaca, ShareGPT, and Multilingual-Thinking formats as JSON or JSONL files. It also generates one-click LLaMA Factory configuration files and supports direct upload to Hugging Face Hub.

Official sources

  1. ConardLi/easy-dataset on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/conardli-easy-dataset.svg)](https://hysenlabs.com/projects/conardli-easy-dataset)
Community notes

Community notes