# ai-data-extraction reads nine assistants' history out of their own storage

> ai-data-extraction is a dependency-free Python toolkit that walks the on-disk storage of nine AI coding assistants, from Claude Code's JSONL sessions to Cursor's content-addressed blob store, and writes every conversation into one JSONL file. It exists because each vendor invented its own layout, and the scripts are versioned against those layouts rather than against a stable API.

**0xSero/ai-data-extraction** — extract all your personal data history from cursor, codex, claude-code, windsurf, and trae

- Repository: https://github.com/0xSero/ai-data-extraction
- Stars: 1,287 · Forks: 110
- Language: Python
- License: not declared
- Published: 2026-09-29 · Updated: 2026-09-29 · Language: en
- Canonical page: https://hysenlabs.com/projects/0xsero-ai-data-extraction

## Nine scripts, because nine vendors invented nine layouts

The problem is not extraction, it is that nobody agrees where the data lives. Claude Code keeps JSONL session files under ~/.claude and a handful of sibling directories. Codex writes rollout JSONL under ~/.codex. Continue keeps JSON sessions in ~/.continue/sessions/. Gemini CLI nests its chats under a hashed path, ~/.gemini/tmp/[hash]/chats/. Cursor uses SQLite, with a second extractor needed for the terminal client. The toolkit's answer is one script per vendor, each searching that vendor's directories, reading the format it actually writes, and emitting the same shape. Running them together gives you one conversation corpus across products that share no format and no export button. The audience is narrow and specific: people assembling machine learning data from their own usage, and people who want their own history out of a tool that has no history page.

## Cursor needs two readers, and one walks a blob tree

Cursor gets more attention than any other target because its storage has changed shape several times. extract_cursor.py covers four generations at once: the old Chat mode in workspace storage, Composer inline storage where messages sit in a composerData array, the Composer separate storage used across the v1.x to v2.0 transition where messages move into bubbleId keys, and the latest Composer and Agent layout in v2.0 and later. It reads state.vscdb and cursorDiskKV. The terminal client is a different storage system again, so extract_cursor_cli.py reads ~/.cursor/chats/<chatId>/<agentId>/store.db, which the README describes as a content-addressed SQLite blob store: a blobs table mapping a sha256 hex id to data, where message blobs are JSON with role, content and id, and tree nodes are protobuf-framed lists of 32-byte child blob references. Message order is reconstructed by walking those references depth first from the latestRootBlobId recorded in the meta table. That is the most intricate code in the repository, and it exists because the GUI and the CLI do not share a database.

## OpenCode is SQLite, JSON and Tauri .dat in one script

extract_opencode.py is the other script with real work in it, because OpenCode ships a CLI and a desktop app with different storage again. It searches the CLI locations on Linux and macOS, the desktop locations under ai.opencode.app, and it handles three formats: the opencode.db SQLite database, JSON session, message and part files, and Tauri .dat files from the desktop build. It also carries the older layouts, reading storage/message, storage/part and storage/session for installs that predate the database, and it pulls sidecar metadata from storage/session_diff, storage/directory-readme, storage/agent-usage-reminder and storage/rules-injector. What lands in the output is a full conversation hierarchy with tool calls, code blocks, token usage, cost tracking, model and provider, agent mode, project directory, and parent and child session links. The practical point for an adopter is that this script is the one most likely to need patching, since it is written against a storage layout that the upstream project is free to change.

## The simpler five map straight onto their file format

The rest are more direct, and worth listing because the search paths are the whole story. extract_trae.py looks in ~/.trae and ~/Library/Application Support/Trae, reading JSONL and SQLite. extract_windsurf.py reads SQLite in a VSCode-like format from ~/Library/Application Support/Windsurf or the equivalent per platform. extract_continue.py reads the JSON session files in ~/.continue/sessions/ and pulls user and assistant messages, tool calls and results, reasoning blocks, context items and workspace information. Claude Code's own script searches five directories, ~/.claude, ~/.claude-code, ~/.claude-local, ~/.claude-m2 and ~/.claude-zai, which suggests the author was working around forks and mirrors as much as the official client. The common pattern across all of them is auto-discovery: detect the operating system, walk the known base directories, find every installation of the target tool, then scan for SQLite files with .vscdb or .db extensions and for JSONL sessions.

## Running it is one command per vendor, and no pip install

There is nothing to install, which is the strongest practical property of the project. The only setup step the documentation gives is a version check:

```bash
python3 --version  # Ensure Python 3.6+ is installed
```

Then you run the script for the assistant you care about, with no arguments:

```bash
python3 extract_claude_code.py
```

The same shape applies to extract_cursor.py, extract_codex.py, extract_trae.py, extract_windsurf.py, extract_continue.py, extract_gemini.py, extract_opencode.py and extract_cursor_cli.py. When you want everything, there is a shell driver at the root:

```bash
./extract_all.sh
```

There is no output flag, no path argument and no filter in anything documented, which means the scripts write where they are run from and take what they find. Run it in a directory you are willing to fill with your own conversation history.

## One JSON object per conversation, with code_context attached

Every extractor converges on the same record, and the shape is close to what Cursor stores, with the code context carried as structured data rather than as text. A message can carry a code_context array whose entries name the file, the code and a range:

```json
{
  "messages": [
    {
      "role": "user",
      "content": "How do I fix this TypeScript error?",
      "code_context": [
        {
          "file": "/Users/user/project/src/index.ts",
          "code": "const x: string = 123;",
          "range": {
            "selectionStartLineNumber": 10,
            "positionLineNumber": 10
          }
        }
      ],
      "timestamp": "2025-01-16T14:30:22.123Z"
    }
  ]
}
```

The documented example continues with an assistant turn carrying suggested_diffs and a model name such as claude-sonnet-4-5, and the conversation record adds source, a name for the conversation, and created_at as an epoch millisecond value. Two details matter if you plan to train on this. The model field tells you which model produced which turn, and the source field tells you which product and mode the conversation came from, so both are worth keeping in your pipeline rather than dropping as metadata.

## Output is a directory of timestamped files, one per run

Nothing is written in place. Each script creates an extracted_data/ directory and drops a timestamped JSONL file into it, so successive runs accumulate instead of overwriting, which matters when you are extracting from several assistants across a week. The documented listing shows the pattern, with a name per vendor and a YYYYMMDD_HHMMSS stamp: claude_code_conversations_20250116_143022.jsonl, cursor_complete_20250116_143045.jsonl, gemini_conversations_20250116_143145.jsonl, codex_conversations_20250116_143102.jsonl, trae_conversations_20250116_143115.jsonl, windsurf_conversations_20250116_143130.jsonl, continue_conversations_20250116_143145.jsonl, opencode_conversations_20250116_143200.jsonl. Because the stamp is in the filename, downstream tooling has to glob rather than read a fixed name. For anything that touches client work, remember that the file paths inside code_context, and the code inside them, are real source from your machine, and that the output directory is plain JSONL with no access control of its own.

## No releases, no licence, and a privacy filter nobody explains

The repository has no GitHub releases, so there is no version to pin and no upgrade path beyond pulling main. The last push was on 2026-09-03, which is recent enough that the scripts are being adjusted as vendors move their storage around, and there is a tests/ directory in the tree, though the documentation does not describe how to run it. There is no LICENSE file at the root either, so no licence identifier is stated, which is worth settling before you ship anything built on it. The most interesting gap is a pair of files the documentation never reaches: filter_privacy.py and a separate requirements-privacy-filter.txt, next to corpus_to_skills.py, which by its name turns a corpus into skills. filter_privacy.py is the one to read first, because a toolkit whose entire output is your personal conversation history across nine products, including file paths, code snippets and diffs, should be run with its filter rather than without it. What it removes is not documented, so treat it as unverified until you have read the code.

## If you use one assistant, read its files instead

The alternative approach is worth weighing before you adopt the toolkit. If you only ever used Claude Code, its history is already JSONL session files, and a line-oriented tool over those files gives you the same records with no schema translation, no OS detection and no guessing about paths. The toolkit earns its keep when the data is not already line-oriented: Cursor's four storage generations, the content-addressed blob store in cursor-agent, and OpenCode's mix of SQLite, JSON and Tauri .dat files are all cases where reading the raw artefacts by hand is a real project. The cost of the normalisation is a permanent maintenance bill. Each script encodes a vendor's private schema, and when that vendor ships a new layout the extractor quietly returns less, with no error you would notice unless you check the output count against what the client shows. For a one-off archive the honest advice is to run the single script you need, count the conversations, and compare.

## Conclusion

Use ai-data-extraction if you are building a training corpus or an archive and you have used more than one of these assistants, because the normalisation across four different Cursor storage generations and OpenCode's SQLite plus JSON layouts is work nobody has written for you. Do not use it if you use exactly one assistant whose history is already JSONL, since a jq pipeline over ~/.claude gets the same records with no schema guesswork, and do not point it at a machine with client code in it before you have read filter_privacy.py, which sits at the repository root unexplained. Verify first that the search paths on your machine match, that the target file is the one you think, and that no release has appeared since the last push on 2026-09-03, since the repository publishes no versions at all.

## FAQ

### Does ai-data-extraction need any pip packages?

No. The README states there are no dependencies and that the scripts use the Python 3 standard library, with the only setup step being python3 --version to confirm Python 3.6 or newer.

### Where does ai-data-extraction write its output?

Each script creates an extracted_data/ directory and writes timestamped JSONL files into it, such as claude_code_conversations_20250116_143022.jsonl, so runs accumulate rather than overwrite.

### Does ai-data-extraction cover the Cursor terminal CLI as well as the editor?

Yes, with two scripts. extract_cursor.py covers Chat, Composer and Agent data from the editor across four storage generations, and extract_cursor_cli.py covers cursor-agent history stored in ~/.cursor/chats/<chatId>/<agentId>/store.db.

### Which assistants can ai-data-extraction read?

Nine: Claude Code and Claude Desktop, Codex, Cursor, the cursor-agent CLI, Trae, Windsurf, Continue, the Gemini CLI, and OpenCode in both its CLI and desktop forms.

### Is there a privacy filter in ai-data-extraction?

A file called filter_privacy.py and a separate requirements-privacy-filter.txt exist at the repository root, alongside corpus_to_skills.py. The documentation does not describe what the filter removes or how to invoke it, so read the code before running an export on a work machine.

## Sources

- [0xSero/ai-data-extraction on GitHub](https://github.com/0xSero/ai-data-extraction)
- [Issues](https://github.com/0xSero/ai-data-extraction/issues)
- [README](https://github.com/0xSero/ai-data-extraction/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/0xsero-ai-data-extraction
