Model or dataset
stardustai/dataset-viewer avatar
stardustai/dataset-viewer

Dataset Viewer: a Tauri desktop app for streaming 100GB+ files from S3, WebDAV, SSH and Hugging Face

A sleek dataset viewer built entirely by AI Agent. Supports streaming large files from WebDAV, S3, SSH, Local or Hugging Face.

1,079 stars74 forksTypeScriptLicense varies

At a glance

What is it?
Dataset Viewer is a cross-platform desktop viewer built with Tauri, React and TypeScript that streams large files from remote and local sources instead of downloading them whole. It is a reasonable fit for log and Parquet triage, and a poor fit if you need a browser-based or self-hosted shared viewer.
Who is it for?
Adopt Dataset Viewer if you work on a desktop and regularly need to open multi-gigabyte CSV, Parquet or log files that live on S3, WebDAV, SSH, SMB or Hugging Face, and you want archive and document preview in the same window.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Dataset Viewer solves, and for whom

The problem is opening a file that does not fit in memory. A 40GB CSV, a Parquet dump, a rotated log file. The usual answer is to download it first, then open it in an editor that either refuses or stalls. Dataset Viewer takes the other route: it keeps the file where it lives and reads it in chunks.

The README frames the target audience explicitly under a section called Perfect For: data scientists exploring large datasets and Parquet files, people searching massive log files without memory constraints, archive management without extraction, remote access over WebDAV, SSH/SFTP, SMB and cloud storage, and performance-critical work. That list is coherent. Every item on it is a case where the bottleneck is I/O and the file is too large to copy casually.

The project describes itself as 100% AI-generated, and the repository carries a badge to that effect. Treat that as a statement about provenance, not about quality. What matters for adoption is that the code is TypeScript on the frontend and Rust in src-tauri, that the build pipeline is real (pnpm workspaces, Biome, Husky, cargo clippy), and that releases exist. The last push to the default branch was on 2026-03-28, which is under six months before today, so the repository is not stale, but the most recent release listed is v1.6.4 from 2025-10-17. That gap between commits and releases is worth knowing about before you plan an upgrade path.

How the streaming and search actually work

Dataset Viewer is a Tauri application, which means the window you see is a webview rendering React, and the file access happens in a Rust process behind it. The repository layout confirms this split: src/ holds the React and TypeScript frontend, src-tauri/ holds the Rust backend, and vite.config.ts drives the frontend build.

The README describes the mechanism in two lines under Technical Highlights: chunked loading with virtual scrolling for millions of rows, and large file chunked transmission with no full extraction needed. The frontend dependencies back that up. @tanstack/react-virtual is listed in package.json, which is the virtualization layer that renders only the visible rows of a table. hyparquet and hyparquet-compressors are listed as well, which is how Parquet is read without a server-side conversion step.

So the data flow is: the Rust side opens a connection to the chosen backend (local path, WebDAV, SSH/SFTP, SMB/CIFS, S3 or Hugging Face Hub), reads a range of bytes, and hands that range to the frontend. The frontend parses the slice it needs and asks for more as you scroll. Search is described as millisecond search with highlighting across massive files, which only works if the search runs over the streamed index rather than a fully materialized copy.

Archive preview follows the same pattern. The README says ZIP and TAR files are browsed without extraction, so the archive central directory is read and individual entries are streamed on demand.

One consequence worth stating plainly: because parsing happens in the webview, the formats that get first-class treatment are the ones with a JavaScript parser available. Parquet has one in the dependency list. Excel, CSV, JSON, Markdown, PDF and Office documents are listed as supported, but the README does not say which of those are parsed in Rust versus in the webview, and that distinction determines how a 20GB Excel file will behave.

Installing Dataset Viewer and opening a remote file

The README does not document a package manager install. It points to the releases page for downloads, and the repository is a pnpm workspace for people building from source. If you want the app, get the binary for Windows, macOS or Linux from the latest release.

To build it yourself you need Node, pnpm and a Rust toolchain, because Tauri compiles a native binary. The root package.json defines the scripts. Install dependencies first:

bash
pnpm install

The postinstall hook prints a confirmation line when it finishes. Then start the app in development mode, which runs the Vite dev server and the Tauri shell together:

bash
pnpm tauri:dev

To produce a distributable build, the same file defines a packaging script that wraps cargo:

bash
pnpm package

There is also a debug variant, pnpm package:debug, and a Windows cross-target variant, pnpm package:windows, which builds for x86_64-pc-windows-gnu. Note the prebuild script: it builds the plugin SDK first with pnpm --filter @dataset-viewer/sdk build and then runs cargo run --bin export-bindings inside src-tauri. That means the frontend cannot be built standalone; the Rust bindings have to be generated first. If you try pnpm build on a clean checkout without a working Rust toolchain, expect it to fail at the prebuild step rather than at the TypeScript step.

Once the app is running, the first real use is a connection. The README's screenshots show a connection setup screen with multiple storage types. You pick one (for example S3 or WebDAV), enter the endpoint and credentials, and the app lists the objects or directories it can see. Selecting a file opens it in the appropriate viewer: a table for CSV and Parquet, a collapsible tree for JSON, a syntax-highlighted editor for code, an archive browser for ZIP and TAR. The README does not document the exact credential fields for each backend, so the connection screen itself is the reference.

Where Dataset Viewer falls short

The most concrete limitation is the absence of a documented licence in the repository metadata. The README carries an MIT badge linking to opensource.org, and the GitHub release badges are present, but the licence field for the repository is unknown. An MIT badge in a README is not the same as a LICENSE file at the root, and the top-level entries listed for this repository do not include one. Before you ship this inside a company, confirm the licence from the repository itself rather than from the badge.

Second, this is a desktop application, not a service. There is no documented server mode, no Docker image and no documented way to host it for a team. If your workflow is "send a colleague a link to the dataset," Dataset Viewer does not do that.

Third, the search claim needs a boundary. The README says millisecond search across massive files, but it does not say whether that applies to every supported format or only to the text and code viewers. Searching inside a Parquet file and searching inside a 10GB log file are different operations, and the README does not separate them.

Fourth, the plugin system is documented through a wiki and an npm package, @dataset-viewer/sdk, but the repository declares that SDK as a workspace dependency with version workspace:*, which means the SDK is built from this repository. The npm package is published separately. If you write a plugin against the npm SDK and the app expects a different binding version, the mismatch will surface at load time. The README does not document a compatibility matrix.

Finally, the project is entirely AI-generated by its own description. That is neither a defect nor a guarantee. It does mean you should read the code for the parts you depend on rather than assume conventions from a long-lived human-maintained codebase.

Dataset Viewer compared with DuckDB and the Hugging Face dataset viewer

The closest thing to Dataset Viewer for querying large local files is DuckDB, and the difference in approach is fundamental. DuckDB is a SQL engine. You point it at a Parquet or CSV file and it executes queries over it with predicate pushdown, returning aggregated results. Dataset Viewer is a viewer. You point it at a file and it renders rows, columns and text, with search and highlighting. If your question is "what is the average value in this column across 200 million rows," DuckDB answers it and Dataset Viewer does not. If your question is "what does row 4,000,000 actually look like, and where does this string appear," Dataset Viewer is the more direct tool. They are complementary, not substitutes.

The other comparison is the Hugging Face dataset viewer, which several of the related searches reference. That is a hosted web feature: you open a dataset page on huggingface.co and the site renders a preview table for you, with no install. Dataset Viewer connects to the Hugging Face Hub as one of its backends, so it can open Hub datasets locally, but it is a desktop binary you have to download and run. The trade is control and file size against zero setup. If you only ever look at small public Hub datasets, the hosted viewer is less work. If you need to open a 100GB file that lives on your own S3 bucket, the hosted viewer has nothing to offer.

Maintenance, upgrades and licence questions to settle

The repository is not archived. The last push was on 2026-03-28. The most recent release listed is v1.6.4 from 2025-10-17, following v1.6.3 and v1.6.2 earlier the same day. So the release history shows bursts of patch releases rather than a steady cadence, and the interval between the last release and the last commit is roughly five months. Plan upgrades around releases, not around commits.

Upgrade cost is dominated by the Tauri and Rust toolchain, not by the JavaScript. The root package.json pins @tauri-apps/api to ^2 and the Tauri plugins to ^2, and the prebuild script regenerates bindings from Rust. Any change to the Rust command surface means regenerating those bindings before the frontend compiles. If you build from source, budget for a working Rust toolchain and for cargo build times, which are not short for a Tauri app. If you use the prebuilt releases, the upgrade is a download, and the risk is that a new release changes the connection or plugin interface in a way the README does not record. The README does not document a changelog or a migration guide.

On licensing: the README displays an MIT badge and links to the MIT text on opensource.org, but the repository's licence field is unknown and no LICENSE file appears in the top-level entries. Treat the badge as a claim to verify, not as a settled fact. If you need MIT terms in writing, check the repository and the release artifacts before you depend on them. This is a factual gap, not legal advice.

Editorial conclusion

Adopt Dataset Viewer if you work on a desktop and regularly need to open multi-gigabyte CSV, Parquet or log files that live on S3, WebDAV, SSH, SMB or Hugging Face, and you want archive and document preview in the same window. Do not adopt it if you need a browser-accessible or self-hosted viewer that several people open from one URL, or if you need a signed and notarized macOS build, since the README documents Windows, macOS and Linux releases but says nothing about code signing. Before committing, verify three things yourself: that the release you download installs and launches on your OS, that your storage backend connects with the credentials you plan to use in production, and that the plugin SDK version published on npm matches the app version you installed, because the repository declares the SDK as a workspace package and the npm package is versioned separately.

Frequently asked questions

How can I visualize a dataset with Dataset Viewer?

Open the app, create a connection to one of the supported backends (WebDAV, SSH/SFTP, SMB/CIFS, S3, local files or Hugging Face Hub), and select a file. Dataset Viewer picks a viewer by format: a virtualized table for CSV, Excel and Parquet, a collapsible tree for JSON, syntax-highlighted text for code, and an archive browser for ZIP and TAR.

How do I analyze a dataset in Dataset Viewer?

Dataset Viewer is a viewer rather than a query engine, so analysis means opening the file and using the built-in search with highlighting across large files, plus filtering and sorting in the CSV and Excel sheet view. The README does not describe aggregation or SQL-style queries.

What is the difference between a database and a dataset in this context?

Dataset Viewer does not connect to databases. Its supported sources are WebDAV, SSH/SFTP, SMB/CIFS, S3, local files and Hugging Face Hub, and it reads files such as Parquet, CSV, Excel, JSON and code. A database would answer queries; Dataset Viewer renders the contents of a file.

Is there a VS Code dataset viewer?

Dataset Viewer is not a VS Code extension. It is a standalone Tauri desktop application for Windows, macOS and Linux, distributed through the GitHub releases page. The README does not mention any editor integration.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. stardustai/dataset-viewer on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stardustai-dataset-viewer.svg)](https://hysenlabs.com/projects/stardustai-dataset-viewer)