Model or dataset
SHITIANYU-hue/SheetasToken avatar
SHITIANYU-hue/SheetasToken

SheetasToken: a two-stage sheet retriever from the Sheet As Token paper

Implementation and resources for Sheet as Token, a graph-enhanced framework for multi-sheet spreadsheet understanding and retrieval.

605 stars106 forksPythonLicense varies

At a glance

What is it?
The repository implements the paper's pipeline: a fine-tuned BGE sheet encoder followed by a gated relational GNN over top-50 candidates. It ships metadata-only data, training scripts and baselines, but no checkpoints.
Who is it for?
Adopt SheetasToken if you are reproducing the Sheet As Token paper or want a reference implementation of query-sheet retrieval followed by graph-based cross-sheet selection, and you can supply your own checkpoints and cell values. Do not adopt it if you need a working spreadsheet search service today: the public release is metadata-only, ships no checkpoints, and its data is a single domain corpus.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 64 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What SheetasToken solves, and for whom

Spreadsheet retrieval is not document retrieval. A workbook has many sheets, each with a name, dimensions and column headers, and the question is which sheet answers a query, not which paragraph. SheetasToken is the code release for the paper Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding. It targets researchers and engineers who need a reference implementation of that retrieval problem rather than a hosted product.

The repository states its scope plainly: it contains the two-stage retrieval pipeline, the Stage 1 sheet encoder, the Stage 2 graph retriever, and the experiment scripts used in the paper. The audience is narrow by design. If you are building a general spreadsheet assistant, most of what you need is absent: there are no checkpoints, no result JSON, and the checked-in sheet files are metadata-only, carrying sheet IDs, names, dimensions and column names with no cell or example values. The README says cell and example values are not used by the current paper model, so this is a deliberate constraint, not an oversight. Anyone expecting to drop in a workbook and get answers is looking at the wrong repository.

The two-stage pipeline: BGE encoder, then a gated relational GNN

Stage 1 fine-tunes BGE for query-sheet retrieval and serializes only the sheet name, dimensions and column headers. That serialization is the sheet token: the representation deliberately excludes cell contents. Stage 2 performs query-conditioned cross-sheet retrieval over a candidate workspace, using a gated relational GNN over the retrieved top-50 candidates. The README describes the current paper model as a fine-tuned BGE Stage 1 plus a gated relational GNN Stage 2 over real full-corpus top-50 candidates.

Two Stage 2 variants exist in the tree. models/stage2/stage2_gtn_baseline.py is the shallow graph retriever, described as the architecture ablation or shadow model. models/stage2/stage2_gtn_v2.py is the enhanced retriever with stronger relational composition and is the full model. Stage 1 has three files, and the distinction matters: biencoder_model_with_example.py uses example-enhanced serialization, biencoder_model_wo_example.py omits column examples, and biencoder_model.py is a legacy baseline the README keeps for reference only, with the note that current paper experiments use the two variants above.

The data flow is therefore query, then Stage 1 candidate sheets, then a Stage 2 graph pass that can pull in sheets related to the candidates through dependency edges. That second hop is the whole point of the graph: a sheet that does not match the query text may still be the right answer because it is connected to one that does.

Installing SheetasToken and running a first training script

There is no package on an index and no install target. The README gives one dependency step and then shell scripts. Clone the repository, create an environment, and install from requirements.txt.

bash
pip install -r requirements.txt

The file pins nothing except lower bounds: torch, transformers, pyyaml, tqdm, scikit-learn, wandb, numpy, matplotlib, sentence-transformers>=3.0, openai>=2.0, pydantic>=2.0, httpx>=0.27 and pytest>=8.0. Unpinned torch and transformers mean the environment you resolve today may differ from the one the paper used; record what you install.

The scripts default to the Hugging Face model name bert-base-uncased, and the README shows overriding MODEL_NAME when you have a local snapshot.

bash
MODEL_NAME=/path/to/local/model bash scripts/stage2/train_enhanced_freeze.sh

Stage 1 training has two entry points, one for each serialization variant.

bash
bash scripts/stage1/train_with_example.sh
bash scripts/stage1/train_wo_example.sh

Stage 2 training freezes Stage 1 and takes a checkpoint path. The README's example passes both MODEL_NAME and STAGE1_CKPT, with the checkpoint named best_model_with_example/classifier.pt.

bash
MODEL_NAME=/path/to/local/model \
STAGE1_CKPT=best_model_with_example/classifier.pt \
bash scripts/stage2/train_enhanced_freeze.sh

Data selection is an environment variable or a flag: point --data-dir or DATA_DIR at data/industrytab_614 or data/industrytab_1k. The top-level data/sheets.json, query.json, train.json and dependency_edges.json are compatibility copies of IndustryTab-1K, so the repository default is the 1,002-sheet corpus. Other overrides the README lists are OUTPUT_DIR, TB_DIR, BEST_MODEL_DIR and FINAL_MODEL_DIR. For a full reproduction, the README points to scripts/reproduce_paper/README.md, which it says covers release validation, both datasets, all comparison methods, three-seed aggregation, sensitivity plots and A40 latency.

What is missing from the public release

The most concrete limitation is stated in the README itself: checkpoints and result JSON files are not included in this public code release. You cannot evaluate the paper model without training Stage 1 and Stage 2 yourself, and the scripts expect a checkpoint path for the frozen Stage 2 step. That turns a reproduction into a compute commitment before you see a single metric.

The data compounds this. Both IndustryTab variants are metadata-only, so the sheet token is built from names, dimensions and column headers alone. The README is explicit that cell and example values are not used, which means the approach cannot distinguish two sheets with identical headers and different contents. Whether that is acceptable depends entirely on your corpus.

There is also a versioning wrinkle worth reading carefully. IndustryTab-1K expands the original corpus rather than defining a disjoint collection, and the 614-sheet corpus retains an obsolete 134-query arXiv snapshot as query_legacy_134.json for provenance. If you compare numbers across the two datasets, you are comparing an expansion against its source, not two independent benchmarks.

Finally, the repository has no license file listed among its top-level entries. The README does not state terms for reuse, and the GitHub metadata does not supply one either. Treat that as unresolved rather than permissive.

Zero-shot baselines and how they differ in approach

The repository ships three no-training comparison systems, and they are the honest alternative to the trained pipeline. Frozen embedding retrieval uses BAAI/bge-base-en-v1.5 cosine similarity over all sheets. The full-corpus LLM selector asks an OpenAI model to pick sheet IDs directly from the complete sheet catalog. The local LLM selector is a controlled-candidate Ollama run with an explicit Q4_K_M 1.5B model as the default.

The difference in approach is worth stating precisely. The trained pipeline retrieves candidates with a fine-tuned encoder and then reasons over cross-sheet structure with a graph network. Frozen embedding retrieval does no training at all and no relational reasoning: it ranks sheets by vector similarity, full stop. The LLM selectors skip retrieval as a separate stage and ask a model to choose from the catalog, which trades the graph's structural signal for the model's prior knowledge and, in the OpenAI case, for per-call cost.

The baselines are runnable through shell scripts, and the README shows the OpenAI key being exported before the run.

bash
bash scripts/baselines/run_embedding.sh

export OPENAI_API_KEY=...
bash scripts/baselines/run_llm.sh

bash scripts/baselines/run_ollama.sh

All three report Precision, Recall, HitRate, MRR and nDCG at K, and the README says they use the same sheet serialization as the main pipeline. That shared serialization is what makes the comparison meaningful: the difference you measure is the model, not the input format. baselines/README.md covers dry runs, cost-safe smoke tests and model overrides.

Maintenance status and upgrade cost

The repository is not archived, and the last push was on 2026-07-29. There are no releases, so there is no version to pin and no changelog to read. Upgrading means pulling the branch, and the practical cost sits in requirements.txt: sentence-transformers>=3.0, openai>=2.0 and pydantic>=2.0 are lower bounds, not pins, and torch and transformers carry no constraint at all. A fresh environment months from now can resolve to a different stack than the one the scripts were written against.

The shell scripts soften this somewhat. MODEL_NAME, DATA_DIR, STAGE1_CKPT, OUTPUT_DIR, TB_DIR, BEST_MODEL_DIR and FINAL_MODEL_DIR are all overridable, so you can redirect outputs and checkpoints without editing the scripts. That helps with path drift, not with dependency drift.

On licensing, the repository does not state a license in the README, and no license file appears among the top-level entries. Without a stated license, the default position is that no rights are granted, which is a problem for any use beyond reading the code. This is not legal advice; if you intend to build on it, resolve the license with the maintainer first. Note also that the baselines call external services: the OpenAI selector sends sheet catalogs to a third-party API, and the README's example exports OPENAI_API_KEY, so data handling is your responsibility, not the repository's.

Editorial conclusion

Adopt SheetasToken if you are reproducing the Sheet As Token paper or want a reference implementation of query-sheet retrieval followed by graph-based cross-sheet selection, and you can supply your own checkpoints and cell values. Do not adopt it if you need a working spreadsheet search service today: the public release is metadata-only, ships no checkpoints, and its data is a single domain corpus. Before anything else, verify the license, since the repository does not state one, and run the validation commands in data/README.md to confirm the corpus counts match.

Frequently asked questions

Does SheetasToken include trained checkpoints?

No. The README states that checkpoints and result JSON files are not included in this public code release, so both Stage 1 and Stage 2 must be trained from the provided scripts.

What data does SheetasToken ship with?

Two metadata-only dataset variants: data/industrytab_614 with 614 sheets and a 1,453-query workload, and data/industrytab_1k with 1,002 sheets and 1,797 queries. The checked-in sheet files contain sheet IDs, names, dimensions and column names, with no cell or example values.

What is a sheet token in SheetasToken?

Stage 1 serializes only the sheet name, dimensions and column headers, and fine-tunes BGE over that representation for query-sheet retrieval. Cell and example values are not used by the current paper model.

Official sources

  1. Issues
  2. README
  3. SHITIANYU-hue/SheetasToken on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shitianyu-hue-sheetastoken.svg)](https://hysenlabs.com/projects/shitianyu-hue-sheetastoken)