# DeTikZify pins torch to one patch series and needs a full TeX Live on top of it

> DeTikZify turns sketches and existing figures into semantics-preserving TikZ programs, refining them with an MCTS search that needs no retraining. The package pins torch to a single patch series, pulls in a web stack by default, and cannot be installed with pip alone because TeX Live 2023, ghostscript and poppler sit outside Python.

**potamides/DeTikZify** — Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.

- Repository: https://github.com/potamides/DeTikZify
- Website: https://nllg-detikzify.hf.space
- Stars: 1,820 · Forks: 93
- Language: Python
- License: Apache-2.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/potamides-detikzify

## The install is one pip line and three system packages

The package installs from git with an extra that is conditional on which model generation you want:

```sh
pip install 'detikzify[legacy] @ git+https://github.com/potamides/DeTikZify'
```

The legacy extra is only needed for the v1 models, and the readme says it can be removed if you only plan to use v2. For running the bundled examples, the alternative is to clone the repository and install editable with the examples extra.

The part pip cannot do is the part that matters most for a figure tool. A full TeX Live 2023 installation is required, along with ghostscript and poppler, and all three have to come from a package manager or another route outside Python. That is not a small footnote. TeX Live is a multi-gigabyte installation with a package manager of its own, and the compile-and-rasterize path the library exposes depends on it being present and complete. Installing the Python package alone gives you a model that can emit programs it cannot check.

Three optional extras extend the base install. The examples extra pulls in the evaluate group plus diffusers, the legacy extra adds timm, and a separate deepspeed extra is available for distributed training work.

## Dependencies are pinned tight enough to collide with a shared environment

The dependency list is where a careful reader should stop. The torch constraint is ~=2.7.1 and torchvision is ~=0.22.1, which for a two-part version means the whole install stays inside one patch series. transformers is held at ~=4.52.4 with the accelerate and tokenizers extras, numpy at ~=2.1.1, Pillow at ~=10.4.0, datasets at ~=3.6.0.

So this is not a library that floats its way to a compatible version. It asks for a specific window and will fight an environment that already resolved a different one, which is the usual trade in a research codebase and the reason an isolated environment or a container is close to mandatory here.

Two of the pins carry an explanatory comment, which tells you they were not chosen casually. fastapi and pydantic are held at specific versions with references to Gradio issues, one about a dependency conflict and one about a breaking change, so the web layer is pinned to what the UI needs rather than to what is current.

A third pin exists for a Python version rather than a package. audioop-lts is declared with a marker for Python 3.13 and above, referencing an issue in this repository, which is the standard workaround for a module removed from the standard library in that release. It is a small dependency, but its presence tells you the web UI path is tested against current Python rather than a subset of it.

## Version numbers come from tags, so the package version and the model version differ

The project metadata declares the version as dynamic, which means the number is not written in the file. It is derived by setuptools_scm from the repository state at build time and written into the package as a generated file, with a parent directory prefix configured so a version can be recovered from an unpacked sdist.

That has a practical consequence. The package version and the model version are two different numbering systems, and the readme's news entries are all about models: a v2 model released in December 2024, adapters in March 2025, and a v2.5 model in June 2025 that was made the default in the hosted space. The model naming on the Hugging Face collection runs separately from the repository tags.

The release tags are much older than the branch. The three most recent are v0.3.0 from 2025-07-07, v0.2.1 from 2025-03-19 and v0.2.0 from 2024-12-04, while the default branch was last pushed on 2026-09-17. So a checkout from the branch is a much later state than any tag, and a version reported by the installed package reflects your checkout rather than a release you can point at.

The build requirements reinforce the pattern. The backend is setuptools with setuptools_scm for version discovery, both constrained to minimum versions, and the package finder includes only the detikzify namespace, so the examples directory is not importable as an installed package even though the examples extra exists to install its dependencies.

## One sample is one attempt, and the search is where quality comes from

The programming interface separates generating from searching, and the difference matters for cost. A single program is one call, and the result carries a check: if the program compiles, it can be rasterized and displayed, and the property that tells you this is exposed on the object rather than returned as an error.

The second path runs a search. The example passes a timeout of 600 seconds, iterates over a pipeline method that yields a score and a figure pair, and collects results into a set, so the natural usage is to spend a fixed wall-clock budget producing several candidate programs and then choose among them. The abstract calls this an MCTS-based inference algorithm and notes that it refines outputs iteratively without additional training.

That design has a direct consequence for anyone planning a pipeline. Throughput comes from search time rather than from a single forward pass, and a ten-minute budget is a real compute cost per figure. It also means the failure mode is different from a generative model that returns one bad output: you get multiple scored candidates, and the question becomes how you select among them.

The compile step is what makes the scores meaningful, since a program that does not build cannot be compared against a reference image. That is also where the TeX Live dependency earns its cost.

## Text conditioning works through adapters and not through the web interface

A second capability sits on top of the base model: synthesis conditioned on a text description rather than on a sketch. Two adapter families are offered, one described as zero-shot text conditioning that plugs directly into the v2 model, and a second with additional end-to-end fine-tuning at a larger size.

The readme is explicit about the access route. Text-conditioned synthesis is currently only supported through the programming interface, so anyone expecting to type a caption in a browser has to write a script.

The examples show the difference in call shape. The text path passes a caption to the sampling call, while the image path passes an image URL. The text-conditioned example uses a caption describing a multi-layer perceptron with two hidden layers, which is the kind of architectural figure the tool is aimed at, and the v2 adapter example loads the base model first and then an adapter on top of it.

Model sizes named across the readme range from a 1b variant to 8b and 10b, and the 1b line matters for a practical reason: the free hosted space is documented as supporting inference for the 1b models only. Anyone planning to run the larger checkpoints needs local hardware.

## The hosted space is the documented fallback, and it has a thirty minute restart

The installation section carries a tip that is really a fallback strategy, and the numbers in it are worth reading before you start debugging a local install. The hosted Hugging Face space is offered for anyone having trouble with installation or inference on their own hardware, and the note that restarting the space can take up to thirty minutes is a real constraint on that route.

Long queues have an answer too, and it costs money. The space can be duplicated with a paid private GPU runtime, or run locally with Docker. A Google Colab demo is offered as well, with the caveat that setting up the environment there takes time and the free tier only supports inference for the 1b models.

So the fallback ladder is: fix your local install, or wait in a shared queue, or pay for a private runtime, or run the Docker image yourself, or use Colab with the smaller models. The fact that a research project documents its own frustration this plainly is a sign of honesty about the hardware requirement, and it is a reminder that the practical entry point for most users is the hosted space rather than the pip line.

The model quality story is separate from the distribution story. The v2.5 model was made the default in that space after release, and the readme points to model cards for the details of each checkpoint.

## A research codebase with a research repository shape

The repository is organized like a research project rather than a library, and the top-level entries reflect that: a gitattributes file, a GitHub directory, a gitignore, a license, a manifest file, the readme, the detikzify package, an examples directory and the project metadata.

The examples directory carries its own readme and a set of scripts with names that map directly to the workflow. There are scripts to run inference, to evaluate, to pretrain, to refine, to sketchify, to train, and a subdirectory for the text-conditioning work. That naming is the clearest statement of the project's scope, because it includes training and evaluation code alongside inference, which a consumer library would not.

The research framing shows in the citations too. The work was accepted at NeurIPS 2024 as a spotlight, and a related system was accepted as a highlight paper at ICCV 2025, with an arXiv preprint linked for the main system and a second one for the adapters.

The license is Apache-2.0, recorded in the metadata as a text field rather than as a modern SPDX expression, which is an older style that some tooling does not read as a license identifier. The project requires Python 3.11 or later within the 3 series, which pairs with the audioop-lts marker for 3.13 and above rather than contradicting it.

## Conclusion

DeTikZify fits a research group that needs figures as editable vector source rather than as images, and that already has a LaTeX toolchain on the machine, because the value it offers is a program you can edit and recompile rather than a picture you cannot. The MCTS search is the part to understand before you plan around it: a single sample is one attempt, and the search spends a wall-clock budget producing several candidate programs with scores, so quality comes from search time rather than from a better model. Two things to check first. Read the pin list in the project metadata before you install, since the torch and transformers constraints are narrow enough that a conflicting package elsewhere in your environment will decide whether this installs at all. And decide how you will run the heavier models, because the free hosted route is a shared space that may sit in a queue for up to thirty minutes and whose free tier covers the smaller models only. Anyone who needs batch figure generation should budget for GPU memory and a TeX installation before starting.

## FAQ

### What does DeTikZify need installed besides its Python package?

A full TeX Live 2023 installation, ghostscript and poppler, none of which come from pip. The Python install is pip install 'detikzify[legacy] @ git+https://github.com/potamides/DeTikZify', where the legacy extra is only needed for the v1 models.

### Which Python and torch versions does DeTikZify support?

The project metadata requires Python 3.11 within the 3 series, and pins torch to ~=2.7.1 with torchvision at ~=0.22.1, so the install stays inside one patch series. transformers is held at ~=4.52.4 and numpy at ~=2.1.1.

### How does DeTikZify produce a better figure than a single generation?

It uses an MCTS-based inference algorithm that refines outputs iteratively without additional training. The pipeline exposes a simulate call that takes a timeout, so a wall-clock budget of several hundred seconds yields multiple scored candidate programs rather than one answer.

### Can DeTikZify generate TikZ from a text description?

Yes, through adapters. TikZero adapters plug into the v2 model for zero-shot text conditioning, and a further fine-tuned variant is offered at a larger size. The readme states this is currently supported only through the programming interface, not the web UI.

### What is the newest DeTikZify release tag?

v0.3.0, published 2025-07-07, after v0.2.1 and v0.2.0. The default branch was last pushed on 2026-09-17, and the package version is derived from the checkout at build time rather than written into the metadata.

## Sources

- [License: Apache-2.0](https://github.com/potamides/DeTikZify/blob/main/LICENSE)
- [potamides/DeTikZify on GitHub](https://github.com/potamides/DeTikZify)
- [Project website](https://nllg-detikzify.hf.space)
- [README](https://github.com/potamides/DeTikZify/blob/main/README.md)
- [Releases](https://github.com/potamides/DeTikZify/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/potamides-detikzify
