VQASynth: MIT in the classifier, Apache-2.0 in the repository, two onnxruntime builds
Compose multimodal datasets 🎹
At a glance
- What is it?
- A reproduction of the SpatialVLM pipeline for generating spatial-reasoning VQA datasets from image collections on the Hugging Face Hub, with three named departures from the original. Its packaging declares two licences at once, its requirements file installs the CPU and GPU builds of the same runtime together, and its in-process example stops one comment before the step it advertises.
- Who is it for?
- Use VQASynth when you already have a GPU to spare and want spatial-reasoning supervision built from an image collection you control, since the pipeline runs through Docker Compose against a config file and the in-process API lets you go one image at a time.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The in-process example stops one comment before the step it promises
The documented single-image path is the part a reader is most likely to copy, and it does not run as printed:
from PIL import Image
from vqasynth.localize import Localizer
from vqasynth.scene_fusion import SpatialSceneConstructor
from vqasynth.prompts import PromptGenerator
image = Image.open("warehouse.jpg").convert("RGB")
# Detect + segment task-relevant objects
masks, _, captions = Localizer(captioner_type="florence").run(image)
# Lift to 3D — VGGT emits per-object point clouds, depth, and intrinsics in one pass
pcd_filepaths, canonicalized, _, _ = SpatialSceneConstructor().run(
"warehouse_0", image, masks, output_dir="./scenes"
)
# Generate spatial VQAs from the reconstructed 3D`PromptGenerator` is imported and never used, and the block ends on the comment announcing spatial VQA generation with no code after it. Two of the three values from the localizer are discarded into throwaway names. The push-to-hub snippet has the same problem, calling `Dataloader(cache_dir).push_to_hub(final_dataset, target_repo_name)` where none of those three names is defined.
The classifier says MIT and the repository metadata says Apache-2.0
Two licence sources, two different answers, and no tie-breaker. The repository metadata carries Apache-2.0, and the tree holds a LICENSE file. The packaging script's classifier list says `License :: OSI Approved :: MIT License`, which is an MIT statement in the one place a package indexer would read it. Notably the setup script has no `license` field at all, so the classifier is doing the work a real declaration would normally do, and it disagrees with the metadata. Neither the README nor the visible documentation mentions licensing at all, so a reader has no third source to consult. The same file declares version 0.0.1 and the repository has no GitHub release, so the package version and any tag you might find are not connected. For anyone planning to build on the code, this needs to be settled with the author rather than inferred.
requirements.txt installs the CPU and GPU builds of the same runtime
The dependency list has nineteen entries and reads like a frozen environment, with a couple of exceptions. Seventeen are pinned to an exact version with `==`, including numpy 1.26.4, transformers 4.48.0, opencv-python 4.8.1.78, accelerate 0.34.2, onnx 1.21.0 and openai 2.24.0. Two use ranges, open3d and onnxruntime. One, bitsandbytes, carries no version at all, so it is the single dependency that can drift between two installs of the same release. The concrete hazard is `onnxruntime` and `onnxruntime-gpu` appearing together. Those are the CPU and GPU distributions of one library, not two libraries, and installing both into one environment means the second overwrites the first's files rather than coexisting with it. It is also the only dependency in the list that reaches for a GPU, which is consistent with the hardware floor the README states.
The published reasoning traces estimate from remembered object sizes
The examples section shows three prompts with full chain-of-thought, and reading them closely undercuts one of the project's own claims. The README lists base responses on consistent references like floors and surfaces as a capability of the trained models. What the traces actually do is reason from remembered dimensions: a standard pallet is around four feet wide and eight feet long, a soccer goalpost extends 2.5 metres beyond the line, a standard office chair is around 60 to 70 centimetres tall and a bookshelf anywhere from 1.2 to 1.8 metres. The distance answers then fall out of those priors, in feet for the warehouse, in metres for the pitch. So the shown reasoning grounds geometry in learned object statistics rather than in the reconstructed scene, even though the pipeline behind it builds per-object point clouds and depth with VGGT.
Four models, four superlatives, and no metric behind any of them
The model list reads like a leaderboard with the numbers removed. SpaceOm is the best overall, SpaceThinker-Qwen2.5VL-3B has the most accurate distance estimates, SpaceQwen2.5-VL-3B-Instruct is the most popular, and SpaceLLaVA 13B is the original. Each label is a different axis, chosen so that no two entries are directly comparable, and none is accompanied by a score, a dataset split or an evaluation protocol. The dataset list has the same shape, five names including SpaceOm, SpaceThinker, OpenSpaces, OpenSpaces_MC_R1 and a spacellava set, with no sizes or counts. That is not a criticism of the models, which may well be as described; it is a limit on what the page lets a reader check. The examples carry real outputs, and those are the only quantitative material on offer.
Three named departures from SpatialVLM, and one of them is an issue link
The project describes itself as an open-source reproduction of SpatialVLM, which supplies a 3D scene reconstruction pipeline and prompt templates, and it enumerates exactly where it diverges. Metric depth estimation comes from VGGT replacing DepthPro, with the stated benefit of better speed and accuracy. In the localization refinement stage, SAM2 replaces SAM. Object-grounded captions come from point prompting with Molmo, and that item is linked to a repository issue rather than to a paper or a commit. Then there is the chain-of-thought layer, multimodal thinking by CoT reasoning, which is what turns a reconstructed scene into the templated VQA chat the description talks about, with the intent being that vision-language models can be instruction-tuned on it with low-rank adapters.
An A10 or larger for the pipeline, an A100 for the notebook
The hardware requirement is stated once and it is firm: Python 3.10 or later, Docker with Compose V2, the NVIDIA Container Toolkit, at least 24GB of VRAM on an A10 or larger, and 16GB of RAM. The Colab route is stricter, requiring an A100 runtime. The batch path is Docker Compose driven, with `run.sh` at the repository root and the dataset chosen by editing `config/config.yaml`, which is the one file to touch to point the pipeline at a different image collection. The package tree sits alongside that: `.gitmodules`, `docker/`, `pipelines/`, `experiments/`, `notebooks/`, `tests/`, eight files under `examples/` including a Gradio app, and an unexplained `.remyx/` directory at the top level. Docker Compose V2 means a `docker-compose.yml`, not the legacy standalone Compose binary.
Editorial conclusion
Use VQASynth when you already have a GPU to spare and want spatial-reasoning supervision built from an image collection you control, since the pipeline runs through Docker Compose against a config file and the in-process API lets you go one image at a time. Do not pip-install it into an existing environment without reading the requirements file first, because it pins seventeen packages exactly, leaves one unpinned, and installs both the CPU and GPU builds of ONNX Runtime in the same set. Two things to settle before you depend on it: the licence, which is declared Apache-2.0 by the repository and MIT by the packaging classifier with no `license` field to break the tie, and the fact that the four listed models are ranked by four adjectives with no metric attached to any of them.
Frequently asked questions
What hardware does VQASynth need?
Python 3.10 or later, Docker with Compose V2, the NVIDIA Container Toolkit, at least 24GB of VRAM on an A10 or larger, and 16GB of RAM. The Colab notebook path is stricter and requires an A100 runtime.
How do I point VQASynth at a different image dataset?
Edit `config/config.yaml`, which the documentation names as the place where you change the dataset being processed. The pipeline then runs through Docker Compose using `run.sh` at the repository root, with authentication to the hub done first via `huggingface-cli login`.
Which licence does VQASynth use?
Two sources disagree and neither settles it. The repository metadata carries Apache-2.0 and the tree holds a LICENSE file, while the packaging classifier list says MIT. The setup script has no `license` field, so nothing in it breaks the tie, and no GitHub release exists against which to check the declared version of 0.0.1.
What does VQASynth change relative to SpatialVLM?
Three named departures: VGGT replaces DepthPro for metric depth estimation, SAM2 replaces SAM in the localization refinement stage, and object-grounded captions come from point prompting with Molmo. Chain-of-thought reasoning is added on top, turning the reconstructed scene into templated VQA data for instruction tuning with low-rank adapters.
Which VQASynth-trained model is best at distance estimation?
The documentation names SpaceThinker-Qwen2.5VL-3B as having the most accurate distance estimates, SpaceOm as the best overall, SpaceQwen2.5-VL-3B-Instruct as the most popular, and SpaceLLaVA 13B as the original. No scores, dataset splits or evaluation protocol accompany those labels in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/remyxai-vqasynth)