# SenseNova-Vision: One 7B Model That Answers Vision Tasks in Text and Images

> SenseNova-Vision turns segmentation, depth, keypoints, OCR and multi-view reconstruction into instruction-following generation rather than task-specific heads. The 7B-MoT checkpoint is Apache-2.0, but the hardware bar is an 80GB GPU.

**OpenSenseNova/SenseNova-Vision** — Vision as Unified Multimodal Generation

- Repository: https://github.com/OpenSenseNova/SenseNova-Vision
- Stars: 703 · Forks: 40
- Language: Python
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/opensensenova-sensenova-vision

## The problem: every vision task ships its own head

Conventional computer vision pipelines fragment. A detector outputs boxes through a regression head, a segmentation network outputs masks through a decoder, a depth model outputs a dense map through yet another branch. Each task carries its own architecture, its own training recipe and its own output decoding. Adding a task means adding a model. SenseNova-Vision attacks that fragmentation directly: it formulates computer vision as unified multimodal generation, expressing heterogeneous visual tasks through the native text and image generation spaces of a unified multimodal model. The audience is teams that already run several vision models side by side and want to collapse them into one instruction-following checkpoint, plus researchers who want a single formulation to extend to new task variants without designing a new head. The repository states that the training requires no task-specific prediction heads, decoders, or architectural branches.

## Text tokens for symbols, image tokens for dense fields

The mechanism rests on a split between two output spaces. Text generation carries symbolic visual records: categories, boxes, points, OCR strings, keypoints and camera parameters. Image generation carries dense spatial targets: segmentation masks, depth maps, surface normals and multi-view point maps. Mixed text-image responses handle compositional tasks that need both. Natural-language instructions plus optional visual prompts specify the task, the target regions or views, the output schema and the decoding convention. That last part matters for evaluation: the README says outputs stay decodable for standard benchmarks, so a textual box string can be parsed back into benchmark-specific structures. The training data follows the same shape. Heterogeneous computer-vision annotations are converted into instruction-response examples to build the SenseNova-Vision Corpus, spanning decodable text, image and mixed targets. Training starts from an off-the-shelf pretrained unified multimodal model and runs primarily on that corpus, with auxiliary multimodal data used to preserve general understanding and generation. The practical consequence is that a new task is largely a prompt-engineering and data-conversion exercise, not an architecture change, provided the target fits either the text or the image space.

## Installing SenseNova-Vision and running a first segmentation request

The repository provides one entrypoint script for examples, single-image inference, interactive inference and benchmark inference. Environment creation happens from the repository root. The setup script takes an environment name, and the README uses sensenova-vision as that name.

```bash
git clone https://github.com/OpenSenseNova/SenseNova-Vision.git
cd SenseNova-Vision
bash setup.sh sensenova-vision
conda activate sensenova-vision
```

After activation you are inside the environment described by requirements.txt, which pins torch 2.5.1, torchvision 0.20.1, transformers 4.49.0 and gradio 5.38.2, with a stated reference runtime of Linux, Python 3.10 and CUDA 12.4. The curated example is the fastest way to confirm the install works end to end.

```bash
bash scripts/run_sensenova_vision.sh example
```

A single inference request takes a task name, an instruction and an image path. The README's own example asks for binary segmentation of a person in examples/images/2.jpg.

```bash
bash scripts/run_sensenova_vision.sh inference \
  binary_seg \
  "person" \
  examples/images/2.jpg
```

For an interactive session, the demo wrapper prints the local URL before starting Gradio, and the model path is supplied through MODEL_PATH.

```bash
MODEL_PATH=/path/to/SenseNova-Vision-7B-MoT \
  bash scripts/run_sensenova_vision.sh demo
```

The README records that the default BF16 web demo ran on 1 x NVIDIA A800 80GB GPU without CPU or disk offload, peaking at 38,226 MiB of process GPU memory. It recommends an 80GB GPU and states that GPUs with less memory have not yet been validated across the full task set.

## The 80GB floor and the untested smaller-GPU path

The memory number is the constraint to plan around. A recorded peak of 38,226 MiB for a single demo request sits comfortably under 80GB but well above the 40GB and 48GB cards many teams have in a workstation. The README is explicit that the 80GB figure is a recommendation and that smaller GPUs are unvalidated across the full task set, which means a 48GB card may work for one task and fail on another without any documented boundary between them. Benchmarking raises the bar further: the README recommends at least one 8 x 80GB GPU machine for the full benchmark, and training requires at least 2 machines with 8 x 80GB GPUs each, with 32 or more such machines recommended. That is a cluster-scale training footprint, not a fine-tuning-on-one-node footprint. A second limitation is documentation depth. The README points to docs/INFERENCE.md, docs/EVAL.md, docs/data_prepare.md, docs/train_data_prepare.md and docs/TRAIN.md, but it does not describe rollback, checkpoint versioning or an upgrade procedure between releases, so a team that pins a working configuration has no documented path back if a later change breaks it. Finally, this is the wrong tool if you need a small, single-purpose model. If your entire requirement is bounding boxes on a fixed class list at low latency, a dedicated detector will be cheaper to host and simpler to reason about than a 7B multimodal model that also carries depth, normals and multi-view geometry you will never call.

## How it differs from task-specific vision stacks

The obvious alternative is the conventional route: a detector such as a YOLO-family model for boxes, a segmentation model for masks, a monocular depth network for depth, and separate code for anything geometric. The difference is not accuracy, which this material does not let me compare. The difference is in what changes when you add a task. In the conventional stack, each new capability means a new checkpoint, a new preprocessing pipeline and a new output format to parse. In SenseNova-Vision, the new capability is expressed as an instruction plus an output schema, and it lands in whichever of the two generation spaces fits: text for symbolic records, image for dense fields. That also changes failure behaviour. A dedicated detector fails in ways you can characterise per task. A unified generative model can produce a malformed text string that fails to parse, or an image output that decodes to something plausible but wrong, and the README's emphasis on decodable outputs suggests the authors treat parseability as a first-class concern rather than an afterthought. The trade-off is real: you gain one model and one interface, and you accept a much larger memory footprint and a prompt format you must get right.

## Licence, maintenance and what an upgrade actually costs

The repository is licensed Apache-2.0, with the LICENSE file at the top level and the licence badge in the README. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, but it also carries notice and attribution obligations: if you redistribute the code or a derivative, you keep the licence text and state what you changed. The model weights and the SenseNova-Vision-Corpus-50M dataset are hosted separately on Hugging Face, so their terms should be checked on those pages rather than assumed from the repository licence. This is a description of the licence text, not legal advice; get counsel if you plan to redistribute. On maintenance, the last push to the repository was on 2026-09-01, and the repository is not archived. The news entries in the README run from 2026-07-08, when the weights, inference code, corpus and technical report first appeared, through 2026-07-22, when the training data preparation workflow was expanded. There are no retrieved releases, so there is no tagged version to pin and no changelog to read. Practically, an upgrade means pulling the branch and re-reading the docs, because requirements.txt pins exact versions of torch, transformers and gradio, and a dependency bump is as likely to break an environment as a code change. Budget for re-validating your task against the pinned stack after any pull.

## Conclusion

Adopt SenseNova-Vision if you need one checkpoint covering structured visual understanding, dense geometric prediction, segmentation and multi-view geometry, and you have 80GB-class GPUs to run it. Do not adopt it if your target is a single narrow task on modest hardware, or if you need a documented rollback path, since the README describes no versioning or upgrade procedure. Before committing, verify three things: that the 7B-MoT weights load under the pinned transformers 4.49.0 and torch 2.5.1 in requirements.txt, that your task appears in examples/ with a working run_sensenova_vision.sh invocation, and that your GPU memory ceiling survives the demo's recorded 38,226 MiB peak. The repository's own note that GPUs with less memory have not yet been validated across the full task set is the boundary to test against first.

## FAQ

### How does vision AI work?

In SenseNova-Vision, natural-language instructions and optional visual prompts specify the task, target regions or views, output schema and decoding convention. The model then responds in text for symbolic records such as boxes, points, OCR strings and camera parameters, in images for dense targets such as masks, depth maps, normals and point maps, or in a mix of both.

### What does vision mean in AI models?

In this repository, vision is treated as a generation target rather than a set of separate prediction heads. The README describes symbolic visual records going into the text space and dense spatial targets going into the image space, with mixed responses for tasks that need both.

### How does vision AI work in models that generate text instead of boxes?

SenseNova-Vision expresses symbolic visual records such as categories, boxes, points, OCR strings and keypoints as text, and dense targets such as segmentation masks, depth maps, surface normals and multi-view point maps as images. The README states that outputs remain decodable for standard benchmarks.

## Sources

- [Issues](https://github.com/OpenSenseNova/SenseNova-Vision/issues)
- [License: Apache-2.0](https://github.com/OpenSenseNova/SenseNova-Vision/blob/master/LICENSE)
- [OpenSenseNova/SenseNova-Vision on GitHub](https://github.com/OpenSenseNova/SenseNova-Vision)
- [README](https://github.com/OpenSenseNova/SenseNova-Vision/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/opensensenova-sensenova-vision
