Model or dataset
TinyLLaVA/TinyLLaVA_Factory avatar
TinyLLaVA/TinyLLaVA_Factory

TinyLLaVA Factory: the training entry points are shell scripts you edit in place, and the metadata still names the retired repo

A Framework of Small-scale Large Multimodal Models

1,009 stars102 forksPythonApache-2.0

At a glance

What is it?
TinyLLaVA Factory is a modular PyTorch and HuggingFace codebase for small-scale multimodal models, pairing a vision tower, a connector and one of six language model families with frozen, full, partial or LoRA recipes. Its documentation is unusually specific about hyperparameters and unusually silent about configuration: the documented training workflow is four files you edit by hand, and the packaging metadata still advertises the codebase it replaced.
Who is it for?
Judge TinyLLaVA Factory on the parts it does pin and the parts it leaves in your hands. What it does well is the recipe contract: a global batch size of 256 at a learning rate of 1e-3 for pretraining and 128 at 2e-5 for finetuning, a chat template selected by a conv_version value per model family, and a connector limited to MLP, Qformer or Resampler.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The package metadata still points at the codebase this one replaced

The announcement feed records that the new codebase shipped on 2024.05.15 and that the previous one, TinyLLaVABench, was moved onto a branch of this repository rather than kept where it was. The packaging metadata was not updated to match:

toml
[project.urls]
"Homepage" = "https://github.com/DLCV-BUAA/TinyLLaVABench"
"Bug Tracker" = "https://github.com/DLCV-BUAA/TinyLLaVABench/issues"

Both URLs resolve to the old organization and the old repository name. Anyone who installs the distribution and then follows its homepage or files a bug through the tracker link reaches the retired codebase, not this one. The same file declares the distribution name as tinyllava at version 1.0.0, so the name on a pip install record does not match either the repository or the project title. Two papers are linked from the README header, one on the framework and one on this modularized codebase, and the metadata carries no reference to either.

scripts/ is excluded from the package that the training command needs

The documented way to train is a shell script inside the repository:

bash
bash scripts/train/train_phi.sh

The build configuration excludes that directory from both the package discovery and the wheel:

toml
[tool.setuptools.packages.find]
exclude = ["assets*", "benchmark*", "docs", "dist*", "playground*", "scripts*", "tests*"]

With the documented editable install, `pip install -e .`, the checkout is on disk and the command works. Install the same distribution from a wheel and scripts/ is not there, so the one training invocation the README gives has nothing to run. The same exclusion list drops assets, docs and tests from the distribution, which is unremarkable for a library and less so for a repository whose primary entry point is a shell script. The top level listing shows the directory present in git, which is why the gap only appears for people who install rather than clone.

The train extra repeats three packages the base dependencies already require

The dependency list is long and hard pinned: torch at 2.0.1 with torchvision 0.15.2, transformers at 4.40.1, tokenizers 0.19.0, numpy 1.26.4, accelerate 0.27.2, peft 0.10.0, bitsandbytes 0.41.0, gradio 3.35.2, timm 0.6.13, httpx 0.24.0, and a pydantic constraint that stops below version 2. Alongside those sits an extra meant to be the lightweight training set:

toml
[project.optional-dependencies]
train = ["deepspeed==0.14.0", "ninja", "wandb"]

Those three are already unconditional. deepspeed, ninja and wandb appear in the main dependency list as well, so the extra installs nothing extra and the marker that separates a training install from an inference one does not exist in practice. The extras mechanism is otherwise unused, which matters more for the inference path than for training: someone who wants a deployment image inherits deepspeed and a distributed training launcher whether or not they train.

flash-attn is installed outside the dependency set, with build isolation off

One package is deliberately kept out of the pinned list and installed as a separate step:

bash
pip install flash-attn==2.5.7 --no-build-isolation

Disabling build isolation means pip does not create an isolated build environment, so the build reads whatever torch is already installed in the current environment, which is why this step follows `pip install -e .` rather than preceding it. A kernel compiled against a different torch version than the one on the machine is the usual outcome of changing either side later. The version pair is the other thing to check before planning hardware: torch is held at 2.0.1 and torchvision at 0.15.2, a combination from an older generation, while flash-attn is pinned four minor releases ahead of it. The installation notes also state that these requirements differ from the ones the wider LLaVA project expects and recommend building the environment from scratch rather than adapting an existing one.

Training is configured by editing four files, with no config file or flag

The training section instructs the reader to open scripts and change values in place: data paths in scripts/train/train_phi.sh, output_dir in scripts/train/pretrain.sh, pretrained_model_path and output_dir in scripts/train/finetune.sh, then GPU ids and per_device_train_batch_size in both the pretrain and the finetune script. Nothing else is offered as a configuration surface. That makes the recipe readable, since the hyperparameters sit next to the stage that uses them, but it means two runs cannot differ without two checkouts, and a modified script is indistinguishable from a change to the repository itself. The values themselves are given as a table: a global batch size of 256 at a learning rate of 1e-3 for pretraining and 128 at 2e-5 for finetuning, with the guidance to keep both except when tuning with LoRA. Global batch size is defined as the number of GPUs times the per-device batch times the accumulation steps, so the table is a target to solve for rather than a value to copy.

The finetune template mapping stops at Qwen-1.5 while the zoo ships Qwen2

Finetuning picks a chat template through a conv_version value, and the pretraining stage uses the same value for every model. The mapping is spelled out for three groups: phi for Phi-2, StableLM and Qwen-1.5, llama for TinyLlama and OpenELM, gemma for Gemma. The model zoo lists five trained checkpoints, and two of them are Qwen2 based rather than Qwen-1.5: TinyLLaVA-Qwen2-0.5B-SigLIP and TinyLLaVA-Qwen2.5-3B-SigLIP. Nothing in the visible documentation says which template those two were trained with. Two more details sit in the same list. The open ELM checkpoint, at 0.89B, and both Qwen checkpoints are hosted on individual accounts rather than on the project organization, so three of the five trained models are not under the maintainers' control, while the Phi-2 and Gemma checkpoints are. The vision tower side is narrower than the language side: CLIP, SigLIP, Dino, or CLIP and Dino combined, and a connector restricted to MLP, Qformer or Resampler.

The benchmark table that carries the headline claim is unreadable past its first cell

The takeaways section claims the best model, TinyLLaVA-Phi-2-SigLIP-3.1B, performs better overall than existing 7B models such as LLaVA-1.5 and Qwen-VL. The support for that claim is a table of ten benchmarks, VQA-v2, GQA, SQA-image, TextVQA, MM-Vet, POPE, MME and MMMU-val among them, keyed by vision tower and language model path. In the shipped README that table stops inside its first cell, partway through the text openai/clip-vi, so no score is readable from the documentation itself, and the project publishes no GitHub releases either. The comparison is the project's own claim and should be treated as one until the numbers are reproduced. The same caution applies to the model identifiers: the 3.1B and 2.4B figures in the zoo names are name parts, and the repository documents no parameter count, memory footprint or training cost for any of the five checkpoints.

A public demo is announced with its password, and the news feed stopped in 2025

The announcement feed includes a line dated 2024.05.04 announcing a hosted demo, and it states the access password in the same sentence. The host is a personal tunnel domain rather than a project subdomain, so the demo's availability depends on one machine staying online. The rest of the feed is a useful record of how the codebase moved: a visualization tool for reading model predictions added in August 2024, the factory paper in May 2024, and the framework paper in February 2024. Nothing has been added to it since the entry announcing a video extension hosted in a separate repository in January 2025, yet the default branch received a commit on 2026-09-29. That gap between the last announcement and the last commit is the honest way to describe activity here: work continues, release notes do not exist, and the announcement feed is not a changelog.

Editorial conclusion

Judge TinyLLaVA Factory on the parts it does pin and the parts it leaves in your hands. What it does well is the recipe contract: a global batch size of 256 at a learning rate of 1e-3 for pretraining and 128 at 2e-5 for finetuning, a chat template selected by a conv_version value per model family, and a connector limited to MLP, Qformer or Resampler. What it leaves to the reader is configuration and provenance. Training is driven by editing shell scripts, so there is no config file to diff and no flag to pass, and the finetune template mapping names Qwen-1.5 while the model zoo ships Qwen2 and Qwen2.5 checkpoints. Before committing a cluster to it, verify four things yourself: that the dependency set still resolves against your CUDA and driver, since torch is held at 2.0.1 and flash-attn is installed outside it, that the benchmarks behind the headline claim are reproducible, since the performance table in the README is unreadable past its first cell, that the weights you want actually live under the project organization, since two of five zoo entries sit on individual accounts, and that you can rebuild the environment from scratch, which the installation notes insist on because these requirements differ from the ones the wider LLaVA ecosystem expects.

Frequently asked questions

How do I train a model with TinyLLaVA Factory?

Replace data paths in scripts/train/train_phi.sh, set output_dir in scripts/train/pretrain.sh and scripts/train/finetune.sh, set pretrained_model_path in the finetune script, then adjust GPU ids and per_device_train_batch_size in both scripts and run bash scripts/train/train_phi.sh. There is no separate configuration file.

Which language models, vision towers and connectors does TinyLLaVA Factory support?

Language models cover OpenELM, TinyLlama, StableLM, Qwen, Gemma and Phi. Vision towers cover CLIP, SigLIP, Dino, and CLIP combined with Dino. Connectors are limited to MLP, Qformer and Resampler, and recipes cover frozen, full and partial tuning plus LoRA and QLoRA.

Why does pip install flash-attn separately in TinyLLaVA Factory?

It is pinned to 2.5.7 and installed with --no-build-isolation after the editable install, so the build compiles against the torch already present in the environment, which the dependency set holds at 2.0.1. It is not listed in the package dependencies.

Where are the TinyLLaVA Factory trained checkpoints hosted?

The Phi-2 and Gemma checkpoints sit on the project's Hugging Face organization, while the OpenELM 0.89B checkpoint and both Qwen checkpoints are hosted on individual accounts. None of the five is published as a GitHub release from the repository.

What Python version does TinyLLaVA Factory need?

The packaging metadata requires Python 3.9 or newer, and the installation notes build a conda environment on Python 3.10 with an editable install of the repository. The metadata carries only a generic Python 3 classifier without per version entries.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. TinyLLaVA/TinyLLaVA_Factory on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tinyllava-tinyllava-factory.svg)](https://hysenlabs.com/projects/tinyllava-tinyllava-factory)