# roboflow/maestro: a fine-tuning harness for Florence-2, PaliGemma 2 and Qwen2.5-VL

> maestro wraps data loading, configuration and the training loop for three vision-language models behind one CLI and one Python entry point. The recipes are narrow, the extras are model-specific, and the project itself is labelled Alpha.

**roboflow/maestro** — streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL

- Repository: https://github.com/roboflow/maestro
- Website: https://maestro.roboflow.com
- Stars: 2,699 · Forks: 224
- Language: Python
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/roboflow-maestro

## What maestro actually takes off your plate

Fine-tuning a vision-language model is mostly plumbing. You need a dataset loader that understands image-text pairs, a configuration object that survives a restart, a training loop wired to a distributed backend, and a metric that means something for structured output. maestro's stated job is to encapsulate those best practices so you write a config dictionary instead of a training script. The README lists configuration, data loading, reproducibility and training loop setup as the four things it handles.

The intended user is an engineer or researcher who has already picked a model and now wants to adapt it to a specific task. The README's own examples are concrete: object detection with Florence-2, JSON data extraction with PaliGemma 2 and Qwen2.5-VL. That is a narrower audience than a general training framework serves, and deliberately so. If your task is not one of captioning, detection, VQA or structured extraction, the recipes give you less to start from.

## One CLI, one JSONL format, three model cores

The architecture visible in the repository is a set of per-model core modules under maestro/trainer/models, each owning its own configuration and training routine. The CLI and the Python API are thin shells over those cores. That is why the install instructions are model-specific: `pip install "maestro[paligemma_2]"` pulls the dependency set for PaliGemma 2 rather than a universal one, and the README explicitly recommends creating a dedicated Python environment for each model because some of them have clashing requirements.

Data flows in one shape. The 1.0.0 release notes describe a consistent JSONL format for data handling, and the core modules take care of data preparation from that format. The training loop itself is built on Lightning, pinned at 2.6.1 in pyproject.toml, with torch and accelerate listed as keywords and evaluate, nltk and sacrebleu as dependencies for metrics. The 1.0.0 notes also list LoRA, QLoRA and graph freezing as the memory-reduction options, which is what makes the smaller GPUs in the free Colab cookbooks plausible.

The consequence of this design is that the CLI is a façade, not an abstraction layer. `maestro paligemma_2 train` and `maestro qwen2_5_vl train` are separate code paths that happen to share a command structure. A bug in one model's core is not fixed by a change to the shared CLI.

## Installing maestro and running a first fine-tune

The README's install step is a single pip command with a model extra. Because the extras pull different dependency sets, install into a fresh virtual environment rather than into whatever you already have, which is what the README recommends.

```bash
pip install "maestro[paligemma_2]"
```

Once installed, the CLI takes the dataset path, epoch count, batch size, optimization strategy and metric as flags. The README gives this exact invocation for PaliGemma 2 with QLoRA and edit distance as the metric.

```bash
maestro paligemma_2 train \
  --dataset "dataset/location" \
  --epochs 10 \
  --batch-size 4 \
  --optimization_strategy "qlora" \
  --metrics "edit_distance"
```

For more control, the Python API exposes the same parameters as a dictionary passed to the model's train function. The README shows this for PaliGemma 2.

```python
from maestro.trainer.models.paligemma_2.core import train

config = {
    "dataset": "dataset/location",
    "epochs": 10,
    "batch_size": 4,
    "optimization_strategy": "qlora",
    "metrics": ["edit_distance"]
}

train(config)
```

Note the parameter names differ between the two surfaces: the CLI uses `--batch-size` with a hyphen, the Python dictionary uses `batch_size` with an underscore, and `--metrics` takes a string on the command line but a list in Python. If you want to see the whole thing run before touching your own data, the README links four Colab notebooks covering Florence-2 object detection with LoRA, PaliGemma 2 JSON extraction with LoRA, Qwen2.5-VL 3B JSON extraction with QLoRA, and Qwen2.5-VL 7B object detection with QLoRA. The Florence-2 and Qwen2.5-VL 7B detection recipes are marked experimental in the README's table.

## Where the recipes stop being enough

The packaging is the first real constraint. pyproject.toml declares `requires-python = ">=3.9,<3.13"`, so Python 3.13 is out. The dependency list pins `lightning==2.6.1` exactly and caps `requests` at `<=2.32.5`, which means maestro has to be installed into an environment you are willing to shape around it, not the other way round. Anyone maintaining a shared environment with other Lightning-based projects should expect to resolve that pin by hand.

The second constraint is scope. There is no generic model registration path described in the README. The documented surface is Florence-2, PaliGemma 2 and Qwen2.5-VL, and the repository topics mention phi-3-vision and qwen2-vl, but the README's recipes and the release notes stop at the three named models. If your model is not one of them, maestro gives you a training scaffold you would have to fork and extend, at which point you are maintaining a fork of an Alpha-stage project.

The third is maturity signalling. pyproject.toml carries the classifier `Development Status :: 3 - Alpha`, and the only release listed is 1.0.0 from 2025-02-05, while the package version in pyproject.toml is 1.1.0rc3. The repository's default branch is develop, not main. The last push was on 2026-09-14. None of that is disqualifying for a research tool, but it does mean you should read the core module for your model before trusting the defaults, because the README does not document what each training run writes to disk, how to resume an interrupted run, or how to export the adapter afterwards.

## How maestro differs from writing the loop yourself with transformers and PEFT

The obvious alternative is the stack underneath: Hugging Face transformers for the model, PEFT for LoRA or QLoRA, and your own script for the loop. That combination supports far more architectures than three, and it is what most teams reach for when the model is unusual.

The difference in approach is where the code lives. With transformers and PEFT you assemble the pieces per project: you write the dataset class, you pick the collator, you decide how the metric is computed, and you own the resulting script. maestro inverts that. The dataset format is fixed at JSONL, the metric is a flag value like `edit_distance`, and the loop is Lightning's. You trade flexibility for not writing the plumbing, and you inherit maestro's dependency pins along with it.

That trade is worth it when your task matches a recipe and your data can be expressed as JSONL. It is the wrong trade when you need a custom loss, a non-standard evaluation, or a model outside the three. It is also the wrong tool if you need a stable API: the version in pyproject.toml is a release candidate, and the package classifier says Alpha, so pinning to a specific version and reading its core module is the only way to know what you are building on.

## Conclusion

Adopt maestro if you already work with Florence-2, PaliGemma 2 or Qwen2.5-VL and want LoRA, QLoRA or graph freezing without writing your own training loop, and if you are comfortable with a package whose own metadata says Development Status 3 - Alpha. Do not adopt it as a general fine-tuning framework for other vision-language models: the recipes are per-model extras, and the README advises a dedicated Python environment for each because their requirements can clash. Verify three things before committing: that the model you want has an extra named in the README, that the dataset you have matches the JSONL format the core modules expect, and that your hardware fits the strategy you pick, since the four published cookbooks are the only worked examples of that pairing.

## FAQ

### How do I install maestro?

Install the extra for the model you want, for example `pip install "maestro[paligemma_2]"`. The README recommends a dedicated Python environment per model because some of the models have clashing requirements.

### How do I use maestro to fine-tune a model?

Use the CLI with the model name and the train subcommand, passing the dataset path, epochs, batch size, optimization strategy and metrics, or call the model's train function from Python with the same values in a config dictionary. The README shows both forms for PaliGemma 2.

### How do I install maestro on Windows?

The README does not give platform-specific install steps. The package metadata lists Windows, Linux and macOS as supported operating systems, and installation is the same pip command with a model extra.

## Sources

- [License: Apache-2.0](https://github.com/roboflow/maestro/blob/develop/LICENSE)
- [Project website](https://maestro.roboflow.com)
- [README](https://github.com/roboflow/maestro/blob/develop/README.md)
- [Releases](https://github.com/roboflow/maestro/releases)
- [roboflow/maestro on GitHub](https://github.com/roboflow/maestro)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/roboflow-maestro
