NVlabs/Eagle: a vision-language model family you adopt as a codebase, not a package
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
At a glance
- What is it?
- Eagle is NVIDIA's Apache-2.0 research repository holding four generations of vision-language models, from mixture-of-encoders Eagle to the LocateAnything grounding model. It is for teams that want to fine-tune or reproduce, not for teams that want a pip install.
- Who is it for?
- Adopt Eagle if you have GPU capacity, a Python environment you control, and a task that matches one of the four models: grounding and detection with LocateAnything, long-context image and video understanding with Eagle 2.5, or image understanding with Eagle 2.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 97 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Eagle actually ships, and why it is four projects under one name
The repository is not one model. It is a directory tree with three top-level model folders, and the README presents them as a family. Eagle, in the Eagle/ directory, explores mixture-of-encoders designs for vision-centric VLMs. Eagle 2, documented inside the same directory, moves to post-training data strategies for image understanding. Eagle 2.5, in Eagle2_5/, adds long-context multimodal understanding over images and video. LocateAnything, in Embodied/, is a generalist grounding model built on Eagle that handles detection, pointing, GUI grounding, OCR and document understanding.
The README also lists downstream adoptions: Eagle 2 and 2.5 served as the VLM backbone for NVIDIA Isaac GR00T N1, N1.5 and N1.6, and a native-resolution Eagle variant is the backbone of GR00T-N1.6. That matters for a reader deciding whether to build on it. The code you are reading is the same lineage that NVIDIA used in its own robotics stack, which is a stronger signal than any benchmark table in the README.
Who this is for: a team with GPUs that needs a vision-language backbone it can fine-tune, or a research group reproducing a data-centric training recipe. Who it is not for: anyone who wants a hosted API or a single importable library. There is no pip package named eagle in this repository, and the README never claims one.
How the code is organised, and where the entry points live
The layout tells you most of the workflow. Each model generation keeps its own folder with its own README, its own shell scripts under a shell/ subdirectory, and its own documentation. The Embodied/ directory holds LocateAnything, and the release notes point to Embodied/shell/locate-anything-lora-visual-prompt.sh for LoRA fine-tuning with visual prompts. Eagle2_5/ keeps its onboarding material under Eagle2_5/document/, with 0.onboarding.md as the first file.
The data flow is the standard one for a VLM fine-tune: a pretrained checkpoint from Hugging Face, a dataset in whatever format the shell scripts expect, a training entry point invoked through a shell script, and an inference path. LocateAnything adds a runtime detail worth noting. The June 2026 release note says batch inference runs on a pure FlashAttention runtime and is aimed at A100, RTX 4090 and other non-Hopper, non-Blackwell GPUs. That is a deliberate constraint: the fast path does not require the newest hardware generation.
The decoding mechanism is the most concrete design choice in the README. LocateAnything uses Parallel Box Decoding, described as predicting each bounding box atomically in a single forward pass, contrasted with quantized coordinate decoding. The README frames this as a throughput difference. It does not publish numbers in the text, so treat the speed claim as a design description rather than a measured result.
Installing Eagle: clone the repository, then follow the per-model onboarding
There is no install command in the top-level README. The README points to three getting-started documents instead, one per model line, and the actual setup steps live in those files. Start by cloning the repository and reading the onboarding document for the model you want.
git clone https://github.com/NVlabs/EAGLE.git
cd EAGLE
cat Eagle2_5/document/0.onboarding.mdFor LocateAnything, the equivalent entry point is Embodied/README.md, and the fine-tuning example is a shell script in the same tree.
cat Embodied/README.md
ls Embodied/shell/What you should see is a directory listing that includes locate-anything-lora-visual-prompt.sh, the LoRA visual-prompt fine-tuning script named in the release notes. The script is the concrete artifact to read before you touch a config, because it carries the argument names and paths the maintainers actually use.
Weights are separate. The README links a Hugging Face collection at huggingface.co/collections/nvidia/eagle, and the Eagle 2.5 model is published as nvidia/Eagle2.5-8B. There is also a Hugging Face Space for LocateAnything at huggingface.co/spaces/nvidia/LocateAnything if you want to see the model behave before installing anything. The repository does not document a container image, a conda environment file, or a pinned dependency list in the files it ships here.
The licence split is the first thing to check, not the last
The code licence and the model licence are different, and the README badges say so explicitly. The code badge reads Apache 2.0 and links to the LICENSE file at the repository root. The model badge reads NVIDIA License and links to Eagle2_5/LICENSE_MODEL.
That split is common for released weights, but it changes what you can do. Apache-2.0 covers the training and inference code in this repository. The checkpoints you download from the Hugging Face collection are governed by the NVIDIA model licence, and the repository places that text under Eagle2_5/. Whether your intended use, including commercial deployment or redistribution of a fine-tuned derivative, is permitted depends on that file and on the licence attached to the specific checkpoint you pull. Read it for the model you actually plan to use rather than assuming one file covers the whole family. Nothing here is legal advice, and the repository does not summarise the model licence terms in the README.
Where Eagle is the wrong tool
The repository is a research codebase, and the documentation is organised accordingly. The top-level README is a family overview with a resource index and a release timeline. The operational detail is pushed down into per-model documents, and those documents are not uniform: Eagle 2.5 has a numbered onboarding file, LocateAnything has a directory README, and Eagle 2 is documented inside the Eagle/ directory alongside the older mixture-of-encoders work. If you need one consistent installation procedure across all four models, this repository does not provide it.
The second limitation is hardware. The LocateAnything release note describes its fast batch inference path as targeting A100, RTX 4090 and other non-Hopper, non-Blackwell GPUs. That is an explicit statement about which accelerators the pure FlashAttention runtime was built for. If your fleet is something else, or if you were counting on a TensorRT path, note that the README mentions Torch-TRT support for Eagle 2 only, in a September 2025 update, and does not extend that claim to the later models.
The third case is the simplest. If your task is text-only, or if you need a managed endpoint with an SLA, none of the four models here is the right fit. The family exists to explore data-centric strategies for multimodal understanding, and the repository is built for people who will run the training and inference code themselves.
What to compare against: a served multimodal API
The realistic alternative for many teams is not another open repository but a hosted multimodal API from a commercial provider. The difference is structural rather than a matter of model quality. With a hosted API you send an image or a video and receive text; you never see the checkpoint, you cannot fine-tune on your own data through this repository's scripts, and your latency and cost scale with someone else's infrastructure. With Eagle you download weights, run the shell scripts in Embodied/shell/ or the Eagle2_5 document's instructions on your own GPUs, and own the fine-tuning loop.
That trade runs in both directions. Eagle gives you the training recipe and the ability to adapt the model to a narrow domain such as GUI grounding or dense detection. It also gives you the operational work: provisioning GPUs, managing checkpoints, and reading per-model documentation that is not written as a product manual. A hosted API gives you none of the adaptation and all of the operational simplicity.
If your comparison set is other open VLM repositories, the distinguishing feature here is the data-centric framing. The README describes the family as exploring data strategies across multimodal understanding, long-context reasoning and embodied applications, and the reports linked from the header (Eagle.pdf, Eagle2.pdf, Eagle2.5.pdf, and the LocateAnything report) are where that argument is made. The repository is the code that goes with those reports.
Editorial conclusion
Adopt Eagle if you have GPU capacity, a Python environment you control, and a task that matches one of the four models: grounding and detection with LocateAnything, long-context image and video understanding with Eagle 2.5, or image understanding with Eagle 2. Before you commit, check the model licence file for the checkpoint you intend to use, since the Apache-2.0 code licence does not cover the weights, and read the onboarding document for the specific model rather than the top-level README, which is a directory of links.
Frequently asked questions
Is there a pip package to install NVlabs/Eagle?
The repository does not document one. The README points to per-model getting-started documents, and the workflow shown in the release notes is cloning the repository and running shell scripts such as Embodied/shell/locate-anything-lora-visual-prompt.sh.
Which licence applies to the Eagle model weights?
The README shows a code licence of Apache 2.0 and a separate model licence of NVIDIA License, linked to Eagle2_5/LICENSE_MODEL. The weights are published through the Hugging Face collection at huggingface.co/collections/nvidia/eagle, so the model licence governs the checkpoints rather than the Apache-2.0 code licence.
What hardware does LocateAnything batch inference target?
The June 2026 release note states that batch inference uses a pure FlashAttention runtime and is efficient on A100, RTX 4090 and other non-Hopper, non-Blackwell GPUs.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-eagle)