# FlagEmbedding: The BGE Toolkit for Retrieval and RAG

> FlagEmbedding is a Python library from Beijing Academy of Artificial Intelligence (BAAI) that packages the BGE series of embedding, reranking, and retrieval models. It covers a wide range of retrieval scenarios from dense text embedding to multimodal image search, with finetuning support built in.

**FlagOpen/FlagEmbedding** — Retrieval and Retrieval-augmented LLMs

- Repository: https://github.com/FlagOpen/FlagEmbedding
- Website: http://www.bge-model.com/
- Stars: 12,171 · Forks: 917
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/flagopen-flagembedding

## What FlagEmbedding Provides and Who Uses It

FlagEmbedding is the official Python package for BAAI's BGE (Beijing General Embedding) family of models. It targets ML engineers and researchers who need production-grade text embeddings, rerankers, or retrieval models for building search systems and retrieval-augmented generation pipelines.

The README positions it as a one-stop retrieval toolkit: it covers the full pipeline from encoding queries and documents into embeddings, through reranking candidates with a cross-encoder, to finetuning those models on domain-specific data. The package version at the time of the last release was 1.4.2, as recorded in setup.py. BAAI maintains a dedicated documentation site at www.bge-model.com for centralized information about the BGE model family.

FlagEmbedding is MIT-licensed. The setup.py specifies that Python 3.8 or later is required. Core dependencies include PyTorch, HuggingFace Transformers (version 4.44.2 or later but below 6.0.0), HuggingFace Datasets, Accelerate, sentence-transformers, and PEFT for parameter-efficient finetuning.

## BGE-M3 and Its Three Retrieval Modes

BGE-M3 is the flagship model in the BGE family. The README describes M3 as standing for Multi-linguality, Multi-granularities, and Multi-Functionality. It supports over 100 languages with input lengths up to 8192 tokens, which is substantially longer than many embedding models.

The multi-functionality aspect is the distinctive technical characteristic: BGE-M3 supports three retrieval methods in a single model. Dense retrieval produces a fixed-size vector for each text, suitable for approximate nearest neighbor search with a standard vector database. Lexical retrieval operates like a sparse model, assigning weights to specific terms, similar to how BM25 operates but learned rather than rule-based. Multi-vector retrieval, referred to as ColBERT-style in the README, stores token-level embeddings and computes late interaction scores.

The README states BGE-M3 achieved state-of-the-art results on multi-lingual (MIRACL) and cross-lingual (MKQA) benchmarks at the time of release. Being the first model to unify all three retrieval methods in one checkpoint is the architectural claim made in the README.

## Installing FlagEmbedding and Understanding the Finetuning Extras

The package installs from PyPI using the name declared in setup.py:

```bash
pip install FlagEmbedding
```

This installs the core package with all standard dependencies including torch, transformers, datasets, accelerate, sentence-transformers, peft, ir-datasets, sentencepiece, and protobuf. The transformers version constraint is >=4.44.2,<6.0.0, which may conflict with environments already pinned to a different version of that library.

Finetuning support requires additional dependencies not included in the base install. The setup.py defines a finetune extras group containing deepspeed and flash-attn. Both of these packages have system-level build requirements: flash-attn requires a compatible CUDA version and GPU, while deepspeed has its own compiler dependencies.

A tutorials directory was added to the repository in September 2024 and the README notes it will be continuously updated. The examples/ directory contains subdirectories for evaluation, finetuning, and inference use cases.

## BGE-VL and the Multimodal Embedding Extension

The March 2025 news entry in the README announces BGE-VL, described as state-of-the-art multimodal embedding models for visual search. BGE-VL supports text-to-image, image-to-text, image-and-prompt-to-image, and text-to-image-and-text search tasks. The README states these models are released under the MIT licence and are free for both academic and commercial use.

BGE-VL is based on the MegaPairs dataset, a large-scale synthetic dataset for multimodal retrieval. The MegaPairs repository and paper are separate from FlagEmbedding but linked from the README.

This multimodal extension moves FlagEmbedding beyond pure text retrieval into visual search scenarios. For teams that only need text embeddings, BGE-VL is an optional component and does not affect the base installation.

## Version Pinning and Dependency Conflicts to Anticipate

FlagEmbedding's setup.py pins several dependencies at specific ranges. The transformers constraint (>=4.44.2,<6.0.0) means the library will not install alongside any package that requires an older transformers version. The datasets pin is >=2.19.0, and accelerate is >=0.20.1. These ranges are not overly restrictive for a new environment, but they can cause resolver conflicts in an existing environment that was set up with earlier versions of these packages.

The peft dependency, which provides LoRA and other parameter-efficient finetuning adapters, is an unconditional install dependency rather than an extras dependency, even though many users may not need finetuning. This adds to the total install footprint.

The most recent tagged release is v1.4.2, dated 2026-08-24. The last commit to the repository was also on 2026-08-24. The repository uses GitHub tags for release tracking, and three releases in the v1.4.x series were published within a two-day window in August 2026.

## FlagEmbedding vs Sentence-Transformers for Embedding Use Cases

sentence-transformers is the most widely used Python library for computing sentence embeddings. It provides a straightforward interface for encoding text into fixed-size vectors using any model compatible with its framework. Many BGE models are also available through sentence-transformers, meaning the distinction between the two libraries is primarily about what else each one provides.

FlagEmbedding bundles finetuning support, reranker training, the multi-vector ColBERT-style retrieval mode from BGE-M3, and the BGE-VL multimodal models. sentence-transformers is a general-purpose library that works with many model families; it does not include BGE-specific finetuning pipelines or the dense-plus-sparse retrieval setup that BGE-M3 offers.

For a team that only needs to encode sentences into vectors for cosine similarity search, sentence-transformers is simpler and has fewer dependencies. For a team that wants to finetune a BGE model on their own data, use BGE-M3's multi-retrieval modes, or add reranking to their pipeline, FlagEmbedding is the more complete option.

## Conclusion

Python developers building retrieval pipelines or RAG systems who need multilingual embeddings, cross-lingual search, or dense-plus-sparse retrieval in a single model should evaluate BGE-M3 through FlagEmbedding. Teams that need image-text retrieval can add BGE-VL. Before switching an existing sentence-transformers pipeline, verify compatibility with the transformers>=4.44.2,<6.0.0 version constraint in FlagEmbedding's setup.py, since that pin may conflict with other library versions in the same environment. The most recent release was v1.4.2, published on 2026-08-24.

## FAQ

### How do you install FlagEmbedding?

Run pip install FlagEmbedding in a Python 3.8 or later environment with PyTorch already installed. The base install includes transformers, datasets, accelerate, sentence-transformers, peft, and several other packages. Finetuning requires additional dependencies (deepspeed and flash-attn) not included in the base install.

### What is FlagEmbedding?

FlagEmbedding is a Python library from BAAI that packages the BGE series of text embedding and reranking models. It includes BGE-M3 for multilingual dense, lexical, and multi-vector retrieval, rerankers built on large language models, BGE-VL for multimodal image-text search, and tools for finetuning these models on custom data.

### How does FlagEmbedding compare to sentence-transformers?

sentence-transformers is a general-purpose encoding library that works with many model families. FlagEmbedding is specific to the BGE model family and adds finetuning pipelines, BGE-M3's multi-vector and lexical retrieval modes, and reranker support that sentence-transformers does not provide out of the box.

## Sources

- [FlagOpen/FlagEmbedding on GitHub](https://github.com/FlagOpen/FlagEmbedding)
- [License: MIT](https://github.com/FlagOpen/FlagEmbedding/blob/master/LICENSE)
- [Project website](http://www.bge-model.com/)
- [README](https://github.com/FlagOpen/FlagEmbedding/blob/master/README.md)
- [Releases](https://github.com/FlagOpen/FlagEmbedding/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/flagopen-flagembedding
