# AnglE: Training Sentence Embeddings with an Angle Loss

> AnglE is a Python library for training and running BERT and LLM based sentence embeddings. Its angle-optimized loss produced models that reached the top of the MTEB leaderboard, but the README is thin on training mechanics and the package pins a specific transformers version.

**SeanLee97/AnglE** — Train and Infer Powerful Sentence Embeddings with AnglE | 🔥 SOTA on STS and MTEB Leaderboard

- Repository: https://github.com/SeanLee97/AnglE
- Website: https://arxiv.org/abs/2309.12871
- Stars: 575 · Forks: 37
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/seanlee97-angle

## What AnglE solves and who it is for

Sentence embeddings turn a piece of text into a vector so that similar meanings sit close together. The hard part is training: a model that scores well on one similarity benchmark often collapses on another, and the loss function is usually where that goes wrong. AnglE is a Python library built around the angle-optimized loss from the paper AnglE: Angle-optimized Text Embeddings, accepted at ACL 2024. The README describes it as allowing you to train state-of-the-art BERT or LLM based sentence embeddings with a few lines of code.

The audience is narrow and identifiable. You are a machine learning engineer who has a labelled or weakly labelled sentence-pair dataset and wants a domain embedding model, for example medical similarity or code similarity. The README lists WhereIsAI/pubmed-angle-base-en and WhereIsAI/UAE-Code-Large-V1 as models trained in exactly that way. If you only need to call an embedding API, this library is more machinery than you need. If you want to control the training objective, it is aimed at you.

## The angle loss and the training and inference split

The library has two jobs and keeps them separate. Inference is a wrapper around Hugging Face transformers: AnglE.from_pretrained loads a checkpoint, .encode turns text into vectors, and pooling_strategy decides which hidden state becomes the sentence vector. The README shows cls for BERT models and last for LLM models, which matters because picking the wrong one silently degrades every similarity score you compute.

Training is where the angle loss lives. The README lists four losses: AnglE loss, contrastive loss, CoSENT loss, and Espresso loss, the last described as ICLR 2025 work also called 2DMSE. Backbones span BERT, RoBERTa and ModernBERT on one side and LLaMA, Mistral and Qwen on the other. LLM training uses LoRA weights through peft, which is why pyproject.toml pins peft==0.18.1. The repository layout shows examples/Training/, examples/FSDP/, examples/NLI/ and examples/UAE/, so the intended path is to copy an example config rather than start from a blank script.

Multi-GPU training is listed as a feature, and examples/multigpu_infer.py exists for the inference side. The README does not document how the distributed launch is wired, so treat the example directory as the real specification.

## Installing angle-emb and encoding your first sentences

The README gives two install paths, uv and pip. Both install the same PyPI package, angle-emb, and both use the -U flag to upgrade an existing install.

```bash
uv pip install -U angle-emb
```

The pip equivalent is:

```bash
pip install -U angle-emb
```

Python 3.10 or newer is required according to pyproject.toml. Once installed, the shortest useful program loads a BERT checkpoint, encodes a query with a prompt and documents without one, then compares them. Prompts.C is a retrieval prompt and the README notes that prompts use {text} as a placeholder.

```python
from angle_emb import AnglE, Prompts
from angle_emb.utils import cosine_similarity

angle = AnglE.from_pretrained('WhereIsAI/UAE-Large-V1', pooling_strategy='cls').cuda()
qv = angle.encode(['what is the weather?'], to_numpy=True, prompt=Prompts.C)
doc_vecs = angle.encode([
    'The weather is great!',
    'it is rainy today.',
    'i am going to bed'
], to_numpy=True)
for dv in doc_vecs:
    print(cosine_similarity(qv[0], dv))
```

You should see three numbers, with the first document scoring higher than the third. For LLM checkpoints the call changes: you pass pretrained_lora_path, pooling_strategy='last', is_llm=True and torch_dtype=torch.float16, and the README states that is_llm must always be set for LLM models. The package also installs a console script, angle-trainer, mapped to angle_emb.angle_trainer:main, which is the entry point for training runs.

## Where AnglE gets in your way

The dependency pins are aggressive. pyproject.toml requires transformers==5.3.0, peft==0.18.1 and bitsandbytes==0.49.2 as exact versions. If another library in your stack needs a different transformers release, you are resolving that conflict by hand. That is the cost of tracking fast-moving training code, and it is worth knowing before you build a pipeline around it.

The README is also uneven. It covers inference thoroughly, with Colab notebooks for both BERT and LLM paths, but the training section is a feature list rather than a walkthrough. There is no documented data format for a training run, no description of what angle-trainer expects on the command line, and no stated rollback or checkpoint-resume procedure. A MIGRATION_GUIDE.md exists at the repository root, which suggests the API has moved between versions, but the README does not summarise what changed.

Finally, this is the wrong tool if you have no labelled data and no appetite to create any. The angle loss is a training objective; it does not remove the need for sentence pairs. If your only goal is to embed documents for a retrieval index, a frozen checkpoint plus a vector database gets you there without this library in the loop.

## How AnglE differs from sentence-transformers

The natural comparison is sentence-transformers, the widely used library for the same problem. The difference is in the objective, not the interface. sentence-transformers defaults to a contrastive or triplet setup and its training recipes are built around that family of losses. AnglE's contribution is the angle-optimized loss, which the paper frames as addressing the optimisation difficulty of cosine-based objectives.

In practice the two libraries overlap heavily on inference: both load Hugging Face checkpoints and both return vectors. The divergence appears when you train. AnglE ships Espresso loss and CoSENT loss alongside its own, and it explicitly targets LLM backbones with LoRA, which is a heavier setup than a typical sentence-transformers fine-tune. The README also points to BiLLM for bi-directional LLM backbones, a separate project. If your team already has sentence-transformers training scripts that work, switching buys you a different loss and a new set of version pins. The reason to switch is a measurable gain on your own evaluation set, not the leaderboard history.

## Licence and maintenance cost

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice is kept. Note the mismatch: pyproject.toml carries the classifier License :: OSI Approved :: BSD License while the repository ships a LICENSE file and the project metadata here says MIT. Confirm which text actually applies before you rely on either, and route the question to your legal team rather than treating this paragraph as an answer.

The last push to the default branch was on 2026-03-22. The most recent tagged release in the repository is v0.6.0 from 2025-10-19, while pyproject.toml declares version 0.6.1, so there is unreleased work on main. The repository is not archived. Plan for upgrades by reading MIGRATION_GUIDE.md, since the presence of that file implies breaking changes have happened before. The exact pins mean an upgrade is rarely a one-line version bump.

## Conclusion

Adopt AnglE if you need to train or fine-tune sentence embeddings and want a loss that the paper reports as state of the art on STS, or if you want to run the official UAE and angle-llama checkpoints through one API. Do not adopt it if you only need a frozen off-the-shelf embedding model and never intend to train, since sentence-transformers covers that ground with a larger ecosystem. Before committing, verify that your Python version satisfies the >=3.10 requirement in pyproject.toml, that you can live with the pinned transformers==5.3.0, and that the pooling_strategy you pass matches the checkpoint you load.

## FAQ

### How do I install AnglE?

Install the angle-emb package from PyPI, either with uv pip install -U angle-emb or pip install -U angle-emb as the README shows. Python 3.10 or newer is required according to pyproject.toml.

### What is AnglE used for?

It trains and runs sentence embeddings, turning text into vectors for similarity, retrieval and RAG use cases. The README describes it as both a training library for the angle-optimized loss and a general inference framework for transformer-based embeddings.

### Which pooling strategy should I use with AnglE?

The README uses pooling_strategy='cls' for BERT-based checkpoints such as WhereIsAI/UAE-Large-V1 and pooling_strategy='last' for LLM-based models. Matching the strategy to the checkpoint is what keeps the similarity scores meaningful.

### Does AnglE support LLM backbones like LLaMA?

Yes. The README lists LLaMA, Mistral and Qwen among supported backbones and shows loading a LoRA weight with pretrained_lora_path, is_llm=True and torch_dtype=torch.float16. The README states is_llm must always be set for LLM models.

### Is AnglE free to use commercially?

The repository is MIT licensed, which permits commercial use if the copyright notice is retained. pyproject.toml lists a BSD classifier instead, so check the LICENSE file and confirm with your legal team.

## Sources

- [License: MIT](https://github.com/SeanLee97/AnglE/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2309.12871)
- [README](https://github.com/SeanLee97/AnglE/blob/main/README.md)
- [Releases](https://github.com/SeanLee97/AnglE/releases)
- [SeanLee97/AnglE on GitHub](https://github.com/SeanLee97/AnglE)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/seanlee97-angle
