# Metarank: an open source Learning-to-Rank service that reranks search results from click events

> Metarank is a Scala-based ranking service that turns click and purchase events into a LambdaMART reranker for search results and recommendations. It is Apache-2.0, stateless with Redis-managed state, and ships a standalone mode that trains and serves in one command.

**metarank/metarank** — A low code Machine Learning personalized ranking service for articles, listings, search results, recommendations that boosts user engagement. A friendly Learn-to-Rank engine

- Repository: https://github.com/metarank/metarank
- Website: https://metarank.ai
- Stars: 2,447 · Forks: 109
- Language: Scala
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/metarank-metarank

## What Metarank solves, and who it is for

Most teams that run search or recommendations already have a working ranker. It might be BM25, a hand-tuned scoring function, or a vendor's default. What they usually lack is a way to fold observed user behaviour back into that ordering. Metarank is aimed squarely at that gap. The README describes it as "a low code Machine Learning personalized ranking service for articles, listings, search results, recommendations". The audience is engineers who own an existing search or recommendation surface and want to rerank its output, not people building a search engine from scratch.

The project frames the value as making an existing system smarter rather than replacing it. Three named capabilities matter here: integrating customer signals like clicks and purchases into the ranking, tracking a visitor profile so results adapt to a session, and using LLMs in bi-encoder and cross-encoder mode to help search interpret query meaning. The README also lists AutoML features: automatic feature generation and model retraining. That combination, event ingestion plus feature extraction plus model training plus a serving API, is what distinguishes Metarank from a plain ranking library.

## How the ranking pipeline actually works

Metarank is a service, not a library you call inline. The README describes it as a stateless cloud-native service with state managed by Redis, which is the architectural decision that shapes everything else. Because the process holds no local state, replicas can be added horizontally and the README states it can process thousands of RPS that way. The Kubernetes deployment guide is the reference for running it in that shape.

Data flows in two directions. Ranking signals arrive from streaming systems, and the README points to a data sources reference listing the supported integrations. Those events, clicks, purchases and similar interactions, feed feature extractors. The README says Metarank computes dozens of typical ranking signals out of the box, naming CTR, referer, User-Agent and time as examples, with a full list in the feature extractors reference. A separate user-session feature extractor maintains a visitor profile for real-time personalization. On the serving side, the trained model reranks a candidate set supplied by your existing search engine. The README claims reranking latency of 10-20ms for large result sets, pointing to a benchmarks page rather than stating the measurement conditions inline. Treat that number as the project's own claim and check the benchmark page for the hardware and dataset behind it.

The model itself is LambdaMART for the personalization path, described in the README as "secondary reranking". Recommendations use two other approaches: trending and similar-items, the latter built on matrix factorization with ALS. So Metarank is really three ranking modes sharing one event and feature pipeline, not a single algorithm.

## Installing Metarank and running a first rerank

The README's quickstart is a three-step Docker walkthrough using the ranklens dataset, the same data behind the hosted demo. Step one downloads the event file. It is gzipped JSONL, one event per line, and the file is large enough that the download takes a moment.

```bash
curl -O -L https://github.com/metarank/metarank/raw/master/src/test/resources/ranklens/events/events.jsonl.gz
```

Step two fetches the matching configuration file. The README notes this config uses an in-memory store, so no external dependency is needed for this first run. That is a deliberate choice for the tutorial: the production path swaps in Redis.

```bash
curl -O -L https://raw.githubusercontent.com/metarank/metarank/master/src/test/resources/ranklens/config.yml
```

Step three starts Metarank in standalone mode, which the README describes as combining training and running the API into one command. The current directory is mounted at /opt/metarank so the container can read both files.

```bash
docker run -i -t -p 8080:8080 -v $(pwd):/opt/metarank metarank/metarank:latest standalone --config /opt/metarank/config.yml --data /opt/metarank/events.jsonl.gz
```

Expect visible progress output while Metarank imports the events and trains the model. Once it settles, the API listens on port 8080 and you send ranking requests to it. The README's text is truncated right at the point where it would show those requests, so the exact request shape is not reproduced here; the quickstart tutorial in the docs is where that continues. Two scripts in the repository root, run_demo.sh and run_quickstart.sh, appear to wrap this flow.

## Where Metarank is the wrong tool

The first constraint is data. Metarank learns from interaction events, and the README's own framing is about integrating clicks and purchases into ranking. If you have no event stream, or your events are not captured in the schema the feature extractors expect, the pipeline has nothing to learn from. There is a cold-start angle here: the project's blog links include a post on solving a search cold-start problem with aggregated CTR, which acknowledges that sparse interaction data is a real obstacle rather than a solved one.

The second constraint is state. The README describes Metarank as stateless with state managed by Redis. That means Redis is not optional in a realistic deployment; the standalone quickstart sidesteps it with an in-memory store, but that store does not survive a restart and does not share state across replicas. Anyone planning horizontal scaling is planning a Redis deployment as well.

The third is scope. Metarank reranks a candidate set. It does not retrieve documents, so it does not replace your search engine, and it does not replace your recommender's candidate generation. If your problem is recall rather than ordering, this is the wrong layer. The README is explicit that the intended use is making an existing search and recommendation system smarter, not building one.

Finally, the README's feature list is not uniformly finished. The semantic neural search entry is marked TODO, and two guide links, the semantic search guide and the collaborative filtering recommendations guide, point at TODO placeholders. The LLM bi-encoder and cross-encoder extractors are documented in the feature extractors reference, but the top-level guides promised in the README are not there yet.

## How Metarank differs from a Learning-to-Rank library

The closest alternative in kind is a Learning-to-Rank plugin embedded in your search engine, for example the OpenSearch LTR plugin. The difference is architectural. A plugin runs inside the search process, scores candidates at query time using features you define and store yourself, and leaves model training, feature computation and event ingestion to you. Metarank inverts that: it runs as a separate service, owns feature extraction and training, and is called after your search engine returns candidates.

That inversion buys the things a plugin cannot easily give you. Session-aware features come from a user-session extractor rather than from code you write. Model retraining and automatic feature generation are part of the product. Model serving supports multiple models for A/B testing, per the configuration reference. And because state lives in Redis rather than in the search node, the ranking tier scales independently of the search tier.

It costs you a network hop and an operational dependency. A plugin adds no latency beyond its own scoring; a separate service adds a call, plus Redis, plus a deployment to monitor. The README's 10-20ms figure is the project's answer to that objection, and it is worth reading the benchmarks page before accepting it for your result set size. The other cost is conceptual: you now have two systems that both need to agree on what a document is and what a candidate set looks like.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-09, which is recent. Release cadence is visible in the tags: 0.7.11 shipped on 2025-06-24, then 0.8.0 on 2026-08-25 and 0.8.1 on 2026-09-07. That is a long gap followed by two releases in two weeks, so the project's own history suggests bursts of activity rather than a steady drumbeat. Version numbers are still in the 0.x range, which in practice means the configuration format and API surface can change between minor releases. Anyone pinning a version should read CHANGELOG.md before upgrading rather than assuming backward compatibility.

Upgrade cost is dominated by the config file. The quickstart config is a YAML document describing events, feature extractors and models, and the repository carries a .scala-steward.conf, which indicates automated dependency update pull requests are part of the maintenance routine. If you pin the Docker image tag rather than using latest, upgrades are a deliberate act: change the tag, re-read the changelog, restart.

The licence is Apache-2.0, which permits commercial use, modification and redistribution, and includes a patent grant. It also requires that you preserve copyright and licence notices and state significant changes. This is a summary of the licence text, not legal advice; the LICENSE file in the repository is the authoritative document, and organisations with strict review processes should route it through their own counsel. Nothing in the README suggests an open-core split or a feature gated behind a commercial tier.

## Conclusion

Adopt Metarank if you already run a search engine or recommender and have a stream of click or purchase events to learn from. Do not adopt it if you have no event data, no Redis, or no appetite for a JVM service in the request path. Verify first that your events match the expected schema, because the whole pipeline is built around them.

## FAQ

### What is Metarank?

Metarank is an open-source ranking service, written in Scala and licensed under Apache-2.0, that reranks search results and recommendations using machine learning. The README describes it as a low code personalized ranking service and a Learn-to-Rank engine.

### What is a meta ranking system?

The README does not define the term meta ranking. It describes Metarank as a ranking service that reranks a candidate set produced by an existing search engine or recommender, rather than retrieving candidates itself.

### What is RankNet?

RankNet is not mentioned in the README or the repository material. Metarank's personalization path uses LambdaMART for secondary reranking, and its similar-items recommendations use matrix factorization with ALS.

### What is an ML ranking model?

The README does not give a general definition. In Metarank's case, the ranking model is trained from imported interaction events and then served by the API, with LambdaMART used for the reranking path and ALS for similar-items recommendations.

## Sources

- [License: Apache-2.0](https://github.com/metarank/metarank/blob/master/LICENSE)
- [metarank/metarank on GitHub](https://github.com/metarank/metarank)
- [Project website](https://metarank.ai)
- [README](https://github.com/metarank/metarank/blob/master/README.md)
- [Releases](https://github.com/metarank/metarank/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/metarank-metarank
