Metarank: an open source Learning-to-Rank service that reranks search results from click events
A low code Machine Learning personalized ranking service for articles, listings, search results, recommendations that boosts user engagement. A friendly Learn-to-Rank engine
At a glance
- What is it?
- Metarank is a Scala-based ranking service that turns click and purchase events into a LambdaMART reranker for search results and recommendations. It is Apache-2.0, stateless with Redis-managed state, and ships a standalone mode that trains and serves in one command.
- Who is it for?
- Adopt Metarank if you already run a search engine or recommender and have a stream of click or purchase events to learn from. Do not adopt it if you have no event data, no Redis, or no appetite for a JVM service in the request path.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Scala, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Metarank solves, and who it is for
Most teams that run search or recommendations already have a working ranker. It might be BM25, a hand-tuned scoring function, or a vendor's default. What they usually lack is a way to fold observed user behaviour back into that ordering. Metarank is aimed squarely at that gap. The README describes it as "a low code Machine Learning personalized ranking service for articles, listings, search results, recommendations". The audience is engineers who own an existing search or recommendation surface and want to rerank its output, not people building a search engine from scratch.
The project frames the value as making an existing system smarter rather than replacing it. Three named capabilities matter here: integrating customer signals like clicks and purchases into the ranking, tracking a visitor profile so results adapt to a session, and using LLMs in bi-encoder and cross-encoder mode to help search interpret query meaning. The README also lists AutoML features: automatic feature generation and model retraining. That combination, event ingestion plus feature extraction plus model training plus a serving API, is what distinguishes Metarank from a plain ranking library.
How the ranking pipeline actually works
Metarank is a service, not a library you call inline. The README describes it as a stateless cloud-native service with state managed by Redis, which is the architectural decision that shapes everything else. Because the process holds no local state, replicas can be added horizontally and the README states it can process thousands of RPS that way. The Kubernetes deployment guide is the reference for running it in that shape.
Data flows in two directions. Ranking signals arrive from streaming systems, and the README points to a data sources reference listing the supported integrations. Those events, clicks, purchases and similar interactions, feed feature extractors. The README says Metarank computes dozens of typical ranking signals out of the box, naming CTR, referer, User-Agent and time as examples, with a full list in the feature extractors reference. A separate user-session feature extractor maintains a visitor profile for real-time personalization. On the serving side, the trained model reranks a candidate set supplied by your existing search engine. The README claims reranking latency of 10-20ms for large result sets, pointing to a benchmarks page rather than stating the measurement conditions inline. Treat that number as the project's own claim and check the benchmark page for the hardware and dataset behind it.
The model itself is LambdaMART for the personalization path, described in the README as "secondary reranking". Recommendations use two other approaches: trending and similar-items, the latter built on matrix factorization with ALS. So Metarank is really three ranking modes sharing one event and feature pipeline, not a single algorithm.
Installing Metarank and running a first rerank
The README's quickstart is a three-step Docker walkthrough using the ranklens dataset, the same data behind the hosted demo. Step one downloads the event file. It is gzipped JSONL, one event per line, and the file is large enough that the download takes a moment.
curl -O -L https://github.com/metarank/metarank/raw/master/src/test/resources/ranklens/events/events.jsonl.gzStep two fetches the matching configuration file. The README notes this config uses an in-memory store, so no external dependency is needed for this first run. That is a deliberate choice for the tutorial: the production path swaps in Redis.
curl -O -L https://raw.githubusercontent.com/metarank/metarank/master/src/test/resources/ranklens/config.ymlStep three starts Metarank in standalone mode, which the README describes as combining training and running the API into one command. The current directory is mounted at /opt/metarank so the container can read both files.
docker run -i -t -p 8080:8080 -v $(pwd):/opt/metarank metarank/metarank:latest standalone --config /opt/metarank/config.yml --data /opt/metarank/events.jsonl.gzExpect visible progress output while Metarank imports the events and trains the model. Once it settles, the API listens on port 8080 and you send ranking requests to it. The README's text is truncated right at the point where it would show those requests, so the exact request shape is not reproduced here; the quickstart tutorial in the docs is where that continues. Two scripts in the repository root, run_demo.sh and run_quickstart.sh, appear to wrap this flow.
Where Metarank is the wrong tool
The first constraint is data. Metarank learns from interaction events, and the README's own framing is about integrating clicks and purchases into ranking. If you have no event stream, or your events are not captured in the schema the feature extractors expect, the pipeline has nothing to learn from. There is a cold-start angle here: the project's blog links include a post on solving a search cold-start problem with aggregated CTR, which acknowledges that sparse interaction data is a real obstacle rather than a solved one.
The second constraint is state. The README describes Metarank as stateless with state managed by Redis. That means Redis is not optional in a realistic deployment; the standalone quickstart sidesteps it with an in-memory store, but that store does not survive a restart and does not share state across replicas. Anyone planning horizontal scaling is planning a Redis deployment as well.
The third is scope. Metarank reranks a candidate set. It does not retrieve documents, so it does not replace your search engine, and it does not replace your recommender's candidate generation. If your problem is recall rather than ordering, this is the wrong layer. The README is explicit that the intended use is making an existing search and recommendation system smarter, not building one.
Finally, the README's feature list is not uniformly finished. The semantic neural search entry is marked TODO, and two guide links, the semantic search guide and the collaborative filtering recommendations guide, point at TODO placeholders. The LLM bi-encoder and cross-encoder extractors are documented in the feature extractors reference, but the top-level guides promised in the README are not there yet.
How Metarank differs from a Learning-to-Rank library
The closest alternative in kind is a Learning-to-Rank plugin embedded in your search engine, for example the OpenSearch LTR plugin. The difference is architectural. A plugin runs inside the search process, scores candidates at query time using features you define and store yourself, and leaves model training, feature computation and event ingestion to you. Metarank inverts that: it runs as a separate service, owns feature extraction and training, and is called after your search engine returns candidates.
That inversion buys the things a plugin cannot easily give you. Session-aware features come from a user-session extractor rather than from code you write. Model retraining and automatic feature generation are part of the product. Model serving supports multiple models for A/B testing, per the configuration reference. And because state lives in Redis rather than in the search node, the ranking tier scales independently of the search tier.
It costs you a network hop and an operational dependency. A plugin adds no latency beyond its own scoring; a separate service adds a call, plus Redis, plus a deployment to monitor. The README's 10-20ms figure is the project's answer to that objection, and it is worth reading the benchmarks page before accepting it for your result set size. The other cost is conceptual: you now have two systems that both need to agree on what a document is and what a candidate set looks like.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-09, which is recent. Release cadence is visible in the tags: 0.7.11 shipped on 2025-06-24, then 0.8.0 on 2026-08-25 and 0.8.1 on 2026-09-07. That is a long gap followed by two releases in two weeks, so the project's own history suggests bursts of activity rather than a steady drumbeat. Version numbers are still in the 0.x range, which in practice means the configuration format and API surface can change between minor releases. Anyone pinning a version should read CHANGELOG.md before upgrading rather than assuming backward compatibility.
Upgrade cost is dominated by the config file. The quickstart config is a YAML document describing events, feature extractors and models, and the repository carries a .scala-steward.conf, which indicates automated dependency update pull requests are part of the maintenance routine. If you pin the Docker image tag rather than using latest, upgrades are a deliberate act: change the tag, re-read the changelog, restart.
The licence is Apache-2.0, which permits commercial use, modification and redistribution, and includes a patent grant. It also requires that you preserve copyright and licence notices and state significant changes. This is a summary of the licence text, not legal advice; the LICENSE file in the repository is the authoritative document, and organisations with strict review processes should route it through their own counsel. Nothing in the README suggests an open-core split or a feature gated behind a commercial tier.
Editorial conclusion
Adopt Metarank if you already run a search engine or recommender and have a stream of click or purchase events to learn from. Do not adopt it if you have no event data, no Redis, or no appetite for a JVM service in the request path. Verify first that your events match the expected schema, because the whole pipeline is built around them.
Frequently asked questions
What is Metarank?
Metarank is an open-source ranking service, written in Scala and licensed under Apache-2.0, that reranks search results and recommendations using machine learning. The README describes it as a low code personalized ranking service and a Learn-to-Rank engine.
What is a meta ranking system?
The README does not define the term meta ranking. It describes Metarank as a ranking service that reranks a candidate set produced by an existing search engine or recommender, rather than retrieving candidates itself.
What is RankNet?
RankNet is not mentioned in the README or the repository material. Metarank's personalization path uses LambdaMART for secondary reranking, and its similar-items recommendations use matrix factorization with ALS.
What is an ML ranking model?
The README does not give a general definition. In Metarank's case, the ranking model is trained from imported interaction events and then served by the API, with LambdaMART used for the reranking path and ALS for similar-items recommendations.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/metarank-metarank)