Model or dataset
treygrainger/ai-powered-search avatar
treygrainger/ai-powered-search

AI-Powered Search: Code Examples for the Manning Book on Modern Retrieval

The codebase for the book "AI-Powered Search" (Manning Publications, 2025) and associated "AI-Powered Search: Modern Retrieval for Humans & Agents" Maven course

411 stars122 forksJupyter NotebookLicense varies

At a glance

What is it?
This repository contains the Jupyter Notebook codebase for the Manning Publications book 'AI-Powered Search' by Trey Grainger, Doug Turnbull, and Max Irwin. It covers semantic search, retrieval-augmented generation, learning to rank, and other machine-learning-driven search techniques, all running inside Docker containers against Apache Solr by default.
Who is it for?
This repository suits engineers working through the AI-Powered Search book who want runnable code for each chapter. It is not a standalone search framework; without the book, the notebooks lack the context that explains why each technique is structured the way it is.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What This Repository Contains and Who It Is For

The repository is the companion codebase for the book AI-Powered Search, published by Manning Publications in 2025 and written by Trey Grainger, Doug Turnbull, and Max Irwin. Its primary audience is engineers who have purchased the book and want to run the code examples interactively rather than reading static listings.

The README describes the book as teaching modern, data-science-driven search techniques including semantic search with dense vector embeddings, retrieval-augmented generation, question answering and summarization combining search and LLMs, fine-tuning transformer-based models, personalized search based on user signals, collecting behavioral signals for ranking models, semantic knowledge graphs, machine-learned ranking, and generative and hybrid search.

The last push was on 2026-09-21. All code is in Python, with PySpark for data processing tasks. The default search engine used throughout the examples is Apache Solr, but the README states that most techniques are abstracted away from the specific engine and that swappable implementations for other search engines and vector databases are planned.

Repository Layout and How to Navigate It

The top-level directory separates concerns clearly. The chapters/ directory holds the Jupyter notebooks that correspond to specific book chapters; starting at chapters/welcome.ipynb gives a table of contents for all live code examples. The engines/ directory contains Dockerfile-based builds for each supported search backend, including Solr, OpenSearch, and others. The aips/ directory is a Python library shared across the notebooks. The data/ directory holds datasets used by the examples. The semantic_search/ and ltr/ (learning to rank) directories group notebooks by technique area.

The docker-compose.yml defines named profiles. The default profile starts the notebooks container (Jupyter on port 8888), the Solr container (port 8983), and ZooKeeper (used by Solr in cloud mode). An OpenSearch profile starts an alternative backend instead. Additional profiles exist for other engines listed in engines/README.md.

One practical constraint: the docker-compose.yml binds the repository root as a volume at /tmp/notebooks/ inside the notebooks container. This means edits to notebooks on the host are reflected immediately inside the container, which helps when working through the exercises, but it also means the entire repository directory (including datasets) is mounted and must fit within Docker's disk allocation.

Running the Codebase

The README emphasizes that Docker is the only required setup step. Clone the repository and start the containers:

bash
git clone https://github.com/treygrainger/ai-powered-search.git
bash
cd ai-powered-search
docker compose up

Once the containers are running, visit http://localhost:8888 to open the Welcome notebook. The README notes that the first build may take a while, particularly for the Solr and Spark components. Appendix A of the book provides full step-by-step instructions for troubleshooting.

For the OpenSearch backend, the docker-compose.yml shows the service requires environment variables including OPENSEARCH_INITIAL_ADMIN_PASSWORD, with DISABLE_SECURITY_PLUGIN set to true for the development configuration. The OpenSearch container is configured for a single-node cluster with 1 GB of Java heap. The Spark Master UI runs at port 8082 (remapped from the default 8080 to reduce port conflicts), the Spark Worker UI at 8081, and a search webserver at port 2345.

Techniques Covered and Their Organization

The book and repository cover the full arc from traditional relevance tuning to modern generative approaches. The semantic_search/ directory groups the vector embedding examples. Semantic search in the book uses dense embeddings from foundation models, including techniques for fine-tuning transformer-based LLMs on domain-specific data.

The retrieval-augmented generation chapters combine a search retrieval step with an LLM summarization step. This is distinct from pure generative search: the search step grounds the LLM response in indexed content, which the book treats as a technique for reducing hallucination in question-answering applications.

The ltr/ directory contains learning-to-rank examples. These cover collecting user behavioral signals (clicks, dwell time) and building click models that automate the training of machine-learned ranking functions. Semantic knowledge graphs for domain-specific learning appear as a separate technique, distinct from general vector similarity.

Personalized search using user signals and vector embeddings appears as a separate chapter area. The README states most techniques are portable to other engines, but the default Solr implementation is the most thoroughly tested path through the material.

Supported Engines and Portability Limits

Apache Solr is the default engine. The engines/ directory contains build configurations for additional backends, and the repository structure abstracts most technique demonstrations away from engine-specific APIs. The README invites search engine and vector database providers to contact the authors about adding support for their technology.

The README explicitly notes that swappable implementations for most popular search engines and vector databases will be available soon, which means as of the last push date (2026-09-21) that work is ongoing. Engineers who need to run these examples against Elasticsearch or Weaviate, for instance, should check engines/README.md for the current status rather than assuming full parity.

PySpark handles the data processing tasks throughout. The docker-compose.yml runs a Spark Master on port 7077, a Spark Master UI on port 8082, and a Spark Worker UI on port 8081. Teams unfamiliar with Spark should expect a steeper setup curve when troubleshooting the data pipeline notebooks compared to the simpler retrieval examples.

License, Dependencies, and a Practical Limitation

The README states that all code in the repository is open source under the Apache License 2.0 unless otherwise specified.

A notable license note in the README: executing the code may pull in additional dependencies that follow alternate licenses, and some datasets may be derived from AI models or web crawls subject to fair-use considerations. The README states datasets are published for demonstrating the concepts in the book, and their associated licenses may change over time.

The practical limitation that matters most in day-to-day use: this codebase is not a self-contained tutorial. The notebooks carry code but not the explanatory prose from the book. Without access to the book chapters, some notebooks will be opaque about why a particular embedding approach is chosen or why a specific ranking signal is weighted the way it is. The README explicitly recommends purchasing a copy of the book to get the needed context and insights.

Editorial conclusion

This repository suits engineers working through the AI-Powered Search book who want runnable code for each chapter. It is not a standalone search framework; without the book, the notebooks lack the context that explains why each technique is structured the way it is. Before setting up, confirm that Docker can allocate enough memory to run Solr, Spark, and the Jupyter service concurrently, because the docker-compose.yml spins up multiple containers including ZooKeeper, and resource constraints on a laptop frequently cause individual containers to fail silently.

Frequently asked questions

What is AI-Powered Search and what does it teach?

AI-Powered Search is a Manning Publications book by Trey Grainger, Doug Turnbull, and Max Irwin that teaches machine-learning-driven search techniques including semantic search with vector embeddings, retrieval-augmented generation, learning to rank, and personalized search. This repository holds the companion Jupyter Notebook codebase for all the examples in the book.

How do you use the AI-Powered Search codebase?

Clone the repository, then run docker compose up from the ai-powered-search directory. Once the containers are built and running, open http://localhost:8888 to access the Welcome notebook, which provides a table of contents for all the chapter examples. Docker is the only required installation step.

How do you build an AI-powered search engine from scratch?

The repository covers the full pipeline: indexing documents into Apache Solr, applying dense vector embeddings for semantic search, collecting user behavioral signals for learning-to-rank models, and combining retrieval with LLM summarization for RAG. Each technique is demonstrated in a Jupyter Notebook that can be run with a single docker compose up command.

Official sources

  1. Issues
  2. Project website
  3. README
  4. treygrainger/ai-powered-search on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/treygrainger-ai-powered-search.svg)](https://hysenlabs.com/projects/treygrainger-ai-powered-search)