llm-answer-engine: a Perplexity-style answer engine you assemble from API keys
Perplexity Inspired Answer Engine
At a glance
- What is it?
- developersdigest/llm-answer-engine is a Next.js reference implementation that chains Groq, Brave Search, Serper and OpenAI embeddings into cited answers, images, videos and follow-up questions. The code is MIT-licensed and readable; the running cost is four third-party accounts.
- Who is it for?
- Adopt it if you want a working, readable skeleton of a citation-first answer engine and you are willing to own four API accounts and the per-query spend that comes with them. Do not adopt it if you need a supported product with a documented upgrade path, or if you cannot use third-party search and inference services at all.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 155 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What llm-answer-engine actually solves
A plain chat model answers from its weights. Ask it about yesterday's release and it either refuses or invents. The fix is retrieval: search the web, pull the top pages, embed the text, keep the chunks closest to the question, and put those chunks in the prompt with links. llm-answer-engine packages that whole loop as a runnable Next.js app, which is why the README calls it "an ideal starting point for developers interested in natural language processing and search technologies" rather than a product.
The audience is narrow and specific. You are a developer who wants to see how a Perplexity-shaped pipeline is wired in TypeScript, or you want a base to fork for an internal research tool. The README lists the stack plainly: Groq and Mixtral for inference, Langchain.JS for text splitting and embeddings, Brave Search for sourcing content and images, Serper for video and image results, OpenAI embeddings for vectors, Cheerio for HTML parsing. Nothing here is novel research. The value is that the glue already exists and you can read all of it.
The retrieval pipeline and where each key is spent
The flow implied by the configuration file is a four-stage chain. A query arrives at the Next.js app. A search provider returns result URLs; the README and .env.example both treat Serper as the default and Brave as the alternative, and the example file also carries optional GOOGLE_SEARCH_API_KEY and GOOGLE_CX entries. The app then fetches those pages, and Cheerio parses the HTML so the text can be extracted.
Stage three is chunking and embedding. app/config.tsx exposes textChunkSize: 800, textChunkOverlap: 200, numberOfSimilarityResults: 2 and numberOfPagesToScan: 10. Those four numbers define the cost and the quality ceiling at the same time: ten pages scanned, split into 800-character windows with 200 characters of overlap, and only the two closest chunks reach the model. Stage four is generation, with inferenceModel set to 'mixtral-8x7b-32768' and nonOllamaBaseURL pointing at https://api.groq.com/openai/v1, so Groq serves an OpenAI-compatible endpoint by default.
Two switches matter more than they look. useFunctionCalling: true enables the beta tool routing described in the README, which covers Serper Locations, Serper Shopping, a TradingView stock widget and Spotify. useRateLimiting and useSemanticCache both default to false, so a fresh clone has no throttle and no cache in front of paid calls. The .env.example pairs them with UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN, and the package manifest carries @upstash/ratelimit, @upstash/redis and @upstash/semantic-cache.
Installing llm-answer-engine and running one query
The repository offers two paths. Both need keys from OpenAI, Groq, Brave Search and Serper, per the prerequisites section. The Docker path clones the repo inside the image, so you edit docker-compose.yml rather than a local checkout. The compose file defines a single service on port 3000 and passes the four keys as environment variables.
services:
llm-answer-engine:
build: .
environment:
- OPENAI_API_KEY=your_openai_api_key
- GROQ_API_KEY=your_groq_api_key
- BRAVE_SEARCH_API_KEY=your_brave_search_api_key
- SERPER_API=your_serper_api_key
ports:
- 3000:3000After filling in real values, start it with docker compose up -d. The README notes the older docker-compose up -d form for Compose v1. Note that the Dockerfile's ENTRYPOINT is bun run dev, so the container starts the development server, not a production build.
The non-Docker path is more conventional and gives you a config file you can edit. Clone the repo, then install and create the environment file:
npm install
# or: bun installOPENAI_API_KEY=your_openai_api_key
GROQ_API_KEY=your_groq_api_key
BRAVE_SEARCH_API_KEY=your_brave_search_api_key
SERPER_API=your_serper_api_keyPut those four lines in a .env file at the project root and run npm run dev or bun run dev. The package scripts confirm dev maps to next dev. The server listens on the port Next.js prints in the terminal. Open it, type a question, and the expected output is a streamed answer with source links, plus image, video and follow-up sections when the function-calling paths fire. If you want local inference instead of Groq, set useOllamaInference and useOllamaEmbeddings to true in app/config.tsx and point OLLAMA_BASE_URL at your server; the .env.example shows both http://localhost:11434/v1 and a LAN example.
Four keys, no cache, and a dev server in the container
The first real limitation is economic, not technical. A single question can trigger a search call, up to ten page fetches, embedding calls for every chunk, and a generation call. With useSemanticCache: false and useRateLimiting: false out of the box, nothing deduplicates or throttles that. The Upstash integrations exist and are documented, but they are opt-in, and the README does not state what a semantic cache hit costs or how it interacts with fresh results.
The second is the container. The Dockerfile ends with ENTRYPOINT ["/usr/local/bin/bun", "run", "dev"], which runs the Next.js development server. That is fine for a demo and wrong for anything public: dev mode compiles on demand and is not the build the framework expects you to ship. The README does not document a production Docker target.
The third is dependency drift. The manifest pins next to 14.1.2 but pulls @langchain/community and ai at "latest". Two of the libraries closest to the retrieval and streaming logic float. There are no retrieved releases, and the README documents no upgrade or rollback procedure, so a fresh install months apart can resolve different code for the same commit.
Finally, the app is only as good as the pages it can parse. Cheerio reads static HTML. A search result behind a JavaScript-rendered shell, a paywall, or a bot check yields little or no text, and the README does not describe a fallback. The codebase does include puppeteer-core and jsdom in the manifest, but the documentation does not explain when either is used instead of Cheerio.
Compared with Perplexica and other self-hosted answer engines
Perplexica is the closest well-known alternative, and the difference is architectural rather than cosmetic. Perplexica runs its own search backend, SearXNG, and pairs it with local models through Ollama, so the retrieval layer and the inference layer can both stay on your hardware. llm-answer-engine assumes the opposite: Serper or Brave for search, Groq or OpenAI for inference, with Ollama listed as an optional mode in app/config.tsx and OLLAMA_BASE_URL.
That trade is real in both directions. The hosted-API route gives you better default answer quality without a GPU, and the README's Vercel deploy button makes the whole thing a one-click clone with environment variables. The self-hosted route gives you no per-query bill and no data leaving your network, but you own the search index and the model weights. If your constraint is "no third-party calls," llm-answer-engine is the wrong starting point; if your constraint is "no GPU and I want to see it working today," it is the faster one. A third option is to skip both and call a hosted answer API directly, which removes the code you are here to read.
Licence, maintenance and the cost of keeping it running
The repository is MIT-licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission text are preserved. That covers the code in the repository. It does not cover the services the code calls: Groq, OpenAI, Brave Search, Serper, Upstash, Portkey and fal.ai all have their own terms, and the .env.example lists FAL_KEY for Stable Diffusion 3 and Portkey keys for Bedrock routing. Review those separately before shipping anything customer-facing; this is a description of the licence file, not legal advice.
On maintenance, the last push to the default branch was on 2026-04-29. The repository is not archived. There are no retrieved releases, so there is no tagged version to pin and no changelog to read between commits. In practice that means you should treat your fork as the artifact you maintain: pin the floating "latest" dependencies yourself, and expect to reconcile LangChain and Vercel AI SDK changes on your own schedule rather than someone else's.
Editorial conclusion
Adopt it if you want a working, readable skeleton of a citation-first answer engine and you are willing to own four API accounts and the per-query spend that comes with them. Do not adopt it if you need a supported product with a documented upgrade path, or if you cannot use third-party search and inference services at all. Before committing, run the Docker path with real keys and watch what the /api/chat route does when Brave or Serper returns nothing; the README documents configuration but is silent on failure handling and on rollback between versions.
Frequently asked questions
What is an answer engine?
In this project's terms it is a system that searches the web for a query, extracts and embeds the page text, and generates a written answer with the sources attached, rather than returning a list of links. The README describes llm-answer-engine as returning sources, answers, images, videos and follow-up questions from a single query.
Is there an open source AI-based search engine available?
llm-answer-engine is one: it is MIT-licensed TypeScript, and the README frames it as a starting point for developers working with natural language processing and search. It is not self-contained, though, since it depends on external search and inference APIs such as Serper, Brave Search and Groq.
Which API keys does llm-answer-engine require?
The prerequisites list OpenAI, Groq, Brave Search and Serper. The .env.example adds optional entries for Ollama, Upstash Redis, Google Search, Portkey, Spotify, AWS Bedrock and fal.ai.
Can llm-answer-engine run without OpenAI?
Yes, according to the README's Vercel deploy note, which says that if you use Groq instead of OpenAI you can enter a random string in the OpenAI key field so the build does not error. The configuration file also lets you switch useOllamaInference and useOllamaEmbeddings on and point OLLAMA_BASE_URL at a local or LAN server.
How do I change the model or chunk size in llm-answer-engine?
Edit app/config.tsx. It exposes inferenceModel, embeddingsModel, textChunkSize, textChunkOverlap, numberOfSimilarityResults, numberOfPagesToScan and nonOllamaBaseURL, along with the useOllamaInference, useRateLimiting, useSemanticCache and usePortkey flags.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/developersdigest-llm-answer-engine)