# WilmerAI's router is a JSON workflow, not a classifier, and it ships three web servers

> A GPL-3.0 Python application that reads the whole conversation before deciding where to send it, then dispatches into node-based workflow files exposing OpenAI and Ollama compatible endpoints. Its documentation is candid about model-dependent behaviour and less candid about its own Python floor, which it states twice in different ways.

**SomeOddCodeGuy/WilmerAI** — WilmerAI is one of the oldest LLM semantic routers. It uses multi-layer prompt routing and complex workflows to allow you to not only create practical chatbots, but to extend any kind of application that connects to an LLM via REST API. Wilmer sits between your app and your many LLM APIs, so that you can manipulate prompts as needed.

- Repository: https://github.com/SomeOddCodeGuy/WilmerAI
- Stars: 832 · Forks: 51
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/someoddcodeguy-wilmerai

## The decision reads the conversation, then becomes a workflow file

The starting complaint is that most routers classify on the most recent message, so a follow-up like "What do you think it means?" gets filed as small talk when the conversation before it was about the Rosetta Stone. WilmerAI's answer is a node-based workflow engine, and the routing decision is itself a workflow: a JSON file defining a sequence of steps, or nodes, where each node's output becomes the next node's input. Routing happens at two levels, once at the start of a conversation to pick a specialised workflow such as coding, factual or creative, and again inside a running workflow through conditional if and then logic, so a process can choose its next step from what the previous node produced. The whole history can be used, bounded by whatever context limits the workflow configures. A node can call a different model, run a tool, execute a custom script or call another workflow, and the client sees one ordinary API call.

## The front door is OpenAI and Ollama compatible, which is how clients attach

What makes this usable by existing software is the endpoint surface, described as an adaptable API gateway that exposes OpenAI and Ollama compatible routes so current front ends and tools connect without modification. That choice also shapes the model story. Because nodes can each point at a different endpoint, a workflow can summarise with a small fast local model and reason with a large cloud one inside the same run. The documentation pushes that further into distributed inference: you can have as many models working together in a single call as you have machines to host them, spare hardware running 3 to 8b models can be put to work as worker nodes, and one prompt may reach five or more computers including proprietary APIs depending on the workflow you built. The recorded demonstration shows Open WebUI wired to two instances, one talking straight to Mistral Small 3 24b and the other calling an offline Wikipedia text API before the same model, and that recording predates multi-user support, which a single instance now has.

## Memory is four stores, and only one of them is a vector database

Conversation memory is described as a four-part system rather than a single index, and the parts have different jobs. There is a chronological summary file, a rolling summary of the entire chat that is updated continuously, a searchable vector database, and a state document describing the current state of the conversation, also maintained continuously. The vector database is the part people expect, and its default is worth noting: search is keyword based, with embedding-based semantic or hybrid search available as an option. So the default retrieval path is not semantic at all, and the difference between the two modes is a configuration choice rather than a property of the library. This memory arrangement arrived recently, since the release notes for the two most recent patch versions both advertise new memories alongside embedding endpoint support and bug fixes.

## Two patch releases share one release note, and the tag title says nothing

The tag history is short and slightly untidy. v0.7.2 was published on 2026-07-20 and v0.7.2.1 on 2026-08-09, and both carry the same release title text, embedding endpoint support, new memories, bug fixes and so on, with only the version number differing. That means the patch release did not get its own notes, so a reader diffing 0.7.2 against 0.7.2.1 has nothing to read to find out what changed. The newest tag, v0.7.3 on 2026-09-08, is titled with the bare version string and nothing else, so the informative description stops one release earlier. The repository's last push carries the same timestamp as that newest tag. Nothing here suggests the project is idle, and the disclaimer is about warranty and availability rather than activity, but for anyone who wants to know what a given release changed, the tags will not tell them.

## The Python floor is stated twice, with two minor versions apart

The maintainer notes are the first thing after the disclaimer, and they read oddly. The minimum Python version is said to have been bumped up to 3.13.14 or 3.14.5, and then, two sentences later, anyone not intending to use the web fetcher should be fine on 3.11.10 or 3.12.13, with the web fetch described as the only item that really requires the higher versions. So the stated minimum and the stated workable version differ by two minor releases, and the difference is decided by whether you need one feature. Note also the shape of the requirement: these are patch-level numbers, so a minimum of 3.13.14 is a claim about one specific build rather than about the 3.13 series, and a machine on 3.13.9 is outside it by a patch release. The repository root does carry a `.python-version` file, which is where a tool like a version manager would take its answer from instead.

## Flask with two production servers and four launch scripts

The dependency list is fully pinned, including things that are usually transitive, with urllib3 pinned alongside requests. The web stack is where the choices accumulate: Flask 3.1.3 with Jinja2 3.1.6, plus eventlet 0.41.1 and waitress 3.0.2, two production WSGI servers installed together and used as alternatives rather than together. The root directory matches that with several entry points, a server.py plus run_eventlet.py, run_waitress.py, a run_macos.sh and a run_windows.bat, so which server you get depends on which file you start. Around them sit a Middleware directory, a Public directory for the interface, a Scripts directory, a Tests directory with its own requirements-test.txt and a pytest.ini, and a ThirdParty-Licenses directory. Two dependencies explain features the README does not emphasise: the Model Context Protocol package at 1.28.1, which is how the application speaks that protocol, and PySocks at 1.7.1, the long-standing SOCKS proxy library, which fits a router that forwards to endpoints reached over a proxy.

## Tool calling behaves differently on each backend, and says so

Tool support is described with more precision than most of the document, and it separates what the application does from what the model host permits. OpenAI-style tool calling is passed through, including multi-round tool loops driven by authored-prompt workflows through a function called `appendNativeToolExchange`. When a node demands a tool call, meaning the choice is forced or marked required, that only works on backends offering constrained decoding, and the list is llama.cpp, Ollama, LM Studio, vLLM and OpenAI. Elsewhere a forced choice is only as reliable as the model. The same split applies to structure: an LLM node can ask for a JSON Schema through `structuredOutputFile` for routing, extraction or classification, with support and enforcement again dependent on the backend. Read that as the routing advice it is. A workflow whose correctness depends on guaranteed tool use or guaranteed schema output is portable only across the listed hosts.

## The demonstration section is four images and four captions

The examples are the weakest part of the document, and it is presented that way, under a heading that calls them not so pretty pictures. Four subsections exist, a simple assistant workflow using the prompt router, how routing might be used, group chat to different LLMs, and a user experience workflow where someone asks for a website, and each contains only a caption and an image with no prose. Everything a reader can learn about the routing behaviour therefore lives outside the text, which means it cannot be searched, quoted or checked against a version. The disclaimer at the top is more substantive: the project is still under development and provided as is without warranty, the work is done by the maintainer and contributors in their free time on personal hardware, and the maintainer states plainly that they are not taking contract, freelance or collaboration work. That last point is the most useful line in the document for anyone planning to depend on it.

## Conclusion

WilmerAI is for the case where a keyword router keeps sending the wrong follow-up question to the wrong model, because reading the conversation before choosing is the whole point of it, and expressing the choice as a JSON workflow you can read beats a prompt you cannot inspect. Judge the parts that vary by backend for yourself: forced tool calls and JSON schema output only do what the model host enforces, so llama.cpp, Ollama, LM Studio, vLLM and OpenAI will not behave identically for the same node. Pin a tag rather than master, since the two most recent patch releases share one release note, and read the maintainer notes about Python before you rebuild an environment: the stated minimum and the stated workable version differ by two minor releases, and the web fetcher is what decides which applies.

## FAQ

### What is WilmerAI?

An application for semantic prompt routing and task orchestration. Routing decisions read the whole conversation, and the decision is itself a workflow: a JSON file of nodes where each node can call a model, a tool, a script or another workflow, with the whole process presented to the client as one API call.

### What Python version does WilmerAI need?

The maintainer notes say the minimum was bumped to 3.13.14 or 3.14.5, and also that 3.11.10 or 3.12.13 is fine if you do not intend to use the web fetcher, which is called out as the only item needing the newer versions. A .python-version file sits at the repository root.

### Can existing clients connect to WilmerAI unchanged?

That is the design intent. The application exposes OpenAI and Ollama compatible API endpoints as its front door, including OpenAI-style tool calling passthrough, and multi-round tool loops run through authored-prompt workflows.

### How does WilmerAI remember a conversation?

With four stores: a chronological summary file, a rolling summary of the whole chat kept up to date, a searchable vector database that defaults to keyword search with optional embedding-based semantic or hybrid search, and a state document describing the current state of the conversation.

### Can WilmerAI use models running on my own machines?

Yes. Each node can point at a different endpoint, and the documentation suggests putting spare hardware running 3 to 8b models to work as worker nodes. It notes that a single prompt can reach five or more computers, including proprietary APIs, depending on how the workflow is built.

## Sources

- [Issues](https://github.com/SomeOddCodeGuy/WilmerAI/issues)
- [License: GPL-3.0](https://github.com/SomeOddCodeGuy/WilmerAI/blob/master/LICENSE)
- [README](https://github.com/SomeOddCodeGuy/WilmerAI/blob/master/README.md)
- [Releases](https://github.com/SomeOddCodeGuy/WilmerAI/releases)
- [SomeOddCodeGuy/WilmerAI on GitHub](https://github.com/SomeOddCodeGuy/WilmerAI)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/someoddcodeguy-wilmerai
