aimock: one local port that answers for every AI provider your app calls
Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies.
At a glance
- What is it?
- aimock is a TypeScript mock server from CopilotKit that stands in for LLM APIs, MCP tools, A2A agents, AG-UI streams, vector databases and search. The single-port design is the real idea here, and the record-and-replay fixtures are the part worth judging before you commit.
- Who is it for?
- Adopt aimock if your test suite already touches more than one AI surface and you are tired of hand-rolling a fake per provider: one port, one package, and fixtures that keep the same bytes across runs. Skip it if you only need a single OpenAI chat stub, since a plain HTTP interceptor is less machinery.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The test-suite problem aimock was built around
An AI application rarely talks to one service. A chat feature calls an LLM, retrieves context from a vector database, may invoke an MCP tool, and streams events to the frontend over AG-UI. Each of those has its own SDK, its own base URL override, and its own idea of what a fake looks like. Teams end up maintaining a pile of hand-written stubs, and every provider SDK upgrade can shift a response shape without anyone noticing until a test goes red for the wrong reason.
aimock's answer is to put all of those surfaces behind one process listening on one port. The README frames it as "Mock infrastructure for AI application testing" and the pitch is explicit: point your SDK at a local port and every provider, protocol and service answers deterministically. The audience is engineers writing integration and end-to-end tests for AI features, not people unit-testing a prompt string. If your test only needs to assert that a function returns a string, this is more machinery than the job requires.
One port, many protocols: how the suite is composed
The package ships several mock classes rather than a single monolith. LLMock covers provider HTTP surfaces: OpenAI Chat, Responses and Realtime, Claude, Gemini REST and Live, Bedrock, Azure, Vertex AI, Ollama, Cohere, OpenRouter and ElevenLabs TTS. MCPMock handles MCP tools, resources and prompts with session management. A2AMock speaks the agent-to-agent protocol over SSE. AGUIMock emits AG-UI event streams for frontend tests. VectorMock exposes Pinecone, Qdrant and ChromaDB compatible endpoints. A services layer covers Tavily search, Cohere rerank, OpenAI moderation and ElevenLabs TTS.
You can run the whole set from a config file, or import only the pieces you need. The subpath exports in package.json mirror that split: the root entry, plus ./mcp, ./a2a and ./vector stubs. That structure matters in practice, because it means a test that only needs an MCP server does not have to boot the provider mocks. The Dockerfile shows the same composition from the other direction: it exposes port 4010 and runs the CLI with --fixtures ./fixtures --host 0.0.0.0, so the container serves a fixture directory rather than a programmatic setup.
Installing aimock and running a first mocked call
The README gives a two-step quick start. Install the package, then construct LLMock with a port and register a response for a message.
npm install @copilotkit/aimockThe example below is taken from the README. Note the import name: the class is still called LLMock for backwards compatibility after the v1.7.0 rename from @copilotkit/llmock to @copilotkit/aimock. Passing port 0 asks the operating system for a free port, and mock.url reports the address it landed on.
import { LLMock } from "@copilotkit/aimock";
const mock = new LLMock({ port: 0 });
mock.onMessage("hello", { content: "Hi there!" });
await mock.start();The ordering constraint is the part to read twice. The README warns that the base URL environment variables must be set before the provider client is imported or constructed, because many SDKs cache the base URL at construction time. If the client is built first, it talks to the real API instead of aimock.
process.env.OPENAI_BASE_URL = `${mock.url}/v1`;
process.env.OPENAI_API_KEY = "mock"; // SDK requires a value, even when base URL is mocked
// ... run your tests ...
await mock.stop();The API key still has to be a non-empty string, because most SDKs refuse to construct a client without one. When the test finishes, await mock.stop() releases the port. The README does not document what happens to in-flight streaming connections at that point, so treat shutdown behaviour during a live SSE stream as something to check against your own test harness.
Record and replay, and what the fixtures actually preserve
The feature that separates aimock from a static stub library is the record-and-replay path. You proxy real APIs, save the traffic as fixtures, and replay it without network access afterwards. The documentation is specific about what replay keeps and what it does not.
Recorded fixtures capture per-frame arrival timestamps. Replay uses those timings to approximate the original pace, based on recorded time-to-first-token and inter-frame cadence, with a configurable --replay-speed multiplier. The README is honest that this is approximate: the replay chunk count may differ from the recording, so TTFT and average pace are preserved but per-token fidelity is not. If your test asserts on the exact number of streamed chunks, that assertion will not survive a re-record.
Usage accounting is handled differently, and more convincingly. Recording captures the final usage frame of a streaming completion, so replayed fixtures serve real prompt_tokens and completion_tokens values rather than a length estimate. For OpenRouter, the provider-reported usage.cost and its cost_details and *_tokens_details breakdowns are captured as well. That makes an application that bills from real provider cost testable end to end from a tape, which is a narrow but genuinely hard case to fake by hand.
Multi-turn traces are matched through predicates rather than a script: turnIndex, hasToolResult, toolCallId, toolResultContains, sequenceIndex, systemMessage, or a custom predicate. The systemMessage gate is aimed at host-supplied agent context, so a fixture can respond differently depending on what the agent scaffold injected. That is a lot of matching surface, and the cost is fixture design work: a trace that matches too loosely will answer turns it should not.
Where aimock is the wrong tool
The zero-dependency, deterministic framing has a boundary, and it is the same boundary every mock server has: fixtures encode your assumptions about provider behaviour, not the behaviour itself. The repository does carry drift tests and a DRIFT.md file, which suggests the project tracks divergence between its mocks and the real APIs, but the README does not describe how drift is detected or what happens when it is found. If your integration depends on a provider response shape that the mocks do not model, a green test tells you nothing.
The second limit is environmental. The README's own warning about base URL caching is a real failure mode with a costly symptom: a client constructed before the environment variables are set will call the live API and incur charges, and the test may still pass. Any SDK that reads its base URL at import time rather than at request time needs the mock started in a setup file that runs before the client module is loaded. That is a test-harness design constraint, not something the library can enforce for you.
Finally, aimock is a Node package with engines set to node >=20.15.0. A team whose tests run in a browser-only environment, or on a runtime without Node 20, is outside the supported path. The README does not present a browser build, so do not plan on one.
How aimock differs from WireMock and MSW
The closest comparison is WireMock. WireMock is a general HTTP mock server: you define request matchers and canned responses, and it will happily stand in for any REST endpoint. It has no concept of an SSE token stream, an MCP session, or a provider-specific usage frame. With WireMock you would model an OpenAI streaming response by hand, including the chunked framing and the terminal usage object, and you would rebuild that model for each provider.
MSW takes the opposite approach. It intercepts requests inside the process through a service worker or a Node interceptor, so there is no port and no separate server. That is lighter for a frontend test, but it means every test process carries the interception layer, and cross-language or cross-process tests (a Python client hitting your mock) are not the model. aimock sits at a real port, which is why the Dockerfile can run it as a standalone service and why a non-JavaScript client can point at it.
The trade is that aimock only knows the surfaces it ships mocks for. WireMock is provider-agnostic; aimock is provider-deep. If your test needs to mock an internal REST API alongside the LLM, you will likely run aimock for the AI surfaces and something else for the rest, or use the record-and-replay proxy to capture that internal API too.
Maintenance, licence and upgrade cost
The repository is not archived and the last push was on 2026-09-10. Releases have been frequent and versioned in the 1.x line: v1.38.0 on 2026-08-04, v1.39.0 on 2026-08-19, v1.40.0 on 2026-09-09. Minor releases at that cadence mean you should pin the version in CI rather than floating on a caret range, because a fixture format or matching predicate can change between minors. The package.json in the repository shows version 1.42.0, ahead of the newest listed release, so the main branch carries work that has not shipped yet.
The licence is MIT, which permits commercial use and modification with the copyright notice retained. That is a permissive choice with few obligations, but it also means there is no warranty and no support commitment attached to the code. Nothing here is legal advice; if you redistribute aimock inside a product, read the LICENSE file in the repository rather than this summary.
The upgrade cost is concentrated in fixtures, not in code. Because replayed fixtures encode recorded timings and usage frames, a fixture recorded against one provider version may not match the next. The README's note that replay chunk count can differ from the recording is the kind of detail that turns into a flaky assertion six months later. Budget for re-recording tapes when you bump provider SDKs, and keep the recording scripts in the repository so that is a command rather than an archaeology project.
Editorial conclusion
Adopt aimock if your test suite already touches more than one AI surface and you are tired of hand-rolling a fake per provider: one port, one package, and fixtures that keep the same bytes across runs. Skip it if you only need a single OpenAI chat stub, since a plain HTTP interceptor is less machinery. Before wiring it into CI, verify the one thing the README treats as a footgun: that OPENAI_BASE_URL and the other base URL variables are set before the provider client is constructed, then check whether your SDK reads that variable at import time or at request time.
Frequently asked questions
What is aimock and what does it mock?
aimock is a mock server for AI application testing. According to the README it stands in for LLM provider APIs, MCP tools, A2A agents, AG-UI event streams, vector database endpoints, and services such as Tavily search, Cohere rerank and OpenAI moderation, all on one local port.
How do I install aimock?
The README's quick start is a single npm install of @copilotkit/aimock, followed by importing the LLMock class from the same package name. The class kept the LLMock name after the v1.7.0 rename from @copilotkit/llmock.
Why does my test still hit the real OpenAI API when using aimock?
The README warns that many SDKs cache the base URL at construction time, so OPENAI_BASE_URL and OPENAI_API_KEY must be set before the provider client is imported or constructed. If the client is built first, it talks to the real API instead of the mock.
Does aimock replay streaming responses with the original timing?
It approximates them. Recorded fixtures capture per-frame arrival timestamps, and replay uses the recorded time-to-first-token and inter-frame cadence with a configurable --replay-speed multiplier, but the README states the replay chunk count may differ from the recording, so per-token fidelity is not preserved.
What is a mock API?
In aimock's case it is a local server that returns deterministic responses in place of a real provider, so tests run without keys, network access or provider bills. The README describes running the whole suite on one port with npx @copilotkit/aimock --config aimock.json.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/copilotkit-aimock)