# mgrep: semantic search over code, PDFs and images from the command line

> mgrep is a TypeScript CLI from Mixedbread that indexes a repository and answers natural-language queries against it. It is a complement to grep, not a replacement, and it depends on Mixedbread's hosted service.

**mixedbread-ai/mgrep** — A calm, CLI-native way to semantically grep everything, like code, images, pdfs and more.

- Repository: https://github.com/mixedbread-ai/mgrep
- Website: https://demo.mgrep.mixedbread.com
- Stars: 4,409 · Forks: 176
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/mixedbread-ai-mgrep

## The problem mgrep targets: you cannot grep for intent

grep matches characters. If the function you are hunting is named handleSessionToken, grep finds it. If it is named doThing, or if the relevant logic lives in a PDF design note rather than a source file, grep has nothing to match on and you fall back to guessing naming conventions. The README frames this as the core motivation: "if you're looking for deeply-buried critical business logic, you cannot describe it." The consequence it describes is an agent burning its context window on hundreds of failed patterns.

mgrep is aimed at that gap. You describe what you want in natural language, and the tool retrieves passages by meaning rather than by literal match. The audience is developers working in unfamiliar or large repositories, and, judging by the install commands, coding agents acting on their behalf. The README is explicit that this is a complement: "We designed mgrep to complement grep, not replace it."

## How mgrep works: a local watcher over a hosted Mixedbread store

The architecture has two halves. Locally, mgrep runs a file watcher. The README states that mgrep watch "performs an initial sync, respects .gitignore, then keeps the Mixedbread store updated as files change." So the local side is responsible for deciding what is in scope and for detecting changes.

Remotely, the search itself happens in Mixedbread Search, described in the README as "our full-featured search solution" combining semantic retrieval models with context-aware parsing and optimized inference. That is where the embeddings and the query matching live. The practical implication is that mgrep is not a self-contained binary that indexes on your machine: the retrieval step is a network call against a hosted service, and the index is a Mixedbread store rather than a local file.

The supported content types today are code, text, PDFs and images, with audio and video listed as coming soon. The watcher is what makes the model feel like grep: you index once, then query repeatedly without a re-index step in the loop.

## Installing mgrep and running a first semantic query

mgrep is published on npm as @mixedbread/mgrep and ships a single binary named mgrep. Install it globally:

```bash
npm install -g @mixedbread/mgrep
```

Authentication is required before indexing. The default path opens a browser and walks you through Mixedbread authentication:

```bash
mgrep login
```

For CI or headless machines, the README gives an API key alternative that bypasses the browser flow entirely:

```bash
export MXBAI_API_KEY=your_api_key_here
```

With credentials in place, move into the repository you want searchable and start the watcher. It syncs once, then tracks changes:

```bash
cd path/to/repo
mgrep watch
```

Queries are then plain strings. The README's own example is a question rather than a pattern, and a second path argument narrows the search:

```bash
mgrep "where do we set up auth?" src/lib
mgrep -m 25 "store schema"
```

Searches default to the current working directory when no path is given, and -m caps the number of results. If you would rather not have an agent start the watcher for you, running mgrep watch /path/to/your/project explicitly is the documented way to keep control over when indexing happens.

## The background sync and the limits you inherit with it

The agent integrations are where mgrep's behavior changes most. The README carries a caution block stating that when mgrep is installed with a coding agent, it "runs a background process that syncs your files to enable semantic search," that this process starts when a session begins and stops when the session ends, and that usage is visible in the Mixedbread platform.

That is a meaningful shift in what a coding session does on your behalf. Files leave the machine during the session without an explicit command in your shell history. The README tells you where to look (the platform usage view) and gives you the manual alternative, but it does not document a per-project opt-out from the automatic path beyond running the watcher yourself.

Two defaults also shape what gets indexed. The README states a maximum file size of 1MB per file and a maximum file count of 1,000 files per directory, and notes that both can be changed through CLI flags (--max-file-size, --max-file-count), environment variables, or config files. In a repository with a large generated directory or a vendored dependency tree, those limits are the difference between a useful index and a partial one, and the truncation is not something you would notice from a query result alone.

## Where mgrep is the wrong tool

Exact-match work is the clearest case against it. Finding every call site of a symbol, locating a specific error string, or enumerating files that import a module are all jobs grep and ripgrep do deterministically, offline, and without an account. Semantic retrieval returns ranked passages, which is the wrong shape for a question whose answer is a complete list.

The second limitation is the hosted dependency. The README does not describe an offline or self-hosted mode, so an air-gapped machine, a repository under a policy that forbids sending source to third parties, or an environment where an API key cannot be provisioned are all disqualifying. The README also does not document rollback or how to remove a store once files have been synced, so deletion behavior is something to confirm before pointing mgrep at anything sensitive.

The third is content coverage. Audio and video are listed as coming soon, so a repository whose important context lives in recorded design reviews is not yet served by this tool. And the file-size and file-count ceilings mean very large artifacts can silently fall outside the index.

## mgrep vs ripgrep and grep: different retrieval models, not competing speeds

The comparison that matters is not speed. grep and ripgrep walk files and test regular expressions against lines; the cost is proportional to the corpus and the result is exact. mgrep embeds content into a hosted store and matches a query against those embeddings; the cost is a network round trip and the result is a ranked set of passages that may or may not contain the literal terms you typed.

That difference decides the split. ripgrep is the right tool when you know the string. mgrep is the right tool when you know the concept but not the vocabulary, or when the answer is in a PDF or a screenshot rather than a source file. The README's own framing puts both in the toolkit: "use grep for exact matches, mgrep for semantic understanding and intent." Running both is not redundancy; it is two different questions.

The other named alternative is simply not using mgrep: grep plus a human who knows the codebase. That has no account, no sync, and no file-size ceiling, and for a small repository with consistent naming it is often faster than any index.

## Licence, maintenance and the upgrade path

mgrep is licensed Apache-2.0, and the LICENSE file sits at the repository root. Apache-2.0 permits commercial use and modification with the usual notice and patent terms; it says nothing about the hosted Mixedbread service that the CLI talks to, which is governed separately. Treating the CLI's licence as covering the search backend would be a mistake. This is a description of the licence text, not legal advice.

The repository is not archived. The last push was on 2026-04-25, and the most recent release, v0.1.13, is dated the same day, following v0.1.12 on 2026-04-14 and v0.1.11 on 2026-04-08. The version number is still in 0.1.x, so the surface area can move between releases.

Upgrade cost is mostly the npm global install: npm install -g @mixedbread/mgrep pulls the new version, and the build script copies the plugins directory alongside the compiled entry point, which is how the agent integrations ship. Because the retrieval logic lives server-side, some behavior changes will arrive without a CLI upgrade at all, and the README does not describe a version pinning mechanism for the backend.

## Conclusion

Adopt mgrep if your searches are conceptual ("where do we set up auth?") rather than literal, and if sending repository content to a hosted service is acceptable for that repository. Do not adopt it if you need fully local search, if you work in an environment where an outbound API key cannot be provisioned, or if your queries are exact strings, where grep and ripgrep remain faster and dependency-free. Before rolling it into an agent workflow, verify two things: which directories mgrep watch is actually indexing, and whether the 1MB per-file and 1,000 files-per-directory defaults cover your tree, since the README states both are adjustable through --max-file-size and --max-file-count.

## FAQ

### How do I install mgrep?

Install it globally from npm with npm install -g @mixedbread/mgrep, or use pnpm or bun. After that, run mgrep login once to authenticate before indexing a project.

### How do I use mgrep?

Run mgrep watch inside a repository to perform the initial sync and keep the index updated, then query in natural language, for example mgrep "where do we set up auth?" src/lib. Searches default to the current working directory unless you pass a path, and -m limits the number of results.

### What is mgrep?

mgrep is a CLI from Mixedbread that provides semantic, natural-language search over code, text, PDFs and images, backed by Mixedbread Search. The README describes it as a complement to grep rather than a replacement.

### Is mgrep safe?

The README states that when mgrep is installed with a coding agent it runs a background process that syncs your files to a Mixedbread store, starting when a session begins and stopping when it ends, with usage visible in the Mixedbread platform. It does not document an offline mode or a rollback procedure, so the decision depends on whether repository content may be sent to that hosted service.

### Is there a free alternative to mgrep?

grep and ripgrep cover exact-match search with no account, no sync and no file-size limit, and the README recommends using them alongside mgrep rather than instead of it. They do not provide semantic or multimodal retrieval, so they are an alternative for literal queries only.

## Sources

- [License: Apache-2.0](https://github.com/mixedbread-ai/mgrep/blob/main/LICENSE)
- [mixedbread-ai/mgrep on GitHub](https://github.com/mixedbread-ai/mgrep)
- [Project website](https://demo.mgrep.mixedbread.com)
- [README](https://github.com/mixedbread-ai/mgrep/blob/main/README.md)
- [Releases](https://github.com/mixedbread-ai/mgrep/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/mixedbread-ai-mgrep
