BlackSnufkin/LitterBox: a self-hosted payload sandbox for red teams, driven by MCP
A self-hosted sandbox for red teams to test payloads against modern detection before deployment. MCP integration lets an LLM agent drive analysis end to end.
At a glance
- What is it?
- LitterBox runs static and dynamic analysis plus EDR-instrumented detonation against a payload and returns a Detection Score. Here is how it installs on Windows and through Docker, what the Whiskers agent and GrumpyCats CLI actually do, and where the design stops.
- Who is it for?
- Adopt LitterBox if you already run an isolated lab and need a repeatable Detection Score before a payload leaves it, and you accept Python 3.11+ on Windows or the roughly one-hour KVM container build on Linux. Do not adopt it for production traffic or as a general malware triage service for outsiders.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 148 days ago.
- What is it written in?
- Mainly YARA, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LitterBox is for, and who it is actually aimed at
The README frames LitterBox as a self-hosted payload-analysis sandbox for red teams: you upload a sample, run static, dynamic and EDR analysis against it, and get a Detection Score with a breakdown of the indicators that fired. The stated goal is a decision about whether a payload is field-ready before it leaves the lab. That is a narrower job than malware triage. The question is not "what is this file" but "how loudly does my implant announce itself to the tools a defender is likely to run".
The README adds that while the project is designed primarily for red teams, blue teams running the same tools in malware-analysis workflows can use it. That second audience is real but secondary. A blue team gets a Detection Score for the same reason a red team does: it summarises which scanners reacted, not what the sample does in the world. Everything the platform produces is a function of the bundled scanners, so its usefulness tracks how current those scanners are.
How the analysis pipeline is put together
The repository layout tells most of the story. litterbox.py is the entry point, app/ holds the Flask application (Flask 3.1.3 and Werkzeug 3.1.8 are pinned in requirements.txt), Scanners/ holds the bundled binaries and rule sets, Config/ holds configuration including EDR profiles, Whiskers/ holds the agent that runs on the detonation VM, and GrumpyCats/ holds the CLI and Python client.
The scanners listed in the README are a mix of in-process memory inspection and rule matching: PE-Sieve and Hollows-Hunter for injected or hollowed code, Moneta for memory anomalies, Patriot, Hunt-Sleeping-Beacons, RedEdr, and YARA with two rule sets, Elastic's protections-artifacts rules and YARA-Forge Extended 0.9.1. The README tracks each scanner's upstream version and last-update date, and states plainly that the last-updated date is the upstream commit or release date, not the local build date. That distinction matters when you audit a stale install.
EDR analysis is a separate path. The README says LitterBox can dispatch payloads to a separate EDR-instrumented Windows VM running Elastic Defend or Fibratus and pull correlated detection alerts back into the results page. You enable that by dropping profile YAMLs into Config/edr_profiles/, which the upload page reads at boot. The wiki has a page called All in One Pipeline that describes running static analysis and every reachable EDR in parallel. The README does not document what happens when a configured EDR VM is unreachable, so treat that as an open question rather than a guaranteed graceful degradation.
Installing on Windows and running your first analysis
The Windows path is a plain Python application. The README requires Python 3.11 or newer and an admin shell. Clone, create a virtual environment, install the pinned requirements, and start the app:
git clone https://github.com/BlackSnufkin/LitterBox.git
cd LitterBox
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt
python litterbox.py # add --debug for verbose loggingThe README says the UI then answers on http://127.0.0.1:1337. The --debug flag is documented inline in the README exactly as shown; nothing else about the startup flags is given.
On Linux the project ships a Docker route instead of a native install. The README describes a setup script that provisions a Windows 10 container with KVM and runs LitterBox inside it:
git clone https://github.com/BlackSnufkin/LitterBox.git
cd LitterBox/Docker
chmod +x setup.sh
./setup.shThe README states the initial build takes about one hour. When it finishes you get three endpoints: the install monitor on http://localhost:8006, RDP on localhost:3389 with credentials in the docker compose file, and the LitterBox UI on http://127.0.0.1:1337. That is a heavier commitment than a Python install, and it is the price of getting a Windows detonation guest without bringing your own.
For non-browser use, the wiki documents three interfaces: the GrumpyCats CLI, the GrumpyCats Python library, and LitterBoxMCP for LLM agents. The README points to the HTTP API Reference page for every endpoint. None of those interfaces are reproduced in the README itself, so you will be reading the wiki before you can script against the platform.
The MCP integration and what it changes about the workflow
The repository description states that MCP integration lets an LLM agent drive analysis end to end. The mechanism is a separate component, LitterBoxMCP, documented on its own wiki page, and the HTTP API underneath it is documented on the HTTP API Reference page. So the agent is not reaching into the Flask process; it is calling the same HTTP surface a script would.
That has a practical consequence. Anything the MCP server can do, the GrumpyCats CLI or library can do too, and the reverse. If you are evaluating LitterBox for agent-driven triage, the thing to check is the HTTP API Reference, not the MCP page, because the API is the contract both clients share. The README does not describe what the agent is expected to decide or how much of the loop it closes, and it does not give an example MCP session. That is a documentation gap, not a missing feature, but it means the integration is something you verify against the wiki before you plan around it.
Limits, and the cases where LitterBox is the wrong tool
The security advisory in the README is unusually direct. It states the platform is for development use only, that production deployment presents significant security risks, that it must run only in isolated VMs or dedicated testing environments, and that it comes with no warranty. It also puts legal compliance on the user. An application that accepts arbitrary payloads, dispatches them to a Windows guest, and exposes an HTTP API on a local port is exactly the kind of thing that should never face a network you do not control. The README does not describe authentication for the UI or the API, so assume there is none and keep it on a lab segment.
There is a second limit that is structural rather than operational. Every verdict LitterBox produces comes from the scanners bundled under Scanners/. A payload that evades PE-Sieve, Hollows-Hunter, Moneta, Patriot, Hunt-Sleeping-Beacons, RedEdr and the two YARA rule sets will score clean even if a defender's actual stack would catch it, and the reverse is equally true: a rule set that is broader than your target environment will inflate the score. The README's scanner table exists precisely so an operator can see how current each component is, and it asks you to update the row when you refresh a binary. If you do not maintain that table, the Detection Score stops meaning anything you can compare across time.
Where LitterBox is the wrong tool: it is not a general-purpose malware analysis service for untrusted third parties, and it is not a substitute for a real detonation environment with the exact EDR product and configuration your target runs. The EDR path gets you closer, but only if the profile you drop into Config/edr_profiles/ matches the deployment you care about.
How LitterBox differs from a full analysis platform
The obvious comparison is CAPE Sandbox, which is the reference open-source automated malware analysis platform. The difference in approach is the unit of output. CAPE is built around behavioural reporting: it detonates a sample in an instrumented guest and produces a report of what the sample did, with signatures, network captures and process trees. LitterBox is built around a score: it runs a fixed set of detection-oriented scanners against a sample and reduces the result to a Detection Score plus the indicators that triggered.
That makes CAPE the better tool when the question is "what does this binary do", and LitterBox the better tool when the question is "will this binary survive contact with a defender". A red team iterating on an implant wants the second question answered quickly and repeatedly; a malware analyst writing a report wants the first. The two are not interchangeable, and LitterBox's own scanner list makes the orientation clear: PE-Sieve, Hollows-Hunter and Hunt-Sleeping-Beacons are detection tools, not behavioural instrumentation.
The cost of that focus is breadth. CAPE's reporting is richer because it observes execution; LitterBox's score is narrower because it observes specific artifacts. If you need both, you run both, and the Docker route means LitterBox will sit alongside other lab services rather than replacing them.
Licence, maintenance and the cost of keeping scanners current
LitterBox is GPL-3.0. If you modify it and distribute it, the GPL's source-availability terms apply to your version. Running it internally is the ordinary case and does not trigger distribution obligations. This is not legal advice; check with counsel if you plan to ship a derivative.
The maintenance cost is not in the Python code, it is in the scanner directory. The README's table pins upstream versions and dates for a dozen components, and the project's own instruction is to replace the binary under Scanners/<Name>/ and update the row when you refresh one. YARA rule sets move fastest: Elastic's protections-artifacts and YARA-Forge Extended are both dated within days of the v5.0.0 release, and rule sets like these change continuously upstream. A LitterBox install that has not had its rules refreshed will under-report detections, and the failure is silent because the score still returns a number.
The repository's last push was on 2026-05-05, and v5.0.0 was released on 2026-05-04. The project is not archived. Beyond that, the README does not publish a support policy, a deprecation schedule or a compatibility matrix for the pinned Python dependencies, so treat an upgrade as a re-read of the scanner table plus a fresh install of requirements.txt.
Editorial conclusion
Adopt LitterBox if you already run an isolated lab and need a repeatable Detection Score before a payload leaves it, and you accept Python 3.11+ on Windows or the roughly one-hour KVM container build on Linux. Do not adopt it for production traffic or as a general malware triage service for outsiders. Verify first that your YARA rule directories under Scanners/Yara/rules/ are populated, that the EDR profile YAMLs you drop into Config/edr_profiles/ are actually picked up on the upload page, and that the bundled scanner versions in the README table match what is on disk.
Frequently asked questions
What is LitterBox and who is it for?
It is a self-hosted payload-analysis sandbox that runs static, dynamic and EDR analysis on a sample and returns a Detection Score with a breakdown of triggering indicators. The README says it is designed primarily for red teams, and is equally useful for blue teams running the same tools in malware-analysis workflows.
How do I install LitterBox on Windows?
The README requires Python 3.11 or newer and an admin shell: clone the repository, create a virtual environment, run pip install -r requirements.txt, then start python litterbox.py. The UI is served on http://127.0.0.1:1337.
Can LitterBox run in Docker on Linux?
Yes. The README directs you to the Docker directory, where setup.sh provisions a Windows 10 container with KVM and runs LitterBox inside it. The README states the initial build takes about one hour, after which the install monitor is on http://localhost:8006 and the LitterBox UI on http://127.0.0.1:1337.
How does an LLM agent use LitterBox?
Through LitterBoxMCP, which the repository description says lets an LLM agent drive analysis end to end. The README links to a dedicated wiki page for it, and the underlying HTTP endpoints are documented separately in the HTTP API Reference.
What licence is LitterBox released under?
GPL-3.0, per the repository's licence identifier. The README does not discuss licence terms beyond the security advisory's statement that users are responsible for legal compliance.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/blacksnufkin-litterbox)