MOZI checks the filesystem before it claims a deliverable, and cannot sign its own app
A custom Agent OS built to be hackable, heavily inspired by OpenClaw.
At a glance
- What is it?
- A personal agent operating system that runs on your machine, executes shell commands and file edits in a project you pick, and verifies every deliverable against disk before reporting done. The macOS build is unsigned and unnotarised, and the compose file binds to loopback for a reason stated in a comment.
- Who is it for?
- Use MOZI if you want an agent that operates on real files in a real repository rather than one that only answers, and if you are the kind of person who will read its source to understand how it works, since that is the stated reason it exists. The filesystem verification of deliverables is the feature worth having, because it is the difference between an agent that claims success and one that cannot.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 57 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Deliverables are checked on disk before the agent reports done
One sentence in the description is the whole reason to care about this project.
Every deliverable it claims is verified against the filesystem before it reports done. No fake success.
That is a mechanism, not a promise, and it is worth being precise about what it does. The agent is given a project folder and a task, and it runs shell commands, edits files, does research and generates documents. At the end, rather than emitting a completion message, the filesystem is consulted. A document that was supposed to be produced either exists or it does not.
This closes the failure mode that makes agent tools untrustworthy in practice. A model asked to write a file will report writing the file. Asking whether it wrote the file is the same prompt twice. Checking the disk is a different question with a different oracle.
The composer is described as the cockpit, and it holds the four choices that determine what the agent is allowed to do: the project, which is any folder or git repository; the branch to work on; the permission level, running from read only to full access; and the model. Those four controls are what turn an agent from a demo into something you can point at real work, and putting them in one place before the task is described is the right sequencing.
The branch switcher refuses to stash and refuses to force
The code capability is described with three refusals, and refusals are the interesting part of any tool that touches a repository.
The agent reads real repositories, writes and edits files, runs tests, and works on the branch you pick. The built in branch switcher performs an honest git switch: it never auto stashes, it never forces, and conflicts abort with git's own message.
Each of those is a decision about whose problem a failure is.
Not auto stashing means uncommitted work is never silently moved aside. The common shortcut is to stash, switch, and hope the user notices. That converts a visible conflict into an invisible one, and the user's work ends up in a stash they did not create.
Not forcing means the agent cannot resolve a conflict by discarding one side. Forcing a switch over local changes destroys them, and an agent doing that on a repository you care about is not something you want running unsupervised.
Aborting with git's own message means the error the user sees is the one git would have produced, not a paraphrase. That matters for two reasons: the message is accurate, and it is greppable, so searching your history for what went wrong still works.
Taken together this is a tool that will stop rather than improvise when it meets a situation requiring a judgement call. For a permissioned agent that is the correct default.
The macOS app is unsigned, and the checksum is not offered as a substitute
The download section contains a warning block, and the second sentence of it is the careful one.
The current macOS downloads are not signed with an Apple Developer ID and are not notarised. Gatekeeper may report the app as damaged, as unable to open, or as being from an unidentified developer. Verify the download's checksum against the published sums file before removing quarantine.
And then: the checksum confirms which artifact you downloaded. It does not replace Apple code signing trust.
That distinction is the part worth internalising, because most projects conflate the two. A checksum answers one question: did the bytes arrive intact, and are they the ones the publisher built. It says nothing about who built them, or whether the build machine was trustworthy. Publishing a checksum next to an unsigned binary is a genuine integrity control and a genuine absence of a trust control, and this documentation names both rather than implying the first covers the second.
The removal step that follows is explicit about the order of operations. Verify first, then remove the quarantine attribute from the verified copy in Applications, then open it:
xattr -dr com.apple.quarantine /Applications/MOZI.app
open -a /Applications/MOZI.appRemoving quarantine from an unverified download is how you get code execution you never checked, and the sequence here puts verification before the irreversible step.
A build-from-source route is given for people who would rather not take that path at all, and the resulting application bundle is a path you drag into Applications yourself.
The Node window has a hole in the middle for two different reasons
The requirements sentence is the most precisely argued piece of documentation in the repository.
Node 22 long term support or newer is needed, and 26 or newer is excluded. The reason is that 26 is too new for the project's native dependencies. The version in between, 23, is rejected by dependency engine pins.
That is two different exclusion mechanisms producing one contiguous forbidden range. The newest major is excluded because the native addons have not caught up. One specific odd major in the middle is excluded because some dependency declares an engine range that does not include it. Neither is a preference and both are checkable.
Then there is a longer sentence about package management, which is unusually helpful. You do not need corepack or a global package manager, because the setup script resolves the pinned one on its own. It finds a compatible Node from several version managers or from Homebrew, without sudo and without a global install. If no usable Node exists it prints the exact install command for your platform. And there is a check flag that reports what it would use without changing anything.
One instruction is repeated three times: avoid the corepack enable command. It needs sudo on the official Node installs, and it no longer exists on recent Node versions. The setup script exists specifically so you never have to run it.
Intel Macs are excluded by a vector database, not by the agent
The server mode section contains one limitation and one piece of configuration guidance, and the limitation is more interesting than it looks.
The supported server architectures are Linux on x64 and arm64, and macOS on Apple silicon. Intel Macs are not supported, and the reason given is that the vector database dependency does not publish a macOS x64 binary.
So the constraint comes from the memory subsystem rather than from the agent itself. That is worth knowing for two reasons. First, it means the boundary may move without the project changing, when the dependency publishes a binary. Second, it tells you the optional vector memory is not as optional as it looks in a packaged context, because there is a whole architecture you cannot run at all.
The default is reassuring on that point: vector memory is optional and defaults to keyword search. So an Intel Mac user with a manual build and keyword search is plausible even though a packaged Intel build is not.
Server mode is the same runtime and the same features with a web interface, and configuration lives in a JSON file under the home directory rather than in YAML, which is a small consistency win given the project also ships a YAML example config for something else. You can inspect or change individual keys from the command line, or re-run the interactive setup with an update flag.
The compose port is loopback, and the comment explains exactly why
The deployment file contains the clearest security reasoning in the repository, and it is in a comment next to the port mapping.
The published port binds to loopback only. The comment says to expose it to the internet through the reverse proxy configuration the project ships as an example, and explicitly not by changing the bind address to all interfaces. The stated reason: in the simplest single-user path, where authentication mode is none, every API endpoint is open.
That is the right reason to give, because it is about the software rather than about the network. An agent with shell execution and file access behind an open API is a remote shell, so the protection has to be the network boundary, and the file makes that boundary deliberate rather than incidental.
The alternative auth mode is documented in the same block. In local mode the application runs its own account system: the first account to register bootstraps as administrator, then register, login and logout with email and password. The comment notes this is the standard flow and that the application gates itself, so no external proxy authentication is needed. And it says to use the no-auth mode only for a trusted single user box that is already fronted by another gate.
State handling is equally deliberate. All runtime state lives under one directory variable mapped to a single mount, described as one mount and one place to back up.
There is a script that verifies the prompts, and it runs first
The verification scripts are arranged in an order that says something about what this project considers risky.
The pre-merge verification runs the prompt contract check first, then the type check, then the unit tests, then the build, then the integration tests, then the end to end tests.
A prompt contract check running before the type checker is not an accident of ordering. In a project where the LLM decides what happens and everything else executes its decisions, the prompt is part of the interface between the two halves, and a prompt that drifts can break behaviour while every other test still passes. So it gets checked first and separately.
The test tiers are unusually well separated. There are separate configurations for unit, integration and end to end runs, each with a report variant that writes machine readable JSON to a reports directory. The continuous integration script runs unit, build, integration and end to end in that order, so the build sits between the fast and slow tiers where a broken bundle is caught before the expensive tests.
There is also a failure replay generator. A script named for generating replays from failures, alongside a watchdog entry point in the built output, suggests the project records enough to reproduce a bad run rather than only to detect it.
Twenty five skills, loaded on demand rather than pasted into context
The skills system is worth reading because of how it loads, not what it contains.
There are twenty five built in skills, described as an adaptation of the official vendor skill catalogue plus some of the project's own. And the loading model is explicit: the model sees a one line catalogue and pulls the full instructions only when a task needs them.
That is progressive disclosure, and it is the difference between twenty five skills and one unusable prompt. If every skill's full text were in the context window from the start, the agent would be carrying a large amount of text describing tasks it is not doing, and every tool description would compete with every other one for attention.
The extension path is a single file. Drop a skill file into your workspace and the model will see it in the catalogue. For a project whose stated purpose is being hackable, that is the right extension surface: one file, no registry, no plugin manifest.
The architecture diagram places skills alongside sub agents under the capabilities layer, below the reasoning layer and above the file and shell primitives. So a skill is something the agent invokes, at the same level as running a command or reading a file, rather than a layer of the system that reconfigures it.
Editorial conclusion
Use MOZI if you want an agent that operates on real files in a real repository rather than one that only answers, and if you are the kind of person who will read its source to understand how it works, since that is the stated reason it exists. The filesystem verification of deliverables is the feature worth having, because it is the difference between an agent that claims success and one that cannot. Do not install the macOS build expecting Apple trust, because it is neither signed nor notarised and the checksum is not a substitute. Four things to check before you run it. That your Node version is inside the supported window, which excludes both the newest major and one specific odd major for different reasons. That you read the compose binding before exposing it, because with the simple single-user auth mode every API endpoint is open and the port is loopback for that reason alone. That you understand what leaves your machine, since a local provider keeps it local but a hosted one sends your project content to it. And that the optional pieces really are optional, since office grade preview, web search and vector memory all degrade or default rather than block. Licence is MIT, and the last push to main is dated 7 August 2026.
Frequently asked questions
What is spytensor/openmozi?
It is MOZI, a personal agent operating system that runs entirely on your machine, described as a self hosted Codex. You pick a project folder or repository, a git branch, a permission level and a model, then give it a task, and it executes shell commands, file edits, web research and document generation. It also has long term memory across sessions, scheduled tasks, and 25 skills that load on demand.
How does MOZI know a task actually finished?
Every deliverable it claims is verified against the filesystem before it reports done, rather than emitting a completion message. That closes the failure mode where a model reports writing a file it did not write, because checking the disk is a different question with a different oracle than asking the model again.
Is the MOZI macOS app safe to install from a release download?
The macOS builds are not signed with an Apple Developer ID and are not notarised, so Gatekeeper may block them. You are told to verify the download's checksum against the published sums file before removing quarantine, and the documentation is explicit that the checksum confirms which artifact you downloaded but does not replace Apple code signing trust.
What Node versions does openmozi support?
Node 22 long term support or newer, and 26 or newer is excluded because the native dependencies have not caught up, while Node 23 is rejected by dependency engine pins. You do not need corepack or a global package manager, since the setup script resolves the pinned one itself, and it explicitly tells you to avoid the corepack enable command.
Can I run MOZI as a server instead of a desktop app?
Yes, with the same runtime and the same features behind a web interface on a local port. Supported architectures are Linux on x64 and arm64 and macOS on Apple silicon; Intel Macs are excluded because the vector database dependency does not publish a macOS x64 binary. Configuration lives in a JSON file under your home directory and can be changed from the command line.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/spytensor-openmozi)