The repository is collab-agent and the software inside it is called OmniAct
A lightweight multi-agent system
At a glance
- What is it?
- A multi-agent system that splits one goal across three specialised workers, one driving Edge, one driving Windows applications, and one driving a Codex CLI session, coordinated by a leader that only dispatches work with clear boundaries. The collaboration protocol is a directory of files, and the browser worker's default session mode copies your Edge profile.
- Who is it for?
- OmniAct suits someone whose real work crosses a browser, a desktop application and a terminal, and who is willing to treat a task directory as the audit trail. Before running it, decide whether inherited browser sessions are acceptable on the machine in question, since that default copies your logged-in profile into an automated browser, and read the licensing note about modified open-source components, since the public repository is scoped to the runtime core only.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 32 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two names, and the repository says almost nothing
The repository is named for a generic description of multi-agent systems. The software inside it is called something else entirely. Every heading, every configuration variable and the example environment name use that other name, while the repository's own description is reduced to a handful of words about a lightweight multi-agent system.
That mismatch is more than cosmetic in practice. A person who finds the repository, reads the configuration file and searches for the package name will not find the project. The documentation is also split by language in the same way the search results are, with the primary readme in Simplified Chinese and a separate English one, so an English-only searcher has to guess which file to open.
There is no recorded homepage, no published release and no tag. The only dated signal available to a reader is the last push, which is 2026-08-31.
None of that makes the project worse. It does mean that judging it requires reading the repository rather than the repository listing, which is an unusual amount of work before you have installed anything.
Windows, Edge, Python 3.11 and two separate model endpoints
The requirements list is short and specific. A Windows environment with Microsoft Edge installed. Conda or another Python environment manager. Python 3.11. And then two model configurations rather than one: an OpenAI-compatible endpoint with native tool calling for the leader, and a separate OpenAI-compatible endpoint with vision capability for the browser worker.
Two endpoints is the requirement worth pausing on. The design splits reasoning from perception, so a deployment that wants a strong reasoning model and a cheap vision model can have both, and a deployment that wants one model has to configure the same endpoint twice.
The third worker needs no model of its own. It drives the Codex CLI, which must be installed separately.
Setup is a conda environment, a requirements install and a copy of the example environment file:
conda create -n omniact python=3.11
conda activate omniact
python -m pip install -r requirements.txt
Copy-Item .env.example .envRunning it is a single example script, and the project's own verification is two standard-library commands, one that byte-compiles the source, tests and examples, and one that runs the test suite with the built-in test runner. No third-party test framework appears in the dependency list, which is consistent.
The default browser session mode copies your Edge profile
One environment variable decides whether the browser worker starts from scratch or from your session, and the default is to start from yours.
The setting takes two values. The clean mode gives the worker a fresh browser with no history from your account. The inherited mode copies the current Edge profile into the worker's own isolated browser instance, so the worker arrives already logged in to whatever you are logged in to.
The stated benefit is real. An agent asked to do something inside a site you already use would otherwise stop at a login wall, and asking a human to log in on every run defeats the purpose of automation. The stated limitation is also real: the copy is of the profile, which is where cookies and stored credentials live, not of the window, so the browser you are using is not taken over but your session is duplicated into something a model is driving.
The distinction the documentation draws is worth repeating carefully. Isolated means the worker does not fight you for the same window. It does not mean the worker is anonymous.
If the machine in question holds sessions to anything sensitive, that default is the setting to look at first, and it is a single line in the environment file.
The collaboration protocol is a directory of files
There is no message bus and no shared memory here. The handoff mechanism is a task directory, laid out like this:
tasks/<task-id>/
├─ task.md
├─ trace.jsonl
├─ trajectory.md
├─ workers/
│ └─ <worker-name>/
│ ├─ request.md
│ ├─ result.md
│ └─ artifacts/
└─ final/
└─ result.mdEach worker owns its execution environment and writes its handoff into that tree. Workers do not need to share internal state and do not directly depend on one another, which means a worker can be paused, inspected or re-run without the others knowing.
The result format is part of the design rather than an afterthought. The stated rule is that workers hand off through a written summary plus concrete artifacts rather than returning a sentence saying the task is done, and the directory makes that enforceable because there is a place the artifact has to land.
The trace file is the other half of the value. A JSON Lines event log plus a readable trajectory document means a run can be reviewed after the fact without instrumenting anything, and the parallel subtask and resumable session features exist to keep long tasks alive across that boundary.
For a human takeover, the example program exposes three commands taking a worker identifier: take it over, resume it, or cancel it.
Thirty-two dependencies, twenty-four of them pinned exactly
The requirements file lists thirty-two packages and mixes two pinning styles in one file without comment. Twenty-four are pinned to an exact version. Eight carry only a floor.
Exact pins on a project with no lock file committed for its own environment means a rebuild six months from now produces a different stack than the one that was tested, for two thirds of the dependencies. The eight that float, including the terminal output library, the YAML parser, the dataframe library, a PDF library, a fuzzy matching library, a web framework and the Windows binding library, drift in the other direction.
The contents are worth reading past the count. There is a full document-processing stack: a spreadsheet library, two PDF readers, a Word document parser, a PDF writer, a dataframe library and an imaging library. There is a Windows automation stack: two UI automation libraries, a COM bridge, the Windows binary binding, a screen automation library and a display geometry library.
Then there are three that do not obviously belong to a browser-and-desktop agent. A one-time password library, suggesting the tool handles second-factor codes somewhere. A web framework, with no documented purpose in the readme. And an ASCII art library.
For a project described as lightweight, thirty-two dependencies including a document pipeline is a stack worth questioning before you install it.
The CLI worker is sandboxed and capped, the browser worker is capped in neither
The two optional workers are disabled by default and configured with a short block each.
The desktop worker takes an enable flag, its own model endpoint and a maximum step count of fifty. Its configuration comments identify it as a wrapper around a Windows desktop automation project with a different name, and the project also ships a third-party notices file for modified open-source components, which fits a worker built on top of someone else's automation library.
The command-line worker also takes an enable flag, plus the path to the Codex executable, a working directory, and a maximum turn count of three. It also sets a sandbox mode for that executable to a workspace-writable scope, which is the single most reassuring line in the whole configuration file: the coding agent it delegates to is confined to writing inside a workspace rather than running with full access.
Now compare the browser worker. It has an endpoint and a session mode, and no step or turn budget anywhere in the environment file.
So the two workers that can act on your machine most directly are capped and sandboxed, while the worker with no explicit budget is the one that inherits your login session. Nothing in the visible documentation says whether the browser worker's iteration limit lives elsewhere, such as inside the execution engine it is said to delegate to.
Recovery is on the roadmap, and the benchmark harness already exists
The roadmap is five items, and they are candid about what is unfinished.
The first is optimisation of the browser and desktop workers driven by benchmarks. The second is better reliability for human takeover and for recovery after a partial failure, which is the same thing the takeover, resume and cancel commands were designed to support and the same thing the readme describes as still in development. The third is scheduled tasks for periodic digital work. The fourth is a local console for submitting tasks and viewing artifacts, which would formalise what is currently a directory you browse yourself. The fifth is more execution engines and reusable skills.
The benchmark item is the interesting one, because the groundwork is already in place. There is a benchmarks directory at the top level, and the example environment file carries an optional benchmark workspace path kept deliberately outside the Git repository, plus a data path for a natural language toolkit that the requirements list does not include.
The open source note is unusually clear about scope. The public repository keeps the runtime core focused on task orchestration, worker boundaries, task artifacts and readable execution traces, with a larger product interface to be added gradually once the execution flow is stable. Released under the MIT licence.
Read that as a promise and a warning at once: what is here is the orchestration layer, and the interface a user would want is on the way.
Editorial conclusion
OmniAct suits someone whose real work crosses a browser, a desktop application and a terminal, and who is willing to treat a task directory as the audit trail. Before running it, decide whether inherited browser sessions are acceptable on the machine in question, since that default copies your logged-in profile into an automated browser, and read the licensing note about modified open-source components, since the public repository is scoped to the runtime core only.
Frequently asked questions
What is OmniAct in the collab-agent repository?
It is a multi-agent system that takes one goal and dispatches it across three specialised workers: one driving a browser, one driving Windows applications, and one driving a Codex CLI session, with a leader agent deciding whether to reason directly or delegate.
What does the collab-agent system need to run?
A Windows environment with Microsoft Edge, Conda or another Python environment manager, Python 3.11, two OpenAI-compatible model endpoints (native tool calling for the leader and vision for the browser worker), and the Codex CLI if you use the command-line worker.
What does BROWSER_WORKER_SESSION_MODE do?
It takes inherited or clean. Inherited, the default, copies the current Edge profile into the worker's isolated browser so it reuses your login state without taking over the window you are using. Clean starts the worker without your profile.
How do the OmniAct workers hand off to each other?
Through a task directory on disk containing the task description, a JSON Lines trace, a readable trajectory, and per-worker request, result and artifact folders, plus a final result. Workers share no internal state and do not depend on each other, and hand off written summaries plus concrete artifacts instead of a completion message.
Is the Codex CLI worker restricted in OmniAct?
Yes. Its configuration sets a workspace-writable sandbox mode for the Codex executable, along with a working directory and a maximum of three turns, and it is disabled by default. The optional desktop worker is similarly disabled by default and capped at fifty steps.
How does collab-agent verify itself?
With two standard-library commands: one that byte-compiles the source, tests and examples, and one that runs the test suite using the built-in unittest discovery. No third-party test framework appears in the dependency list.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/iwtbs27-collab-agent)