# Open Executive: an AI executive team you run yourself

> OpenExecutive from SenteLabsAI wraps eight specialist agents behind one executive voice, with ChromaDB retrieval, SQLite memory and a single-instance scheduler. Here is how to install it, what it does well, and where the design puts limits on you.

**SenteLabsAI/OpenExecutive** — AI-powered virtual executive team — a single coherent executive persona backed by 8 specialist agents (FastAPI + Next.js).

- Repository: https://github.com/SenteLabsAI/OpenExecutive
- Website: https://openexec-ui-dev.fly.dev
- Stars: 5,423 · Forks: 580
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/sentelabsai-openexecutive

## What Open Executive is for, and who it is not for

Most people asking an LLM for business advice get a generic answer, because the model has no view of their company and no memory of last month's decision. Open Executive addresses that gap by putting a retrieval layer and a persistent memory behind a single executive persona. The README describes eight specialists: Chief Strategy Officer, CFO, Chief HR/People Officer, General Counsel, Chief Operating Officer, Chief Marketing Officer, Chief Product Officer, and a Board Communications Director. The stated audience is a founder or operator who wants a senior advisor customized for a specific business, not a chatbot with a business-sounding prompt.

The design decision worth noticing is that the internal agent architecture is never exposed to the user. Every answer arrives in one voice, even when several specialists were consulted. That is a deliberate product choice, and it cuts both ways: you get a coherent advisor, but you lose the ability to see which domain produced a claim. If your workflow requires tracing an assertion back to the specialist that made it, the audit log and the eval harness are the places to look, not the chat transcript.

## The orchestrator, the two retrieval layers, and the memory pass

The README's architecture diagram is short enough to quote in structure. A user message reaches the Executive Orchestrator, which runs on claude-sonnet-5. The orchestrator uses tool calls to fan out to specialists in parallel. Each specialist retrieves context from ChromaDB, then the orchestrator synthesizes one response. The default model for the Executive and most specialists is claude-sonnet-5; the CSO, CFO, GC and Board roles can run on claude-opus-5 with extended thinking for deeper reasoning.

Retrieval is split into two layers per specialist call. The first is built-in MBA-level Markdown under knowledge/builtin/, which is git-tracked and seeded into ChromaDB at startup. The second is your own uploaded company documents, chunked into a separate collection named company_docs. The README states that RAG context is injected into the user turn, never into the cached system prompt. That separation matters for cost: the persona, company profile and knowledge index are cached separately, and the README cites a cache hit rate of up to 85 percent after the first few turns. Any dynamic content placed in a cached block would break that, so the constraint is architectural rather than cosmetic.

Memory works in a second pass. After every response, a background job on claude-haiku-4-5 extracts decisions, initiatives and advice into SQLite. The next session opens with a past_decisions block, so the Executive can refer to what it recommended previously. This is episodic memory, not a fine-tune: the model weights never change, and the recall depends entirely on what the extraction pass chose to keep. If the extractor drops a decision, the next session will not know about it.

## Installing Open Executive and getting a first answer

The README's Quick Start assumes Python 3.11+, Node 22+, and an Anthropic API key. Clone the repository, copy the environment template, and put your key in it. The .env.example file documents the required variable as ANTHROPIC_API_KEY, with a placeholder value of sk-ant-your-key-here.

```bash
git clone https://github.com/SenteLabsAI/OpenExecutive.git
cd OpenExecutive
cp .env.example .env
# Edit .env and add ANTHROPIC_API_KEY=sk-ant-...
```

All configuration lives in that repo-root .env. The Makefile comment explains why both dev recipe lines source it into the process environment: Next.js only auto-loads packages/ui/.env*, so Auth.js would otherwise never see AUTH_SECRET, and the API reads BACKEND_SHARED_SECRET and BACKEND_ALLOWED_ORIGINS from os.environ rather than through pydantic Settings. The same comment warns that .env values must be shell-safe, so quote anything containing spaces or a dollar sign.

```bash
make dev
```

make dev starts uvicorn on port 8000 and the Next.js UI on port 3000. Open http://localhost:3000 to start chatting. The README warns that the first run is slow: uv sync pulls ChromaDB and sentence-transformers with PyTorch, and the first boot downloads an embedding model of roughly 90 MB to build the local vector index. Subsequent starts are fast. If you are not using make, the contributor path is to run uv sync inside packages/core, activate the virtualenv, and start the API directly.

```bash
cd packages/core
uv sync
source .venv/bin/activate
uvicorn openexecutive.api.main:app --reload --port 8000
```

For the web UI's Google sign-in you also need the AUTH_* block filled in. The .env.example notes that AUTH_SECRET can be generated with openssl rand -base64 32, that AUTH_GOOGLE_ID and AUTH_GOOGLE_SECRET come from a Google Cloud OAuth client, and that the authorized redirect URIs must include http://localhost:3000/api/auth/callback/google plus the equivalent path on your deployed host. ALLOWED_EMAILS is a comma-separated allowlist: any Google account not on it is rejected even with a valid login.

## The scheduler runs once, and that is a deployment constraint

The built-in scheduler is the part most likely to surprise an operator. It claims due actions with an UPDATE ... RETURNING statement to prevent two workers from firing the same job. That technique protects against double-firing within a single database, but the README is explicit that the API must run as a single instance, and that you should not horizontally scale it without gating the scheduler first. For a team on Fly.io, where scaling to two machines is a config change away, this is a real boundary. The repository does ship separate Fly configs for dev and QA (fly.api.toml, fly.ui.toml, fly.api.qa.toml, fly.ui.qa.toml) plus an optional fly.honcho.toml for a memory app, so multi-environment deployment is contemplated. Multi-instance API deployment is not.

The second operational detail is the backend gate. BACKEND_SHARED_SECRET is a shared secret between the Next.js proxy and FastAPI: the UI attaches it as an x-api-key header, and the API requires it on every non-exempt request. The .env.example says to leave it unset for make dev, which means the gate is off locally, and to set the same random value on both apps in any deployed environment. Generate it with openssl rand -hex 32. Because the API reads this from os.environ rather than through pydantic Settings, a dotenv file alone will not surface it; the Makefile comment calls this out directly, noting that the failure mode is a silently open API gate.

## Where Open Executive is the wrong tool

Three cases stand out. First, if you need to scale the API horizontally for load, the single-instance scheduler constraint means you must either gate the scheduler behind a separate process or accept one instance. The README does not document a supported way to run the scheduler separately, so this is an open design question rather than a documented workaround.

Second, if you cannot or will not use the Anthropic API, the .env.example states that ANTHROPIC_API_KEY is required unless you run entirely on local or OpenRouter models, controlled by LOCAL_MODELS_ENABLED and OPENROUTER_ENABLED. At least one provider must be configured or the app refuses to start. The README's architecture section, however, describes the orchestrator and specialists in terms of specific Claude models, and the prompt caching design is described in terms of cached system prompt blocks. Whether the local and OpenRouter paths preserve the same caching behavior is not documented in the README, so treat that as something to verify in the code before you plan around it.

Third, if you need per-specialist attribution in the output, the product deliberately hides it. The README says the internal agent architecture is never exposed to the user. That is the point of the single executive voice, but it means the chat interface is not an audit surface. The repository includes an audit/ package and an evals/ directory with an LLM-as-judge runner, which suggests the intended path for verification is the eval harness rather than reading transcripts.

## How it differs from a general assistant with a document index

The closest alternative is a general chat assistant with a retrieval index over your files, which many teams already have. The difference is in the routing. A general assistant retrieves documents relevant to the question and answers in one pass. Open Executive routes the question to named domain roles, has each of them retrieve separately, and then synthesizes. That gives the CFO role a different retrieval context than the General Counsel role for the same question, which is the mechanism behind the multi-perspective answer. It also means more model calls per turn, and the parallel fan-out described in the architecture diagram is the reason prompt caching is treated as a first-class concern rather than an optimization.

A second difference is memory. A retrieval-only assistant starts fresh each session unless you paste context in. Open Executive writes extracted decisions to SQLite and reopens with a past_decisions block. Whether that memory is an asset or a liability depends on the extraction quality, and the README does not describe a review step before a decision is persisted. For a workflow where you want to approve what the system remembers, that is a gap.

## Licence, upgrade cost and what the repository does not say

The README carries an Apache 2.0 badge and the tech stack table lists Apache 2.0, but the repository metadata reports the licence as NOASSERTION, meaning GitHub could not classify the LICENSE file automatically. Check the LICENSE file in the repository before you rely on the badge. Apache 2.0 is a permissive licence with an explicit patent grant, but this is a description of the licence text, not legal advice, and the discrepancy between the badge and the metadata is worth resolving with whoever handles licensing on your side.

On upgrade cost, the repository has no retrieved releases, so there is no published version history to plan against. CHANGELOG.md exists at the top level, and that is where upgrade notes would live. The practical upgrade risks are the ones the README already flags: the first boot downloads an embedding model and builds a local ChromaDB index, so a change to the embedding model or the chunking path could require rebuilding that index. Episodic memory lives in SQLite, and the README does not describe a migration path for the memory schema. Back up the SQLite database and the ChromaDB directory before pulling a new revision.

## Conclusion

Adopt Open Executive if you want a self-hosted executive advisor with retrieval over your own documents and memory that persists between sessions, and you can run one API instance behind Google sign-in. Do not adopt it if you need horizontal scaling of the API, or if you are unwilling to hold an Anthropic API key, since the app refuses to start without a provider. Before committing, verify that your .env sets ANTHROPIC_API_KEY and a real AUTH_SECRET, that ALLOWED_EMAILS contains exactly the accounts you intend to admit, and that your deployment sets AUTH_URL to the public origin so the post-callback redirect does not land on 0.0.0.0.

## FAQ

### How do I install Open Executive and start it locally?

Clone the repository, copy .env.example to .env, add your ANTHROPIC_API_KEY, then run make dev. The API starts on port 8000 and the UI on port 3000, and the first run is slow because uv sync pulls ChromaDB and PyTorch while the first boot downloads an embedding model of about 90 MB.

### Does Open Executive need an Anthropic API key to run?

The .env.example states that ANTHROPIC_API_KEY is required unless you run entirely on local or OpenRouter models, controlled by LOCAL_MODELS_ENABLED and OPENROUTER_ENABLED. At least one provider must be configured or the app refuses to start.

### Can I run more than one instance of the Open Executive API?

The README states that the API must run as a single instance because the scheduler claims due actions with UPDATE ... RETURNING, and it says not to horizontally scale without gating the scheduler first. No separate scheduler process is documented.

## Sources

- [Issues](https://github.com/SenteLabsAI/OpenExecutive/issues)
- [Project website](https://openexec-ui-dev.fly.dev)
- [README](https://github.com/SenteLabsAI/OpenExecutive/blob/main/README.md)
- [SenteLabsAI/OpenExecutive on GitHub](https://github.com/SenteLabsAI/OpenExecutive)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sentelabsai-openexecutive
