# OpenLumara starts at four thousand tokens and lets you delete the rest

> A Python agent framework written from scratch around one idea: everything other frameworks treat as core is a module you can switch off, including shell access, memory, scheduling and the module system itself. The trade is a system prompt that starts around four thousand tokens instead of tens of thousands, and an install that clones without a target directory because the update scripts are git-based.

**Rose22/openlumara** — AI agent framework, written from scratch (not based on openclaw), focused on stripping it down to the bare necessities, optimizing token count, reducing security risks. modular so you can enable only exactly what you need.

- Repository: https://github.com/Rose22/openlumara
- Stars: 494 · Forks: 57
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/rose22-openlumara

## The system prompt is the product, and four thousand tokens is the floor

The headline claim is a number. The system prompt can be as small as around four thousand tokens with normal use, which is what makes the project viable on a local model and dramatically cheaper against hosted APIs. Everything else in the design follows from that figure.

The way the number stays low is that nothing is built in. Memory, scheduling, time awareness and token awareness are each modules that can be switched off, along with components other frameworks consider part of the core. The extreme case is stated directly: you can turn absolutely everything off to the point where the system prompt is empty and you are simply talking to the base model.

Two commands make that budget inspectable rather than theoretical. One shows the current context size in input tokens, and one shows exactly what is being sent. The readme also notes that the model can see your token use, which is a small feature that doubles as a prompt cost warning.

The dependency list is small and tells you what kind of program this is: the OpenAI client, a YAML parser, a compact binary serialiser for memory, two HTTP clients, a request helper, an identifier generator, a file type detector, a PDF reader, a regular expression engine, and two libraries whose only purpose is parsing model output.

Those last two are the most revealing entries. json-repair and partial-json-parser exist because streaming model output is frequently malformed, and shipping both means the framework expects to be repairing truncated structured output on a regular basis rather than occasionally.

## Shell access is a module, and it ships disabled

Most agent frameworks give the model a shell as part of the deal. Here it is a module, and the readme states plainly that it is disabled by default for security.

What you get when you enable it is a sandboxed shell that runs inside a Docker or Podman container, with the container under your control rather than the framework's. You choose the image, and you can cut the container off from the internet. That last option is the one that matters: a shell without a network cannot exfiltrate what it reads, which is the failure mode that turns a coding agent into an incident.

Fine grained control over the container is the stated feature, and the practical reading is that the security boundary is a file you write rather than a policy you trust. Choosing the image decides what binaries exist inside. Deciding whether the container has a network decides whether the model can phone home. Nothing in the readme describes a default posture beyond the module being off.

That combination, off by default and configurable when on, is the reason this framework describes itself as local first and security conscious. It is also why it can claim to strip things down without stripping away capability: the capability is still there, one module away.

The model can also be pointed at any OpenAI compatible backend, named as llamacpp, ollama and koboldcpp among local options and cloud providers otherwise, and the readme notes it pairs well with the first two of those.

## The module that lets the model enable modules is itself disabled

The module system has an unusual shape. You can switch components on and off three ways: a command, an edit to the config file, or by asking the model to do it.

The third way is the interesting one, and it is gated. The module that permits the model to toggle modules is disabled by default for security, so on a default install the model cannot grant itself a capability. That is a real design decision rather than a documentation flourish, because self-escalating agents are the failure mode everyone worries about, and closing it with a switch that is off by default means the safe configuration is also the default one.

The naming is slightly confusing and worth knowing before you read a config file. The module that manages modules is called modules, and the command that lists and toggles them is singular. Both appear in the same section of the documentation.

Turning modules on and off by editing a file is the most direct route and the one to reach for first, since it is the configuration you can review in a diff. The model-facing toggle is for convenience once you have decided you trust the arrangement.

Modules themselves are described as simple Python classes with a few custom functions, and a plugin download system is announced as coming later. That last point sets expectations: today you extend the framework by writing a class into the tree, not by installing one.

## A command path that skips the model exists because the model can get stuck

One command is described as bypassing the AI completely: a restart command that forces the server to restart no matter what the AI is doing.

That is a small feature with a large amount of information in it. The author expects the model to be mid-task, unable or unwilling to finish, and needs a way out that does not depend on the thing that is stuck. Every framework with a long-running agent loop eventually needs this, and most of them leave it to the terminal, which means the person has to know how to find the process and kill it.

Here it is a slash command in the same interface as everything else, which means it works from the web interface, from the terminal channel, and from any chat platform the framework is connected to. The same reasoning applies to the model switch command: the readme notes it is especially useful when you have turned the tools off, since asking an agent to switch its own model requires the capability that switching models implies.

The scheduler is the other half of the control surface. Tasks can be scheduled ahead of time, and the readme compares the design to a competitor's cron feature while stressing that it was written from scratch rather than borrowed.

Taken together, the pattern is a framework that assumes the agent will sometimes be the least reliable component in its own loop, and builds exits for that case rather than trusting the loop to converge.

## Characters replace three products and are scoped to one session

An optional characters module is positioned as a replacement for several existing character chat products by name. Once enabled, you can add, edit and remove characters, switch between them and set a user profile, either by asking the model or with a command.

The behaviour that makes it work as a drop-in is that an active character disables all other prompts, so the system prompt becomes the character and nothing else. That is the same four thousand token budget applied to a persona rather than a toolset, and it is why the character is a module rather than a layer on top: with the module on you get one prompt, with it off you get your normal configuration.

The scoping detail is the one to know about. Characters are tied to the chat session, so a character active in the web interface does not affect a Telegram session, and loading a chat that had a character active brings it back automatically. Session-bound state means the same account can behave differently in different channels without any explicit configuration.

Memory is built on the same modular footing and stores in a compact binary format rather than a text or database file, chosen for size and read speed. The model can save a memory or you can ask it to, which means the write path goes through the same permission model as everything else.

The intended use is stated as personal rather than professional: todos, notes, morning routines and habit tracking, with an explicit note about executive dysfunction. That tells you the design centre, which is a single user with one assistant rather than a team with shared state.

## The clone command has no target directory and the zip breaks updates

Installation is one line:

```bash
git clone https://github.com/Rose22/openlumara
```

No directory argument, so the checkout lands in a folder named after the repository. Then you run a shell script on Linux or a batch file on Windows, and the URL the startup prints is what you open.

The reason the readme insists on cloning rather than downloading the archive is stated immediately: the update scripts, one for Linux and one for Windows, use git, so a zip install cannot be updated by them. That is a small operational decision with a real consequence, and it is the kind of thing that is usually only discovered when someone wants to upgrade six months later.

The runtime scripts come in pairs, and there are three requirement files rather than one. Alongside the main file there is a separate list for the Matrix channel and another for Termux, and there is a third launcher script to match, which implies Android support through Termux. Splitting requirements by channel means the base install stays small and each channel pulls its own dependencies, which is the same modular philosophy applied to packaging.

First run configuration happens in the web interface rather than in a file: open the URL the startup printed, use the gear icon in the top corner to reach settings, enter an API connection and save.

## The core was hand written and most of the channels were not

The provenance statement is unusually precise and worth reading in full rather than as boilerplate. Everything in the core directory was designed and coded by hand. The author used AI to ask how to improve things and how to fix certain bugs, and asked it how to accomplish specific tasks in Python, but no code was inserted without being personally audited and modified. Certain non-core parts, and many of the channels in particular, were mostly generated and then manually audited and edited.

The conclusion drawn is that this is not a vibe coded project while still being an AI assisted one. Both halves of that claim are supported by the specifics: the core is hand written, the periphery is generated and reviewed.

Writing a channel is documented in about a paragraph. You import core, subclass the channel base class to inherit the required behaviour, and implement an asynchronous run method containing the main loop that asks for input and hands it to the model. Naming is automatic: a CamelCase class name is translated to a snake_case name that appears in the config file and everywhere else in the framework, so a class named one way is configured the other way.

Two small entries in the tree support the same picture. A rewrite plan for the web interface is committed at the repository root next to the launcher scripts, which is either useful transparency or unfinished business depending on your taste. And the readme ends with a notice that the project has no mascot or emoji identity, followed by a joke aimed at the agent framework it is written against, which tells you the competitive framing as clearly as a positioning section would.

## Conclusion

OpenLumara is worth reading if you run a model locally and cannot afford a large system prompt, because the token budget is treated as a design constraint rather than a cost to be optimised later, and the introspection commands are unusually honest about what is being sent. Two things to accept. Everything is opt in, so a bare install talks to the base model with nothing else, which means the useful configuration is the work. And the shell module runs a container you configure yourself, so the security boundary is yours to get right rather than the framework's to guarantee.

## FAQ

### How small is the OpenLumara system prompt?

The readme states it can be as little as around four thousand tokens with normal use, because memory, scheduling, time awareness and token awareness are all modules that can be switched off, along with everything else.

### Is shell access enabled by default in OpenLumara?

No. Shell access is a module that is disabled by default for security. When enabled it runs inside a Docker or Podman container with fine grained control, including the ability to cut the container off from the internet and to choose which image runs.

### Can the OpenLumara model enable its own modules?

Only if the module that permits module toggling is enabled, and that module is off by default for security. Otherwise you use the module command or edit the config file.

### How do I install and update OpenLumara?

Clone the repository with git, then run run.sh on Linux or run.bat on Windows and open the URL it prints. Updates run update.sh or update.bat, which use git, so a zip download cannot be updated by them.

### What backends can OpenLumara connect to?

Any OpenAI API compatible backend. Local options named include llamacpp, ollama and koboldcpp, and hosted providers work too. The readme says it pairs well with llamacpp and koboldcpp.

### How do I add a new channel to OpenLumara?

Subclass the channel base class from core, implement an asynchronous run method containing the main loop, and the framework derives the name by converting the CamelCase class name to snake_case everywhere including the config file.

## Sources

- [Issues](https://github.com/Rose22/openlumara/issues)
- [License: GPL-3.0](https://github.com/Rose22/openlumara/blob/main/LICENSE)
- [README](https://github.com/Rose22/openlumara/blob/main/README.md)
- [Rose22/openlumara on GitHub](https://github.com/Rose22/openlumara)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rose22-openlumara
