headlong: A Bash Agent Microharness for Persistent, Always-Running Agents
An open source agent microharness featuring persistent agency and recursive LLMs. Of bash, by bash, for bash; it's shells all the way down.
At a glance
- What is it?
- headlong is an open-source agent microharness written almost entirely in Bash. Its defining design is persistent agency: the agent keeps thinking in a self-guided inner-monologue loop even when no one is talking to it, rather than waiting for a request to trigger a session. It is alpha research software that runs shell commands, and the authors recommend running it in a dedicated container with a spend-capped API key.
- Who is it for?
- Researchers and individual developers who want to experiment with agents that maintain continuous inner-monologue loops, set their own priorities, and collaborate across Slack and Telegram will find headlong the only open-source tool that makes this model easy to try. It is not ready for production deployment: it is alpha software by the authors' own description, the architecture changes frequently, and background thinking costs $1 to $2 per hour at the documented settings.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Shell, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What headlong Is and Who It Is For
Most AI agent frameworks are reactive: an agent wakes up when a message or task arrives, handles it, and then stops. headlong's premise is that this is the wrong model for an agent that should behave like a persistent collaborator rather than a stateless service. A headlong agent is never asleep. It keeps thinking in a self-guided loop even with no external input. A Slack message, a Telegram message, or a chat message does not start a session; it arrives as one more observation in the agent's already-running thought stream, and the agent decides if and when to respond.
The target audience is researchers exploring the persistent-agency model and individual developers who want an agent with a name and personality that they can talk to over time, watch develop its own interests, and share with a small team. The project describes itself explicitly as alpha research software: the authors recommend a dedicated, spend-capped API key and a sandboxed container, and warn that frequent changes should be expected. Teams looking for a stable production agent runtime should look elsewhere.
Persistent Agency and the Inner Monologue Model
The README credits the design to a recursive language model (RLM) concept. The agent operates in a loop: write a shell command, run it, read the output, write the next shell command. This loop is the agent's thought process. Because it is expressed as shell commands running in a Bash environment, no separate tool system is needed. curl is the HTTP client, jq is the JSON processor, and any command-line tool on the system is available.
The thought-stream metaphor is structural, not just descriptive. The agent's trajectory is stored as a directed acyclic graph (DAG) of JSONL files with support for fork and merge. Every thought and action is preserved. Nothing is compacted away in place; compaction and introspection work on the same files with the same tools. Context is a projection of the trajectory: recent entries appear verbatim, and older entries are progressively summarized in tiers that act as an index the agent can retrieve when needed.
Background thinking rate backs off exponentially when nobody is interacting with the agent and resets the moment a message arrives. According to the README, at the settings the maintainers run their own agent, background thinking costs $1 to $2 per hour, depending on how quickly the agent loops and which model backs it.
Installing headlong with One Command
Installation is a single curl command that fetches and runs the installer script:
curl -fsSL https://headlong.ai/install.sh | bashThe installer interviews the user to create an agent, asks for a name and personality, and opens a dashboard where the agent's thoughts can be watched as they run. The dependencies are bash 3.2 or later, git, curl, jq, Python 3, and an API key for Anthropic, OpenAI, Gemini, or OpenRouter. Local models on any OpenAI-compatible server (llama.cpp, Ollama, vLLM, LM Studio) are also supported and require no API key. The dashboard additionally requires uv and bun or node; the installer offers to fetch those.
When Docker is available, the installer offers two sandboxed paths: run the whole agent inside a container, or install on the host with agent commands sandboxed in a container. The README recommends one of these two container paths. An unsandboxed host install is available behind an explicit confirmation but is not recommended, since the agent runs shell commands as the current user. To run the container path manually:
docker run -it --name headlong --restart unless-stopped -p 8080:8080 \
--add-host host.docker.internal:host-gateway buildpack-deps:curl \
bash -c 'curl -fsSL https://headlong.ai/install.sh | bash; exec bash'Once installed, the agent's name becomes a CLI command. For an agent named ada:
ada hello
ada dash
ada stop
ada startThe first sends a single message and waits for a reply. ada dash opens the dashboard. ada stop and ada start pause and resume the agent's thinking loop. headlong-killall stops every headlong process on the machine.
Bash-Native Architecture and the shellm Core
The core of headlong is about 11,000 lines of Bash. The design follows what the README calls Ken Thompson's philosophy of small executables that each do one thing and compose through pipes, files, and environment variables. The main components are shellm (the recursive language model loop), traj (trajectory management), llm (the LLM call wrapper), context (the context projection tool), mem (memory), and skills.
Because the agent thinks by writing shell commands, any shell tool is implicitly available as an agent capability. The README notes that no separate tool system besides Bash is needed: the model writes commands, runs them, reads the output, and writes the next commands. Subagents see their ancestor agents' trajectories, so a subagent can read why it was created, what the parent already tried, and where it fits in the larger effort.
Self-improvement is expressed as a fork-test-merge pattern: the agent forks the headlong codebase and optionally its own trajectory, changes something, runs the changed version, and merges the change back if it worked. If it did not work, the agent and its changes are discarded without needing a rollback mechanism.
Multi-Player Conversations Over Slack and Telegram
A headlong agent is designed to be shared. A whole team can talk to a single agent instance over Slack, Telegram, and a built-in chat application. All conversations land in the agent's single thought stream rather than in separate per-user sessions. The README notes that this has an explicit implication: assume anything you tell the agent is shared with everyone who talks to it. There are no hard privacy walls between users.
The agent can follow what different people are working on, identify connections between their work, and send messages to whoever seems most relevant without being directly addressed. The README describes this as the agent behaving more like a person than a service, which is the design goal. The Slack and Telegram integrations are included in the top-level slack/ and telegram/ directories of the repository.
Alpha Status, Running Costs, and What headlong Is Not
The README includes an important warning: headlong is alpha research software, and the authors expect frequent changes. The project has no GitHub releases, and the last push was on 2026-09-26. There is no stable API surface to program against. Running it means accepting that updates may break an existing agent configuration.
The ongoing API cost is a real constraint. Background thinking at the maintainers' recommended settings costs $1 to $2 per hour continuously. Over a month, that is $720 to $1,440 before any interaction-driven usage. The exponential backoff reduces cost when the agent is idle, but the agent never fully stops. Using a local model (llama.cpp, Ollama, or similar) eliminates the per-token cost but requires local hardware capable of running the model.
headlong is not a tool for building production AI pipelines, API endpoints, or data-processing workflows. It is specifically about the experiment of running an always-on agent with persistent memory and continuous self-directed thought. Teams who want a task-driven agent that they trigger on demand and pay only for the work it does will not get that from headlong's persistent-agency model.
Compared to AutoGen and Other Task-Driven Agent Frameworks
AutoGen is an open-source multi-agent framework from Microsoft Research that coordinates multiple agents around solving a specific task. Agents in AutoGen communicate through structured messages and are active only while processing a task. When the task completes, the agents stop.
headlong's architectural difference is fundamental: there is no task boundary. The agent runs a continuous inner monologue, generates its own sub-goals, and persists those goals between interactions. AutoGen is the right choice when the goal is to have multiple agents collaborate to solve a defined problem and then stop. headlong is for the experiment of running an agent that behaves like a persistent participant in an ongoing context, setting its own agenda and reacting to human messages as interruptions to that agenda rather than as the only trigger for action.
Editorial conclusion
Researchers and individual developers who want to experiment with agents that maintain continuous inner-monologue loops, set their own priorities, and collaborate across Slack and Telegram will find headlong the only open-source tool that makes this model easy to try. It is not ready for production deployment: it is alpha software by the authors' own description, the architecture changes frequently, and background thinking costs $1 to $2 per hour at the documented settings. Before running it, use the Docker container path (docker run -it --name headlong) and set a spend cap on the API key, since the agent runs shell commands and generates model calls continuously.
Frequently asked questions
What does headlong agent software do?
headlong runs an AI agent in a continuous self-guided thinking loop: the agent writes shell commands, runs them, reads the output, and repeats. A message from a user lands in the agent's thought stream as one more observation; it does not start a new session.
What does headlong cost to run?
The README states that at the settings the maintainers run their own agent, background thinking costs $1 to $2 per hour. The cost depends on the agent's loop rate and the model used. Using a local model via llama.cpp, Ollama, or a similar OpenAI-compatible server eliminates the per-token API cost.
Can headlong be run without Docker?
Yes, but the README does not recommend it. An unsandboxed host install is available behind an explicit confirmation, but the installer warns that the agent's commands run directly on the host machine as the current user. The Docker path (full container or host install with container-sandboxed commands) is the recommended approach.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/laude-institute-headlong)