# GenericAgent: a 3K-line self-evolving agent that grows a skill tree instead of shipping one

> GenericAgent (lsdefine/GenericAgent) is a minimal Python agent framework whose core is roughly 3,000 lines plus nine atomic tools. It writes new skills as it solves tasks, and the README claims a context window under 30K tokens. Here is what the repository documents, and where the design runs out.

**lsdefine/GenericAgent** — Self-evolving agent: grows skill tree from 3.3K-line seed, achieving full system control with 6x less token consumption.

- Repository: https://github.com/lsdefine/GenericAgent
- Website: https://github.com/lsdefine/GenericAgent
- Stars: 14,273 · Forks: 1,663
- Language: Python
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/lsdefine-genericagent

## The problem GenericAgent picks, and who it is aimed at

Most agent frameworks ship a fixed tool set and a fixed prompt. GenericAgent starts from the opposite position. The README states the design philosophy directly: "don't preload skills, evolve them." The seed is about 3,000 lines, nine atomic tools and an agent loop of roughly 100 lines. Every time the agent completes a task it has not done before, the execution path is written back as a reusable Skill, so the capability set is a byproduct of use rather than a configuration file.

The intended user is a single person on a single machine who wants an LLM to operate that machine: the browser, the terminal, the filesystem, keyboard and mouse input, screen vision, and Android devices through ADB. The README's own demo list is consumer-flavoured rather than enterprise-flavoured. It shows ordering milk tea in a delivery app, screening stocks for an EXPMA golden cross with turnover above 5 percent, pulling expenses over 2,000 yuan from Alipay through ADB, and sending bulk WeChat messages. These are tasks that live in a logged-in GUI, not behind an API.

That framing matters when you evaluate the project. GenericAgent is not trying to be an orchestration layer for a fleet of services. It is trying to be the smallest thing that can sit in front of a real desktop and get things done, and then remember how.

## Nine atomic tools, a real browser, and a 100-line loop

The architecture is deliberately flat. The repository root holds agent_loop.py, agentmain.py, llmcore.py, TMWebDriver.py, simphtml.py and hub.pyw, alongside memory/, reflect/, plugins/ and ga_cli/. The README describes the loop as roughly 100 lines and the tools as nine atomic operations covering browser, terminal, filesystem, input, screen vision and ADB.

The browser piece is the one worth understanding before you commit. TMWebdriver injects into a real browser rather than launching a fresh headless instance, which preserves login sessions. The README's demo section shows why that choice was made: while configuring a Discord bot, an hCaptcha challenge appeared mid-task and the real browser session passed it, letting the task continue. A headless driver that starts from a clean profile would have stopped there. The cost is that the agent depends on a browser state you maintain by hand, and the README does not document how sessions are isolated between tasks.

Token consumption is the other stated differentiator. The README claims a context window under 30K tokens and contrasts it with the 200K to 1M range it attributes to other agents. The argument given is signal-to-noise: less context means fewer hallucinations and lower cost. That claim is a design target from the README, not a measured figure, and the repository does not include a benchmark table supporting it.

Compatibility is broad on the model side. The README lists Claude, Gemini, Kimi and MiniMax, and the root contains mykey_template.py and mykey_template_en.py as the configuration entry points.

## Installing GenericAgent and running the terminal UI

The README is explicit about the Python version, and it is the first thing to get right: use 3.11 or 3.12, and do not use 3.14 because it is incompatible with pywebview and other dependencies. Note that pyproject.toml declares requires-python = ">=3.10,<3.14", which allows 3.10 even though the README recommends 3.11 or 3.12.

The recommended path is a clone plus an editable install with the ui extra, then copying the English key template into place. Run these three commands from a shell:

```bash
git clone https://github.com/lsdefine/GenericAgent.git && cd GenericAgent
uv venv && uv pip install -e ".[ui]"
cp mykey_template_en.py mykey.py
```

After the copy, open mykey.py and fill in your LLM API key. The dependency split is intentional: the core needs requests plus beautifulsoup4, bottle, simple-websocket-server and aiohttp for TMWebdriver's local server. The ui extra adds Streamlit, pywebview, textual, prompt_toolkit, rich and pillow. The pyproject.toml comment says to choose dependencies by OS and need rather than installing everything, and that missing packages can be installed on demand.

Two frontends are documented. The terminal UI is the recommended one:

```bash
python frontends/tui_v3.py
python launch.pyw
```

The first command starts the scrollback-first TUI built on prompt_toolkit and rich, which the README says supports multiple concurrent sessions and real-time streaming. The second starts the Streamlit web UI. On Windows the README warns that TUI rendering can be flaky and suggests Git Bash over PowerShell or cmd.

There is also a one-line installer that builds a self-contained directory with an isolated Python environment and Git. The script lives in assets/ if you want to read it before running it. On Linux or macOS:

```bash
GLOBAL=1 bash -c "$(curl -fsSL https://raw.githubusercontent.com/lsdefine/GenericAgent/main/assets/ga_install.sh)"
```

The README adds a warning that is easy to skip past: GenericAgent grows its environment through the agent itself, so you should not pre-install everything. The first real use is therefore not a smoke test of a fixed feature. It is handing the agent a task and letting it install what it needs along the way.

## The self-bootstrap claim and what it does not prove

The README makes a strong statement: everything in the repository, from installing Git and running git init to every commit message, was completed autonomously by GenericAgent, and the author never opened a terminal. That is a claim about the project's own development process, and it is the kind of claim a reader cannot verify from the repository contents alone. Treat it as a demonstration of intent rather than as evidence of reliability on your machine.

What the claim does tell you is where the project expects friction. If an agent can bootstrap its own toolchain, then the interesting failure mode is not a missing dependency. It is a bad skill. The README describes skills as crystallized execution paths, and the memory/ and reflect/ directories at the repository root suggest where that state lives, but the README does not document a review step, a diff view, or a way to roll back a skill that encoded a wrong approach. In a framework whose whole premise is that capability accumulates over time, the absence of a documented rollback path is the gap I would want closed before running it unattended.

## Where GenericAgent is the wrong tool

Skip it if you need process isolation between users. GenericAgent's stated purpose is system-level control over a local computer, and the tools include keyboard, mouse, screen and ADB. That is a single-operator design. The README does not describe a sandbox, a permission model, or per-session resource limits, and none of the top-level entries suggest one.

Skip it if you need a stable plugin contract. The plugins/ directory exists, but the README's extension story is skill crystallization, which produces artifacts from task execution rather than a versioned API with compatibility guarantees. Version 0.1.0 in pyproject.toml is honest about the maturity level.

Skip it if you cannot keep a logged-in browser around. TMWebdriver's value comes from injecting into a real session. If your environment is a container with no display and no persistent profile, you lose the feature that the README's CAPTCHA demo is built on.

Finally, consider the maintenance picture. The repository is not archived, and the last push was on 2026-08-24. The desktop-portable release line begins at v0.1.5 on 2026-07-11 and reaches v0.2.0 on 2026-08-24, so the release history is roughly six weeks deep. That is a young project, and the README's own roadmap is the place to check before betting a workflow on it.

## How it differs from LangChain-style orchestration

The README names LangChain in a single line about dependencies: no Playwright, no LangChain, no browser binaries to download. The difference in approach is worth spelling out, because it is not just a dependency count.

LangChain-style frameworks compose. You assemble chains, tools and retrievers from a library, and the behaviour of the system is defined by the graph you build. The framework's job is to make composition expressible. GenericAgent inverts that. It ships nine atomic tools and a short loop, and it expects the composition to be discovered at runtime and then stored as a skill. The framework's job is to make discovery cheap.

That has a concrete consequence for debugging. In a composed framework you can read the graph. In GenericAgent you read the accumulated skills, and their quality depends on how well previous tasks were solved. It also changes the dependency profile: the core needs requests plus four lightweight packages, and there is no browser binary to download because TMWebdriver uses the browser you already have.

If you want a declarative pipeline with explicit stages, GenericAgent is the wrong shape. If you want a small loop that learns your desktop, the shape is the point.

## Licence and the cost of staying current

GenericAgent is MIT licensed, stated in both the LICENSE file and the license field of pyproject.toml. That is permissive: you can use, modify and redistribute it, including commercially, provided the copyright notice and permission notice are retained. This is a description of the licence text, not legal advice, and the README adds a note that GitHub plus gaagent.ai are the official channels and that DintalClaw is the sole authorized commercial partner. If you plan to redistribute, read the LICENSE file yourself.

Upgrade cost is where the design helps and the release cadence does not. The dependency set is small and tiered, so a version bump rarely drags in a large transitive tree. The optional extras are split by frontend, and pyproject.toml comments tell automated installers to pick by need rather than install all. On the other hand, the desktop-portable line moved from v0.1.5 to v0.2.0 in about six weeks, which is a fast enough clip that pinning a version is reasonable. The README does not document a migration guide between releases, so plan on reading the release notes for each bump.

## Conclusion

Adopt GenericAgent if you want a small, readable Python agent core that drives a real browser session and a local filesystem, and you are comfortable letting it write its own skills over time. Do not adopt it if you need a sandboxed multi-tenant runtime, a stable plugin API, or a project with a long release history: the desktop-portable line only starts at v0.1.5 in July 2026, and the repository does not document rollback of a bad skill. Before installing, verify that your Python is 3.11 or 3.12, because pyproject.toml allows >=3.10,<3.14 while the README explicitly rules out 3.14.

## FAQ

### What is GenericAgent?

GenericAgent is a minimal, self-evolving autonomous agent framework written in Python. Its core is about 3,000 lines with nine atomic tools and an agent loop of roughly 100 lines, and it gives an LLM system-level control over a local computer covering browser, terminal, filesystem, input, screen vision and ADB.

### How do I install GenericAgent?

The README recommends cloning the repository, running uv venv and uv pip install -e ".[ui]", then copying mykey_template_en.py to mykey.py and filling in your LLM API key. Use Python 3.11 or 3.12; the README states that 3.14 is incompatible with pywebview and other dependencies.

### What models does GenericAgent support?

The README lists Claude, Gemini, Kimi and MiniMax among the supported major models, and describes the framework as cross-platform. Model credentials are configured through mykey.py, created from one of the mykey_template files at the repository root.

### Is GenericAgent an AI agent like ChatGPT?

They are different kinds of software. ChatGPT is a conversational model you talk to; GenericAgent is a framework that drives a local computer through nine atomic tools, including a real browser session via TMWebdriver, and writes new skills as it completes tasks.

### How does GenericAgent reduce token consumption?

The README states that GenericAgent keeps its context window under 30K tokens, which it contrasts with the 200K to 1M range it attributes to other agents, arguing that less noise means fewer hallucinations and lower cost. The repository does not include a benchmark table supporting the comparison.

## Sources

- [Official documentation](https://github.com/lsdefine/GenericAgent)
- [Official README](https://github.com/lsdefine/GenericAgent#readme)
- [Project repository](https://github.com/lsdefine/GenericAgent)
- [Release notes](https://github.com/lsdefine/GenericAgent/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lsdefine-genericagent
