Model or dataset
OpenBMB/XAgent avatar
OpenBMB/XAgent

XAgent: An Autonomous Agent for Complex Task Solving

An Autonomous LLM Agent for Complex Task Solving

8,549 stars903 forksPythonApache-2.0

At a glance

What is it?
An open-source LLM agent that plans and executes tasks by using tools (file editing, Python code execution, web browsing, shell commands). XAgent runs in Docker with a Dispatcher that routes tasks, a Planner that breaks them into steps, and an Actor that executes them.
Who is it for?
Evaluate XAgent if you need to automate multi-step technical tasks (data analysis, code generation, research workflows) where human guidance is acceptable and cloud API costs are budgeted. The project is experimental; the last push was 2026-07-31 and releases are infrequent, so check the status before relying on it for production work.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 61 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The three-part design

XAgent divides task execution into three components. The Dispatcher receives a task and decides which agent should handle it, allowing multiple specialized agents to coexist in the system. The Planner breaks the task into subtasks and milestones, generating a sequence of steps rather than a single monolithic prompt. This decomposition is critical for complex tasks because it keeps the context window manageable; a 100-step task would exceed token limits if posed as one LLM prompt, but split into subtasks with milestone markers allows the Planner to revise the plan as the Actor reports results. The Actor executes each step by selecting and invoking tools, such as writing a file, running Python code, or querying a website. Tool outputs are captured and fed back to the Actor for the next decision. This separation allows the system to reason about long tasks without hitting token limits on a single LLM call. Each component can be extended: you can add new agents to the Dispatcher, modify the Planner's decomposition strategy, or add new tools to the Actor's toolkit. The architecture is designed to scale from simple tasks (a web lookup) to complex ones (multi-day data pipelines).

ToolServer: the sandbox environment

ToolServer is a Docker container that provides a safe runtime for XAgent to execute arbitrary commands without risking the host system. It avoids security risks by isolating all actions inside the container. The File Editor tool writes and reads files in the workspace, supporting text file creation and modification. The Python Notebook tool runs Python 3.x code, allowing data manipulation, visualization, and testing of ideas without leaving XAgent's execution loop. The Web Browser tool fetches and parses web pages, handling search queries and link following by executing HTTP requests and parsing HTML responses. The Shell tool executes bash commands, capable of installing packages (apt-get, pip), running services, or manipulating files with full shell capabilities. The Rapid API tool retrieves and calls APIs from the RapidAPI marketplace, giving XAgent access to hundreds of external services including data APIs, weather services, and business tools. All tool outputs are captured and fed back into the LLM's context, closing the planning-execution loop. The requirements.txt lists dependencies including FastAPI, SQLAlchemy, Pydantic, Redis, and MySQL for the full deployment stack. The docker-compose.yml orchestrates multiple services: ToolServerManager (port 8080) dispatches tool calls, ToolServerNode runs isolated tasks, MongoDB stores configuration, and XAgentServer (port 5173) runs the web UI. Building and deploying the full system requires at least a few gigabytes of disk space and sufficient CPU and RAM for Docker container isolation.

Install with Docker Compose and Configure OpenAI API Keys

XAgent requires Python 3.10 or later and Docker with Docker Compose. Clone the repository and install dependencies:

bash
pip install -r requirements.txt

The requirements.txt includes the OpenAI client library, FastAPI, Pydantic, Redis, MySQL connectors (PyMySQL), and task scheduling (APScheduler). Build and start ToolServer:

bash
docker compose up

Or, to build from local source:

bash
docker compose build
docker compose up

The docker-compose.yml defines services for ToolServerManager, ToolServerNode, MongoDB for config storage, MySQL for execution records, Redis for caching, and XAgentServer. Run in the background with `docker compose up -d`. Configure XAgent by editing `assets/config.yml` and setting OpenAI API keys using environment variables (DB_USERNAME, DB_PASSWORD). The config must specify at least one key for the primary model (the README recommends `gpt-4-32k`; `gpt-4` works for simpler tasks) and a backup model like `gpt-3.5-turbo-16k`. Do not use `gpt-3.5-turbo` alone; its context length (4K tokens) is too short for XAgent's reasoning. Execute a task:

bash
python run.py --task "put your task here" --config-file "assets/config.yml"

The `--upload-files` argument selects initial files to submit. Outputs are saved to `local_workspace` during execution and copied to `running_records` afterward for review. The setup.py defines XAgent version 1.0.0 as a package, though most users run it directly via run.py rather than installing it as a library.

Execution flow and cost control

Running XAgent is a sequence of LLM calls, each receiving the previous outputs and deciding the next action. The entire exchange (prompts, tool outputs, intermediate reasoning) is logged to `running_records/` with timestamped entries for inspection and debugging. Execution can be expensive: reaching an LLM thousands of times on a complex task (data processing across many files, repeated API calls, code iteration) incurs API costs proportional to token usage. A single 50-step task using gpt-4-32k could cost tens of dollars; simpler tasks using gpt-3.5-turbo-16k cost less but have lower reasoning capability. The config file allows specifying different models for different parts of the task, so you might use a cheaper model for simple subtasks (file reading, shell execution) and reserve gpt-4 for complex reasoning steps. Token consumption is not capped; if the Planner recurses deeply or the Actor explores many wrong paths, the total cost grows linearly. Execution can be reproduced by loading a past run using the `record_dir` setting in the config, which replays the same LLM outputs and tool calls, useful for debugging without re-running API calls. This record-and-replay feature makes it possible to analyze what went wrong without consuming new tokens.

Interaction and control

XAgent can seek human input during execution and respond to guidance. The web interface (default port `5173`) displays the task status, intermediate steps, tool outputs, and LLM reasoning in real time, updated as the agent progresses. This allows pausing execution to provide feedback or correct errors without restarting the entire task. Users can inspect the conversation history, view tool output, and send new instructions inline. The CLI mode (`python run.py`) runs fully autonomously, with all outputs going to stdout and logs for post-execution review. Neither mode requires constant human supervision; the agent decides when it needs help and pauses to ask, typically when it encounters ambiguity or an error it cannot resolve. The dual interface (web UI for interactive work, CLI for batch processing) makes XAgent suitable for both exploratory tasks (research, data exploration) and background automation.

Limitations of the current state

XAgent is experimental and not in active development. The last push was 2026-07-31 (two months before this snapshot) and the most recent release is v1.0.0 from November 2023, indicating the project receives infrequent updates. Error handling is rudimentary: if the LLM makes a mistake (incorrect Python syntax, a wrong API call, misinterpreting a tool output), recovery depends on whether the error message is informative enough for the LLM to learn from. Python execution errors are captured and fed back, but silent failures (e.g., an API call that returns garbage data the LLM misinterprets) are harder to detect. Complex tasks requiring reasoning over hundreds of files may exceed token limits even with gpt-4-32k, forcing manual task subdivision. The Docker container running ToolServer requires significant memory and CPU; a single XAgent instance is not lightweight and concurrent instances quickly exhaust resources. Prompt injection is possible if the task description or file contents contain adversarial text designed to trick the LLM into executing unintended commands. The project does not include defensive measures against such attacks. File permissions within the container are not strictly enforced, so a compromised task could theoretically access all files in the workspace.

Structured Planning Versus AutoGPT's Step-by-Step Execution

AutoGPT and similar projects aim for general task solving but lack the structured planning phase that XAgent's Planner provides. AutoGPT typically uses a single LLM loop for planning and execution, while XAgent separates them, allowing refinement of the plan as the task progresses. Langchain is a library for building agents with predefined workflows using chains and tools; XAgent is a full system with its own UI, agent architecture, and execution runtime. Traditional workflow engines like Airflow or Temporal require explicit DAG definitions written in code; XAgent generates plans dynamically from natural language. The key difference is architectural: Langchain gives you building blocks; XAgent gives you a finished system. None of these systems hide the cost of LLM calls: using XAgent, Langchain, or AutoGPT on a 100-step task calls the LLM approximately 100 times (once per step), which is a fundamental trade-off between autonomy and cost.

Inactive Since July 2026, Apache 2.0 Licensed

The last push was 2026-07-31. The project is Apache-2.0 licensed and welcomes contributions via GitHub. The codebase is Python, with Docker images provided in the dockerfiles/ directory. The website and documentation live in the Markdown_Docs/ directory and at blog.x-agent.net. The repository includes a README in multiple languages (English, Chinese, Japanese), a CONTRIBUTING guide, and a CODE_OF_CONDUCT. The project hierarchy is documented in .project_hierarchy.json. Contact the team at [email protected] for collaboration inquiries.

Editorial conclusion

Evaluate XAgent if you need to automate multi-step technical tasks (data analysis, code generation, research workflows) where human guidance is acceptable and cloud API costs are budgeted. The project is experimental; the last push was 2026-07-31 and releases are infrequent, so check the status before relying on it for production work. Start with simple tasks to verify OpenAI API integration and ToolServer stability in your environment.

Frequently asked questions

What is XAgent used for?

XAgent is an autonomous agent that plans and executes multi-step technical tasks like data analysis, code generation, web research, and file manipulation using LLM decision-making and tool execution.

Can XAgent work offline?

No. XAgent requires OpenAI API access to function. The LLM (gpt-4, gpt-4-32k, or alternatives) must be reachable during execution.

How much does it cost to run XAgent?

Cost depends on the task complexity and LLM model chosen. Each step consumes tokens from your OpenAI quota. A complex 50-step task using gpt-4-32k could cost tens of dollars; simple tasks using gpt-3.5-turbo-16k cost less.

Official sources

  1. License: Apache-2.0
  2. OpenBMB/XAgent on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openbmb-xagent.svg)](https://hysenlabs.com/projects/openbmb-xagent)