InternLM/lagent: a PyTorch-style framework for wiring LLM agents
A lightweight framework for building LLM-based agents
At a glance
- What is it?
- Lagent treats an agent as a stack of layers with message passing between them, and keeps conversation state in an explicit memory object. This review covers the install path, the aggregator and parser hooks, and where the design runs out of road.
- Who is it for?
- Adopt lagent if you want agent control flow expressed as Python objects you can subclass, and if you are already running vLLM or LMDeploy locally, since the README's first example uses VllmModel with a local path. Do not adopt it if you need a maintained release cadence or a documented upgrade path: the newest tag is agentrl_rc0 from 2026-05-19, and the README does not document rollback or migration between versions.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 17 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem lagent solves, and the shape of the answer
Most agent codebases start as a script: a while loop, a prompt string, a call to a chat completion endpoint, and a growing list of dictionaries. That works until you want a second agent, a different memory policy, or a parser that turns model output into a tool call. At that point the loop becomes the architecture, and changing it means rewriting it.
Lagent's answer is to borrow PyTorch's vocabulary. The README states the project is "inspired by the design philosophy of PyTorch", and the analogy is explicit: you create layers and define message passing between them. An agent is an object you call, not a loop you write. The audience is Python developers who already run an inference server and want to compose agents without adopting a heavier orchestration product.
That framing has a cost worth naming up front. The README's own example is a single-turn exchange with a system prompt that restricts the answer to one of three Chinese characters. It is a demonstration of plumbing, not of agent behaviour. Anyone expecting a tutorial on tool use, planning or multi-step reasoning will find the README stops well short of that.
AgentMessage as the unit of everything
Every input and output in lagent is an AgentMessage. The dataclass carries sender, content, formatted, extra_info, type, receiver and stream_state. The README shows the printed result of a call: content is the text, sender is 'Agent', and stream_state is AgentStatusCode.END. Because the same type flows in both directions, an agent can be handed the output of another agent without a translation step, which is the mechanism behind the multi-agent claim.
The second field that matters is formatted. It is reserved for whatever an output_format parser extracts from the raw model text. In the forward method shown in the README, the LLM response is passed to self.output_format.parse_response, and the result is attached to the returned AgentMessage rather than replacing content. So the raw string and the parsed structure travel together. If you have ever lost the original completion because your parser consumed it, this is a deliberate fix for that.
The state model is the part I would flag as the real design decision. Memory is mutated in __call__, not in forward. The README gives the pseudo code: pre_hooks, add_memory of the input, forward, add_memory of the output, post_hooks. The stated reason is that both directions get recorded in each pass. The practical consequence is that forward is a pure-ish function you can override while the surrounding bookkeeping stays consistent, and that anything you do inside forward will not be persisted unless it returns an AgentMessage.
Installing lagent and running the first agent
The README gives one installation path: from source. There is no documented pip install of a released wheel in the README body, even though a PyPI badge is present at the top. Clone the repository, then install in editable mode.
git clone https://github.com/InternLM/lagent.git
cd lagent
pip install -e .The editable install is what the README shows, and it is the right choice if you intend to subclass agents or aggregators, because your local edits take effect without reinstalling. Dependencies are split: requirements.txt pulls in requirements/optional.txt and requirements/runtime.txt, so the optional layer is installed by default when you go through setup.py.
The first real use is a single agent over a local vLLM model. The README uses a Qwen2-7B-Instruct path, an INTERNLM2_META template, tensor parallelism of 1, and a stop word of <|im_end|>. Note that the meta template and the model do not match by name; you supply the template that fits the chat format you are serving.
from lagent.agents import Agent
from lagent.schema import AgentMessage
from lagent.llms import VllmModel, INTERNLM2_META
llm = VllmModel(
path='Qwen/Qwen2-7B-Instruct',
meta_template=INTERNLM2_META,
tp=1,
top_k=1,
temperature=1.0,
stop_words=['<|im_end|>'],
max_new_tokens=1024,
)
agent = Agent(llm, system_prompt='你的回答只能从“典”、“孝”、“急”三个字中选一个。')
bot_msg = agent(AgentMessage(sender='user', content='今天天气情况'))
print(bot_msg)What you should see is an AgentMessage whose content is a single character and whose sender is 'Agent'. If you get an empty string or a template error instead, the mismatch is almost always between the meta_template you passed and the chat format the server is actually using.
Inspecting memory takes two forms, and the difference matters. agent.memory.get_memory() returns live AgentMessage objects. agent.state_dict() returns a plain dictionary under the key 'memory', with each message flattened to primitives. The second is what you would serialize; the first is what you would assert against in a test.
memory = agent.memory.get_memory()
dumped = agent.state_dict()
print(dumped['memory'])
agent.reset()agent.reset() clears the session, and the README notes that session_id defaults to 0. That default is the thing to think about before you put this behind a web server.
Aggregators: where prompt assembly actually happens
The step between memory and the model is the aggregator. DefaultAggregator converts the stored AgentMessage list into OpenAI-style role and content dictionaries, and it is invoked inside forward with the session's memory, the agent name, the output format and the template.
The README's custom example, FewshotAggregator, is worth reading closely because it encodes two decisions. First, it prepends a fixed list of few-shot turns before the real conversation. Second, it merges consecutive user messages into a single user entry rather than emitting two in a row, checking whether the last appended message already has role 'user' and appending to its content if so. That merge is a real constraint of the OpenAI message format, and it is the kind of detail that bites people who build prompt strings by hand.
The trade-off is that aggregation is a full pass over memory on every forward call. Nothing in the README suggests caching or incremental assembly. For short conversations this is irrelevant. For a long-running session with tool output appended repeatedly, you are rebuilding the whole message list each turn, and the cost scales with history length rather than with the new input.
Output parsers and the ToolParser path
Structured output is handled by an output_format object with a parse_response method. The README shows the branch in forward: if output_format is set, the raw LLM response is parsed and the parsed value goes into the formatted field of the returned AgentMessage.
The README begins an example using ToolParser from lagent.prompts.parsers with a system prompt instructing the model to analyse step by step and write Python code. The excerpt is truncated at the point where the parser is constructed, so the full tool-calling flow is not visible in the README itself. That is a documentation gap, not a missing feature: the examples directory contains run_agent_lmdeploy.py, run_agent_services.py and several async variants, which is where the working tool-use code lives.
My read is that this split is intentional. The README teaches the object model, and examples/ teaches the integrations. It is a reasonable division for a library, and a frustrating one for someone evaluating whether tool calling works before they clone anything.
Where lagent is the wrong tool
The default session_id of 0 is the clearest limitation. Memory is stored per session, but if you never pass a session_id, every caller shares session 0. The README documents agent.reset() as clearing "this session" and notes the default, without describing how sessions are isolated in a served deployment. If you are building a multi-user service, you need to confirm from the source how session identifiers are threaded through, because the README does not tell you.
Release cadence is the second issue. The tags listed are v0.5.0rc2 from 2024-11-29, v0.5.0rc3 from 2025-03-04, and agentrl_rc0 from 2026-05-19. Two of the three are release candidates, and the naming changes between the 0.5 line and agentrl_rc0, which suggests a different track rather than a straight continuation. The last push to the default branch was on 2026-08-03. The README does not document rollback, deprecation policy or migration between these versions, so pinning to a tag and reading the diff is the only reliable upgrade procedure available from the published documentation.
Third, lagent assumes you have a model server. The README's example constructs a VllmModel with a local path, and the examples include LMDeploy and OpenAI variants. There is no hosted endpoint, no managed control plane, and no scheduler beyond what you write. If you want an agent runtime that runs somewhere for you, this is the wrong layer.
How lagent differs from LangChain and LlamaIndex
LangChain, the most common comparison point, is a broad integration library: it ships connectors for many model providers, vector stores, retrievers and tool wrappers, and its abstractions are chain- and runnable-oriented. The difference in approach is where the extensibility sits. In LangChain you typically compose existing components and reach for a custom class when nothing fits. In lagent you subclass Agent, Memory or DefaultAggregator, and the README walks through exactly that: FewshotAggregator overrides aggregate, and the agent is constructed with aggregator=FewshotAggregator([...]).
LlamaIndex takes a third position, centred on indexing and retrieval over your documents, with agents layered on top. Lagent has no retrieval layer documented at all. There is no index, no document store, no retriever abstraction in the README. If your problem is "answer questions over a corpus", lagent gives you the agent shell and leaves the retrieval to you.
So the honest comparison is scope. Lagent is smaller and expects you to write more, in exchange for an object model you can read in one sitting. That is a real benefit for teams who have been burned by framework churn, and a real cost for teams who want retrieval and provider integrations on day one.
Licence and the maintenance arithmetic
Lagent is Apache-2.0. The repository carries a LICENSE file at the top level, and the badge in the README points to it. Apache-2.0 includes an explicit patent grant and requires that you preserve notices and state changes you make. If you fork the aggregator or agent classes and redistribute, keep the licence and attribution intact. This is a description of the licence text, not legal advice; check the terms against your own distribution model.
The maintenance arithmetic is straightforward and unflattering. The last push to the default branch was on 2026-08-03, roughly six weeks before this writing, so the repository is not dormant. But the release history is thin and candidate-heavy, and the README does not describe a versioning policy. The upgrade cost you should budget for is not the pip command; it is the time to read the diff between agentrl_rc0 and whatever you pinned, because nothing in the published documentation tells you what changed or what breaks.
Editorial conclusion
Adopt lagent if you want agent control flow expressed as Python objects you can subclass, and if you are already running vLLM or LMDeploy locally, since the README's first example uses VllmModel with a local path. Do not adopt it if you need a maintained release cadence or a documented upgrade path: the newest tag is agentrl_rc0 from 2026-05-19, and the README does not document rollback or migration between versions. Before committing, verify three things in your own checkout: that the optional requirements in requirements/optional.txt install cleanly on your platform, that your chosen model's chat template matches the meta_template you pass, and that agent.reset() gives you the session isolation you expect when you serve more than one user.
Frequently asked questions
How do I install InternLM/lagent?
The README gives one path: clone the repository and run pip install -e . from the project directory. That installs the runtime and optional requirement files, since requirements.txt includes both.
How does lagent keep conversation state?
Each Agent has a memory object, and both the input and the output messages are appended during __call__ rather than inside forward. You can read it live with agent.memory.get_memory() or as plain dictionaries with agent.state_dict()['memory'].
How do I clear an agent's memory in lagent?
Call agent.reset(). The README notes that this clears the session with session_id=0, which is the default.
What is the latest release of InternLM/lagent?
The most recent tag listed is agentrl_rc0 from 2026-05-19. Before that are v0.5.0rc3 from 2025-03-04 and v0.5.0rc2 from 2024-11-29, so two of the three recent tags are release candidates.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/internlm-lagent)