Underthesea: a Vietnamese NLP toolkit that now ships an agent runtime
Underthesea - AI Assistant
At a glance
- What is it?
- Underthesea was a Vietnamese word segmenter and NLP library. Since v9.3.0 it also bundles a multi-provider AI agent with tracing and an A2A server, written against the Python standard library. Here is what that combination does well and where it stops.
- Who is it for?
- Adopt Underthesea if you already handle Vietnamese text in Python and want word segmentation, POS tagging and an agent loop in one dependency, or if you want an agent runtime whose base install pulls in no vendor SDK. Skip it if your stack is JavaScript, if you need a hosted service with an SLA, or if you only need English-language agents, where the Vietnamese modules add weight you will not use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Underthesea solves, and for whom
Vietnamese is a language where whitespace does not mark word boundaries. A string like "Hà Nội hôm nay đẹp trời" arrives as space-separated syllables, and downstream tasks such as classification, search or entity extraction need those syllables grouped into words first. Underthesea's core is that grouping step, packaged with the usual NLP companions: part-of-speech tagging, named entity recognition, sentiment, text-to-speech and a CLI. The project describes itself in pyproject.toml as a "Vietnamese NLP Toolkit", and that is the part with the longest history.
Since v9.3.0 the package also carries an agent runtime. The README says Underthesea is "an open-source Agentic AI Toolkit with built-in Vietnamese NLP capabilities", offering multi-provider agent support alongside the NLP modules. The audience for that half is narrower than it first looks: Python developers building LLM-backed assistants who want the agent loop, tool calling and tracing to come from the same pip install as their Vietnamese text processing. If you are building an English-only agent, the NLP half is dead weight, and the package's real value proposition does not apply to you.
How the agent talks to four LLM providers without vendor SDKs
The design choice that shapes everything else is stated plainly: the agent communicates with LLM APIs using only Python standard library, urllib plus json, with no openai, anthropic or google-genai packages required. The pyproject.toml optional-dependencies table backs this up. There is an agent extra, and it is empty.
Each provider is a class following the Anthropic SDK pattern: OpenAI, AzureOpenAI, Anthropic and Gemini, each constructed with an api_key, and AzureOpenAI also taking endpoint and deployment. The LLM class is the auto-detect path, reading whichever provider environment variable is set. An Agent wraps a provider, a name, an optional instruction string and a list of Tool objects built from plain Python functions whose docstrings become the tool description.
The trade-off is explicit. You get a small dependency surface and no version churn from vendor SDKs. You give up whatever those SDKs add beyond the wire format: retry policies, typed response objects, provider-specific helpers. The README does not document retry behaviour, timeouts or rate-limit handling, so treat those as things you would have to add yourself.
Install and first agent call
Installation is a single pip command, and the README shows no virtualenv step, so create one if you want isolation. The package requires Python 3.10 or newer according to pyproject.toml, and the badge lists support through 3.14.
$ pip install undertheseaBefore the agent can reach a model, export the key for whichever provider you intend to use. The README gives one line per provider; the OpenAI form is:
export OPENAI_API_KEY=sk-...With that set, LLM() auto-detects the provider from the environment. The smallest working example is three lines, and the agent is called like a function:
from underthesea.agent import Agent, LLM
agent = Agent(name="assistant", provider=LLM())
agent("Hello!")The return value is the model's reply. If you would rather see tokens as they arrive, the README shows agent.stream() returning chunks you print in a loop. To skip the environment entirely, pass a provider class directly, for example OpenAI(api_key="sk-..."), or AzureOpenAI with api_key, endpoint and deployment.
Tools, default tools and what they expose
Tool calling turns a Python function into something the model can invoke. You wrap the function in Tool and pass a list to the Agent; the docstring is the description the model sees. The README's weather example returns a dict and the agent answers with the values from it. There is nothing to register globally and no schema file to maintain.
The package also ships a set of ready-made tools. The README counts 12 built-in tools and names calculator, datetime, web search, wikipedia, file I/O, shell and python exec, exposed as default_tools. That list deserves a second read before you copy it into a production agent. Shell and python exec are the two that should give any reviewer pause: a model that can call them is a model with a path to arbitrary execution on the host. The README presents default_tools as a convenience and does not describe a sandbox, an allowlist or a confirmation step. If you use them, the isolation is your responsibility, not the library's.
Tracing, sessions and the A2A server
Every agent call is traced automatically to ~/.underthesea/traces/, and the README shows the console output: a trace id, each generation with model name, latency and token counts, then each tool call with its duration. Setting UNDERTHESEA_TRACE_DISABLED=1 turns it off. The default writes files to the local filesystem, which is convenient during development and a data-handling question in production, since prompts and completions land on disk unencrypted unless your filesystem says otherwise. A LangfuseTracer is available if you install langfuse, and a @trace decorator lets nested functions become child spans.
For long runs, Session wraps an agent with a progress_file and a task list, and run_until_complete(max_sessions=5) drives it across context resets. The README attributes the pattern to Anthropic's writing on harnesses for long-running agents. The server side is where the dependency discipline pays off: serve(agent, port=8000, path="/a2a/math", ui=True) exposes the agent over the A2A protocol as a raw ASGI app, with a bundled chat UI and a discoverable agent card. The base install carries no web framework; uvicorn, starlette and httpx come from the agent-server extra, and make_app() returns a callable you can mount in any ASGI server.
Where Underthesea is the wrong choice
The agent runtime is young. It arrived in v9.3.0 in April 2026, and the releases that followed added tracing and the A2A server. The README documents the happy path and little else: no retry semantics, no timeout configuration, no discussion of what happens when a provider returns a malformed response. An agent library whose error handling is undocumented is one you will be debugging from source.
The other boundary is language. If your text is not Vietnamese, the NLP modules are inert and you are choosing Underthesea purely as an agent runtime, where it competes against libraries with far more surface area. And if you need a managed service with an uptime commitment, this is a library you host yourself. The README points to a live demo and a docs site but describes no hosted API tier.
The alternative: vncorenlp and PyVi
For Vietnamese word segmentation specifically, the comparison people search for is Underthesea against vncorenlp, with PyVi as a third name in the same space. The approaches differ. Underthesea is a pip-installable Python package that pulls its core from underthesea_core and its models through huggingface-hub, so the model download happens at runtime rather than at install. vncorenlp is built around a Java implementation, which means a JVM in your deployment and typically a separate service or a bundled jar. The practical consequence is operational: Underthesea fits a pure-Python container, while a JVM-based segmenter adds a runtime your build has to carry.
That does not make Underthesea the better segmenter on accuracy, and the README makes no accuracy claim either way. It does mean the integration cost is lower if your stack is already Python. PyVi sits closer to Underthesea in packaging but the README does not compare against it, so any quality judgement would have to come from your own evaluation on your own text.
Licence and the cost of keeping up
Underthesea is Apache-2.0, declared both in pyproject.toml and in the repository. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on your own code. The bundled models are downloaded separately through huggingface-hub, and their licences are not stated in the README, so check each model card before commercial use. Nothing in this is legal advice.
Upgrade cost is mostly about the split identity. The NLP modules have been stable for years, but the agent half is moving: the last push to the repository was on 2026-09-09, and three releases landed between April and May 2026. If you depend on the agent API, pin the version and read the release notes before bumping, because the surface is still being extended rather than frozen.
Editorial conclusion
Adopt Underthesea if you already handle Vietnamese text in Python and want word segmentation, POS tagging and an agent loop in one dependency, or if you want an agent runtime whose base install pulls in no vendor SDK. Skip it if your stack is JavaScript, if you need a hosted service with an SLA, or if you only need English-language agents, where the Vietnamese modules add weight you will not use. Before committing, check the docs site for the current CLI subcommands, confirm which of the 12 built-in tools you are willing to expose, and read the trace files that land in ~/.underthesea/traces/ to see whether the default tracing is acceptable for your data.
Frequently asked questions
How do I install Underthesea?
Run pip install underthesea. The package requires Python 3.10 or newer according to pyproject.toml, and the README shows no additional setup step beyond that command.
What is Underthesea?
It is a Python package that the README describes as an open-source Agentic AI Toolkit with built-in Vietnamese NLP capabilities. Since v9.3.0 it has combined Vietnamese text processing modules with a multi-provider AI agent runtime.
What is a good alternative to Underthesea?
vncorenlp is the alternative most often compared with it for Vietnamese word segmentation. The difference is packaging: Underthesea is a pip-installable Python package, while vncorenlp is built around a Java implementation and needs a JVM in your deployment.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/undertheseanlp-underthesea)