YoMo: QUIC-Based Serverless Runtime for LLM Function-Calling Tools
🦖 Serverless AI Agent Framework with Geo-distributed Edge AI Infra.
At a glance
- What is it?
- A Rust binary and CLI that lets teams deploy LLM function-calling tools as serverless processes connected over QUIC, with TLS 1.3 applied to every data packet and a geo-distributed routing model that brings inference closer to users.
- Who is it for?
- YoMo is a good fit for teams that need to deploy LLM function-calling tools as independent, geo-distributed serverless processes and can build their tool implementations in TypeScript (the init template targets TypeScript). Teams that need a Python-first tool ecosystem or that are not prepared to manage a QUIC-based server infrastructure should evaluate LangChain tool use or a simpler REST-based function-calling proxy before committing to YoMo's runtime model.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Teams Build with YoMo
LLM function calling allows a language model to invoke external tools, retrieve data, and take actions. The most common implementation embeds all tools in the same process as the model client, which works for simple use cases but creates scaling problems: a popular tool cannot be deployed independently, all tools share the same failure domain, and latency is determined by the round-trip to the nearest server hosting the full tool stack.
YoMo addresses this by treating each LLM tool as an independent serverless function that connects to a central router over QUIC. The server handles routing between the LLM and the correct tool function. Each tool can be deployed, updated, or scaled independently. Tools run close to users in a geo-distributed setup, reducing the round-trip time for latency-sensitive applications.
The primary target is teams building AI agents that need external capabilities, particularly customer-facing applications where response time directly affects the user experience. The README describes this use case as the core motivation: instant AI inference matters to users, and moving tool execution closer to end users is one way to achieve it.
QUIC Transport and the Serverless Tool Model
QUIC is a UDP-based transport protocol originally developed at Google and standardized as RFC 9000. It provides multiplexed streams, built-in TLS 1.3 encryption, and reduced connection establishment latency compared to TCP with TLS. A QUIC connection requires only one round trip to establish, compared to three for TCP with TLS 1.2. YoMo uses the s2n-quic Rust crate for its QUIC implementation, with the provider-tls-rustls feature enabled.
Every data packet in YoMo is encrypted with TLS 1.3 by design. This is not opt-in; the Cargo.toml specifies s2n-quic with the provider-tls-rustls feature, making TLS mandatory for all tool communications. The certs directory in the repository root holds the certificates used for local development and testing.
The serverless model means tool functions are not HTTP endpoints that accept requests and return responses. Instead, each tool runs as a process that registers itself with the YoMo server and handles function-call arguments as they arrive over the persistent QUIC connection. The server exposes an OpenAI-compatible chat completions endpoint at port 9001 so existing LLM clients can connect without modification.
The Cargo.toml dependency list includes AWS Bedrock runtime and Google OAuth2 dependencies, indicating that cloud provider integrations are built into the server layer rather than requiring external middleware. The serverless directory in the repository contains the serverless execution runtime components.
Installing YoMo and Deploying a First Tool
Install the YoMo CLI:
curl -fsSL https://get.yomo.run | shVerify the installation:
yomo --versionStart the server using Ollama as the LLM provider, pulling the model first:
ollama pull ornith
yomo serveCreate a new tool project:
yomo initRun the tool, naming it get-weather:
yomo run -n get-weather ./appOnce the server and tool are running, send a request using the OpenAI-compatible endpoint:
curl http://127.0.0.1:9001/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "What should I wear for a hike?"}]}'The server also exposes a direct tool endpoint at /tool/{name} that accepts a JSON body with an args field, allowing tools to be tested independently of the LLM routing layer.
Building from Source and the Rust Codebase
YoMo is written in Rust and targets the 2024 edition. The Cargo.toml specifies Tokio as the async runtime, Axum for the HTTP layer, and clap for the CLI. Build from source:
cargo build --release
./target/release/yomo --helpThe tool project template generated by yomo init produces a TypeScript file at ./app/src/app.ts. This means tool implementations are written in TypeScript, compiled and run by the YoMo runtime, while the server itself is Rust. The README shows editing app.ts as the step after yomo init.
Configuration can be passed via a YAML file using the --config flag on yomo serve. The README does not document the full configuration schema, referring instead to the documentation at yomo.run.
The Cargo.toml dependency list includes opentelemetry-otlp with HTTP proto export, tracing-opentelemetry for span collection, and brotli plus flate2 for compression. This indicates the server emits OpenTelemetry traces, which enables integration with observability platforms that support the OTLP protocol.
Where YoMo Is the Wrong Choice
YoMo requires managing a QUIC-based server infrastructure. This is more operational complexity than most teams need for a first LLM function-calling integration. Setting up geo-distributed deployments requires infrastructure work that is only worthwhile when low latency to end users is a hard requirement.
The tool implementation language appears to be TypeScript based on the init template, which generates a project with app/src/app.ts. Teams with existing Python tooling, Python-based data science code, or Python-specific libraries cannot directly reuse that code as YoMo tool functions without a rewrite or a subprocess boundary.
The README does not document a local development mode that runs without a QUIC server, which means every development iteration requires a running server instance. For teams prototyping quickly, this adds friction compared to in-process function calling.
The project has no formal documentation on tool versioning, rollback, or blue-green deployment for live tool updates. The --config flag for yomo serve accepts a custom YAML configuration file, but the full configuration schema is documented at yomo.run rather than in the repository README.
The certs directory contains default certificates for local use, but production deployments require proper certificate management. The README does not document certificate rotation procedures.
How YoMo Differs from LangChain Tool Use
LangChain is a Python framework for building LLM-powered applications, including tool-augmented agents. Tools in LangChain are Python functions or classes decorated with metadata that the framework passes to the LLM as function definitions. Tool execution happens in-process, in the same Python runtime as the agent.
YoMo's architecture is fundamentally different: tools are separate processes that communicate with the server over QUIC. A tool can be deployed on different hardware than the model client, updated without restarting the agent, and co-located with the data it accesses to reduce network latency.
For a team that wants to call external APIs or run Python scripts as LLM tools with minimal infrastructure setup, LangChain is simpler. For a team that needs tool-level scalability, independent deployment of each tool, and low-latency geo-distribution, YoMo provides an architecture that LangChain's in-process model does not support.
License and Release History
The repository declares Apache-2.0 as the license in Cargo.toml. The GitHub license field shows as unknown, but the Cargo.toml entry and the project page on Crates.io are the authoritative source.
Version 2.1.3 was released on 2026-09-23. Two earlier patch releases, v2.1.2 and v2.1.1, were released in August and September 2026. The last push to the main branch was on 2026-09-24. The project is under active development with regular patch releases.
Editorial conclusion
YoMo is a good fit for teams that need to deploy LLM function-calling tools as independent, geo-distributed serverless processes and can build their tool implementations in TypeScript (the init template targets TypeScript). Teams that need a Python-first tool ecosystem or that are not prepared to manage a QUIC-based server infrastructure should evaluate LangChain tool use or a simpler REST-based function-calling proxy before committing to YoMo's runtime model. The license is Apache-2.0. The last push was on 2026-09-24, and v2.1.3 was released on 2026-09-23.
Frequently asked questions
What language are YoMo tool implementations written in?
The yomo init command generates a TypeScript project at ./app/src/app.ts. Tool logic is written in TypeScript, which the YoMo runtime compiles and executes. The YoMo server itself is written in Rust.
Does YoMo require a cloud service or can it run locally?
YoMo can run fully locally. The yomo serve command starts the server on the local machine, and yomo run starts the tool process that connects to it. A cloud deployment is required only when geo-distributed latency reduction is the goal.
Which LLM providers does YoMo support out of the box?
The README shows Ollama as the LLM provider in the getting started example. The Cargo.toml includes AWS Bedrock runtime and Google OAuth2 dependencies, indicating support for AWS Bedrock and Google Cloud providers as well.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yomorun-yomo)