Model or dataset
opensolon/solon-ai avatar
opensolon/solon-ai

Solon AI: a Java LLM, RAG, MCP and Agent framework for JDK 8 through 26

Java AI application development framework (supports LLM-tool,skill; RAG; MCP; Agent-ReAct,Team-Agent). Compatible with java8 ~ java26. It can also be embedded in SpringBoot, jFinal, Vert.x, Quarkus, and other frameworks.

459 stars69 forksJavaApache-2.0

At a glance

What is it?
Solon AI is a Java framework for LLM calls, RAG, MCP and Agent orchestration that runs on JDK 8 through 26 and can be embedded in Spring Boot, Vert.x and Quarkus. This review covers its module layout, its ChatModel and Talent mechanisms, a Maven install path, and where the documentation stops short.
Who is it for?
Solon AI fits Java teams that already run Solon, Spring Boot, Vert.x or Quarkus and want LLM calls, tool invocation, RAG and MCP in one dependency set without leaving the JVM. It is a poor fit if you need a mature Python ecosystem of integrations, or if you want a framework that hides the model dialect behind a single vendor SDK; the README makes dialect selection explicit via provider().
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Solon AI fills for JVM teams

Most LLM orchestration libraries assume Python. A Java team that wants tool calling, retrieval-augmented generation and an agent loop has to either call a Python service over HTTP or assemble the pieces from several unrelated Java libraries. Solon AI is an attempt to put those pieces in one place, and the README states the intent directly: it describes itself as "the same type of development framework as LangChain, LangGraph and LlamaIndex."

The audience is narrower than that sentence suggests. This is for developers who are already writing Java services and want the model layer to live inside the same process, build and dependency graph. The README lists application shapes rather than users: autonomous agents, RAG knowledge bases, multi-agent orchestration, controlled business workflows, document ETL, text-to-SQL dashboards. Two of those, the controlled workflow and the text-to-SQL dashboard, are the ones a JVM shop is most likely to actually ship, because both assume existing business code that the model has to call into.

The compatibility claim is the unusual part. The badges cover JDK 8, 11, 17, 21 and 25, and the repository description extends that to java8 through java26. Supporting JDK 8 in 2026 is a deliberate choice aimed at enterprises that cannot move off older runtimes. It also constrains the implementation: no records, no sealed classes, no virtual threads in the core path.

How ChatModel, Talents and the module tree fit together

The entry point for a model call is ChatModel. The README shows it built from a URL, a provider name, a model name and an optional talent, then used through prompt(). The provider string is the dialect selector, and the README says it is needed "to identify interface style (also called dialect)." That is the central design decision: rather than one client per vendor, there is one interface with per-vendor dialects, so switching from Ollama to DeepSeek to Dashscope means changing the provider value and the endpoint, not the calling code.

Talents are the second mechanism, and they are more interesting than tools. A Talent carries a description, an isSupported predicate and an instruction function. The README example activates an "order_expert" talent only when the user message contains "order", and changes the injected instruction depending on whether prompt metadata marks the user as VIP. That is conditional prompt assembly driven by application state, which is the kind of thing teams usually end up hand-rolling. Whether the predicate and instruction functions are evaluated per turn or cached is not stated in the README.

The repository layout shows how the pieces are separated. There is solon-ai-core for the base abstractions, solon-ai-llm-dialects for vendor adaptation, solon-ai-rag-loaders, solon-ai-rag-repositorys and solon-ai-rag-searchs for the retrieval path, solon-ai-mcp plus mcp-core and mcp-json-jackson2 for the Model Context Protocol, and solon-ai-agent, solon-ai-flow, solon-ai-loop and solon-ai-harness for orchestration. There is also solon-ai-sandbox, solon-ai-router and solon-ai-ui. The README does not describe what the router, sandbox or UI modules do, so their role has to be read from the source.

Installing Solon AI and making a first Ollama call

The README points to Maven Central for artifacts and to the project site for the learning path. The badge group links to a Central search for the org.noear group, which is the namespace the artifacts live under. Add the dependency to your build, then write the call.

The README example targets a local Ollama endpoint on port 11434 and a qwen2.5:1.5b model. The same shape works against any supported provider once the provider value and endpoint change.

java
ChatModel chatModel = ChatModel.of("http://127.0.0.1:11434/api/chat")
        .provider("ollama")
        .model("qwen2.5:1.5b")
        .build();

AssistantMessage result = chatModel.prompt("The weather in Hangzhou today?")
        .options(op -> op.toolAdd(new WeatherTools()))
        .call()
        .getMessage();
System.out.println(result);

What you should see is an AssistantMessage printed to standard output. Because a tool was attached, the model may return a tool call instead of a finished answer, and the framework is expected to run the tool and continue. The README does not spell out that loop, so treat the first run as a way to confirm connectivity and dialect compatibility before wiring in real tools.

For streaming, the README shows a separate call that returns a reactive stream of events:

java
chatModel.prompt("hello").stream();

The README describes the return as Flux<ChatEvent> and calls them "full semantic events", which implies more than raw text chunks. It does not enumerate the event types.

RAG wiring and the Reranking step

The retrieval path in the README is short and concrete: build an EmbeddingModel, optionally build a RerankingModel, create an InMemoryRepository, insert loaded documents, then search. Document loading goes through a PdfLoader in the example, and the repository layout shows a broader solon-ai-rag-loaders module for other formats.

java
EmbeddingModel embeddingModel = EmbeddingModel.of(apiUrl).apiKey(apiKey)
        .provider(provider).model(model).batchSize(10).build();
RerankingModel rerankingModel = RerankingModel.of(apiUrl).apiKey(apiKey)
        .provider(provider).model(model).build();
InMemoryRepository repository = new InMemoryRepository(TestUtils.getEmbeddingModel());
repository.insert(new PdfLoader(pdfUri).load());

List<Document> docs = repository.search(query);
docs = rerankingModel.rerank(query, docs);

The batchSize(10) on the embedding model is the only tuning knob the README shows, and it matters: embedding a large PDF in batches of ten against a hosted API will produce a lot of sequential round trips. There is no mention of concurrency, retry or rate limiting here, which is where most production RAG pipelines actually break. The InMemoryRepository is also the only repository the README demonstrates. The repository layout lists solon-ai-rag-repositorys as a module, so persistent implementations presumably exist, but the README does not name or compare them.

Separating rerank from search is the right call, and it is the part many lightweight Java RAG stacks skip. Reranking costs an extra model call per query, so it is a latency and money trade-off rather than a free improvement.

MCP support and what the README leaves out

Solon AI implements the Model Context Protocol on both sides. The topics list includes mcp-client and mcp-server, the repository has a solon-ai-mcp module alongside mcp-core and mcp-json-jackson2, and the ChatModel example attaches an McpGatewayTalent as a default talent. That last detail is the most informative: MCP tools are surfaced to the model through the same Talent mechanism as local tools, so a remote MCP server and a local Java tool are addressed through one interface.

What the README does not provide is a complete MCP example. There is no server registration snippet, no transport configuration, and no statement about which MCP protocol revision is implemented. For a protocol whose specification has moved through several revisions, that is the first thing to check in the source before committing. The mcp-json-jackson2 module name suggests Jackson 2 is the JSON binding, which matters if your application already pins a different Jackson major version.

The README also links to a separate repository of embedding examples for third-party frameworks, hosted on Gitee, GitCode and GitHub. If you are embedding Solon AI inside Spring Boot, Vert.x or Quarkus, that repository is where the working integration code lives, not the main README.

Where Solon AI is the wrong choice

The honest limitation is documentation depth relative to surface area. The module tree contains roughly twenty top-level modules, and the README gives working code for three of them: ChatModel, Talents and the RAG path. Agent, flow, loop, harness, router, sandbox and UI are named in the layout and absent from the prose. If your project depends on Team-Agent orchestration, you are reading source code and tests, not documentation.

The second limitation is ecosystem gravity. LangChain4j is the comparison Java teams will actually make, and it has a longer list of model and vector store integrations with a larger body of public examples. Solon AI's advantage is that it does not force you onto a particular application framework: the README states it can be embedded in Spring Boot, jFinal, Vert.x and Quarkus, whereas LangChain4j is typically paired with Spring Boot through its own starter. If you are already on Solon, that advantage compounds. If you are not, you are adopting a second framework's conventions for the AI layer alone.

A third case: if your workload is a single chat completion against one vendor, the dialect abstraction and the module set are overhead. A plain HTTP client and a JSON mapper would be fewer moving parts.

Maintenance, releases and licence

The repository is not archived, and the last push was on 2026-09-07. The most recent release is v4.1.0, dated the same day, following v4.0.6 on 2026-08-17 and v4.0.5 on 2026-08-12. That is a release cadence measured in weeks, and the version numbering suggests the project is willing to ship minor bumps rather than holding changes for a major.

The practical consequence is upgrade cost. On a cadence like this, pinning an exact version and reading UPDATE_LOG.md before each bump is cheaper than tracking the main branch. The README does not document a deprecation policy or a compatibility guarantee between minor versions, so the release notes are the only signal available.

The licence is Apache-2.0, per the LICENSE file and the badge. That is a permissive licence with an explicit patent grant, which is generally the least friction option for commercial use. It is not legal advice; if you redistribute the artifacts or modify them, read the NOTICE.template file in the repository root, since Apache-2.0 obligations around attribution are usually handled there.

Editorial conclusion

Solon AI fits Java teams that already run Solon, Spring Boot, Vert.x or Quarkus and want LLM calls, tool invocation, RAG and MCP in one dependency set without leaving the JVM. It is a poor fit if you need a mature Python ecosystem of integrations, or if you want a framework that hides the model dialect behind a single vendor SDK; the README makes dialect selection explicit via provider(). Before adopting, verify two things against the repository rather than the README: which artifact under solon-ai-core or solon-ai-llm-dialects corresponds to your target model, and whether the release notes for v4.1.0 cover any breaking changes to the Talent or MCP APIs you plan to use.

Frequently asked questions

What Java versions does Solon AI support?

The repository description states compatibility with java8 through java26, and the README badges cover JDK 8, 11, 17, 21 and 25. Supporting JDK 8 means the core cannot use records, sealed classes or virtual threads.

Which model providers does Solon AI support?

The README lists OpenAI, Gemini, Claude, Ollama, DeepSeek and Dashscope as dialects selected through the provider value on ChatModel. The repository has a solon-ai-llm-dialects module where those adaptations live.

Can Solon AI be used inside Spring Boot or Quarkus?

Yes. The README states it can be embedded in SpringBoot, jFinal, Vert.x and Quarkus, and links to a separate repository of embedding examples for third-party frameworks on Gitee, GitCode and GitHub.

Does Solon AI support MCP servers and clients?

The topics list includes mcp-client and mcp-server, and the repository contains solon-ai-mcp, mcp-core and mcp-json-jackson2 modules. The README shows an McpGatewayTalent attached to a ChatModel but does not give a full server or client configuration example.

Official sources

  1. License: Apache-2.0
  2. opensolon/solon-ai on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/opensolon-solon-ai.svg)](https://hysenlabs.com/projects/opensolon-solon-ai)