baidubce/app-builder: the Qianfan AppBuilder Python SDK for RAG and agent pipelines
appbuilder-sdk, 千帆AppBuilder-SDK帮助开发者灵活、快速的搭建AI原生应用
At a glance
- What is it?
- Baidu's Apache-2.0 Python client for Qianfan AppBuilder wraps model calls, 40+ capability components, knowledge base management and workflow orchestration behind one package. It is a cloud client, not a self-hosted framework, and that shapes who should use it.
- Who is it for?
- Adopt it if your application already lives on Baidu Smart Cloud and you want the Qianfan AppBuilder console, knowledge base and component catalogue reachable from Python without hand-rolling HTTP calls. Do not adopt it if you need a self-hosted LLM stack, a non-Baidu cloud, or a provider-neutral abstraction layer, because every component routes through Baidu endpoints and a token.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 158 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What baidubce/app-builder actually solves
The repository is the client SDK for Baidu Smart Cloud Qianfan AppBuilder, a hosted platform at appbuilder.cloud.baidu.com. The problem it addresses is not "how do I call an LLM" but "how do I reach a specific vendor's model catalogue, component catalogue and published applications from Python without writing that integration myself". The README describes three jobs: calling models and the 40+ components drawn from Baidu's ecosystem, orchestrating knowledge bases and workflows, and deploying the result. The audience is therefore narrower than the topic list suggests. If you are building on Qianfan and want the console-side assets (knowledge bases, published apps, component quotas) available in code, this is the intended path. If you are evaluating LLM frameworks in general, the design decisions here are vendor-shaped and will read as constraints rather than features.
How the SDK is layered: Playground, components, AgentRuntime
The README exposes three abstraction levels. At the bottom, Playground takes a prompt template and a model name and returns a Message; the example passes model="DeepSeek-V3.1" and shows that the response object carries content plus a token_usage block with prompt_tokens, completion_tokens and total_tokens. Above that sit the capability components, each a class with a run method, such as RagWithBaiduSearchPro, which returns both an answer and a search_baidu array containing title, url, site_name and ref_id for each retrieved page. That ref_id is what makes citation rendering possible without a second lookup. At the top, AgentRuntime is the orchestration layer, described as offering Message, Component and AgentRuntime abstractions and as interoperating with LangChain and OpenAI ecosystem pieces. The data flow is consistently cloud-bound: your process holds a token, the SDK serialises a request, Baidu's service does the work, and the SDK returns a typed Message. Nothing in the README suggests local inference. Tracing and DebugLog are also server-assisted monitoring tools rather than a local profiler.
Installing appbuilder-sdk and making a first Playground call
The README states Python 3.9 or later is required and gives a single pip command. It also points to install.md for Java, Go and Docker usage, so pip is the Python-only path.
python3 -m pip install --upgrade appbuilder-sdkBefore any call you need a personal token, which the README says to replace in the example. It is read from the APPBUILDER_TOKEN environment variable. The README's own snippet sets it inline, which is fine for a scratch script but leaves the key in your source file.
import os
import appbuilder
os.environ["APPBUILDER_TOKEN"] = "your api key"The first real call is Playground. The README defines a prompt template with {role} and {question} placeholders, wraps the inputs in an appbuilder.Message, and passes stream=True. The output is iterable when streaming, and printing output.model_dump_json(indent=4) afterwards gives the full content plus token usage. Note the temperature value in the README example, 1e-10, which is effectively deterministic sampling rather than a typical default.
template_str = "你扮演{role}, 请回答我的问题。\n\n问题:{question}。\n\n回答:"
input = appbuilder.Message({"role": "java工程师", "question": "请简要回答java语言的内存回收机制是什么,要求100字以内"})
playground = appbuilder.Playground(prompt_template=template_str, model="DeepSeek-V3.1")
output = playground(input, stream=True, temperature=1e-10)
for stream_message in output.content:
print(stream_message)What you should see is a streamed answer followed by the JSON dump. If the token is wrong or missing, the failure surfaces at call time, not at import time, so a quick Playground call is the cheapest way to validate credentials.
Component calls need quota, not just a token
The second README example uses RagWithBaiduSearchPro, which combines Baidu search with ERNIE's semantic understanding. The call signature differs from Playground: you construct the component with a model, wrap the query in a Message, and pass an instruction as a second Message to run. The README's own example asks whether 9.11 or 9.8 is larger, and the returned extra.search_baidu array shows the pages the answer drew on.
rag = appbuilder.RagWithBaiduSearchPro(model="DeepSeek-V3.1")
input = appbuilder.Message("9.11和9.8哪个大")
result = rag.run(message=input, instruction=appbuilder.Message("你是专业知识助手"))
print(result.model_dump_json(indent=4))The constraint the README states plainly is that components require a claimed free trial quota before use, and it links to a Baidu console resource page for that. This is the first place a new user gets stuck: the SDK installs cleanly, a Playground call may succeed, and then a component call fails because the account has no entitlement for that component. Treat quota provisioning as a prerequisite step, not an error to debug later.
The RAG pipeline is assembled from named components, not hidden
The README lays out an industrial RAG flow as six stages and maps each to a component class: document parsing (DocParser, DocFormatConverter, DocCropEnhance, ExtractTableFromDoc, GeneralOCR), chunking (DocSplitter), embedding (Embedding), index construction and retrieval (BaiduVectorDBRetriever, BaiduElasticSearchRetriever), and answer generation. The answer-generation stage is where the design gets interesting, because it includes components that are not retrieval at all: QAPairMining, SimilarQuestion, TagExtraction, IsComplexQuery, QueryDecomposition, QueryRewrite, MRC and HallucinationDetection. Listing a hallucination detector as a pipeline stage is an admission that generated answers need a check, and it is more honest than frameworks that present retrieval as sufficient. The trade-off is that this is a component catalogue, not a pipeline. You choose the ordering, handle the intermediate representations, and pay for each stage separately. The cookbook at cookbooks/end2end_application/rag/rag.ipynb is the README's pointer for seeing the stages wired together.
Deployment options and where they stop
AgentRuntime can be served two ways according to the README: as an API service on Flask with gunicorn, or as a conversational front end on Chainlit. There is also an appbuilder_bce_deploy tool that pushes a program to Baidu Cloud to expose a public API and connect to AppBuilder workflows. The setup.py extras make the cost visible: the serve extra pins chainlit~=1.0.200, flask~=2.3.2, flask-restful==0.3.9 and arize-phoenix==4.5.0, while tracing pulls SQLAlchemy==2.0.31 and the LangChain extra pins langchain==0.3.0 and datamodel-code-generator==0.25.8. Installing the full set is a heavier dependency tree than the base package, and the exact pins mean you inherit those versions. The README does not document rollback for the BCE deploy path, nor does it describe what happens to a deployed service when the underlying workflow changes. If you need a documented rollback story, that gap is real and you should not assume one exists.
When this is the wrong tool, and what to use instead
The clearest mismatch is self-hosting. Every mechanism described here assumes Baidu Smart Cloud endpoints and a personal token; the README never describes running models or vector stores locally, and pymochow in requirements.txt points at a Baidu-managed vector service rather than an embedded store. If your constraint is data residency on your own hardware, this SDK cannot satisfy it. The natural alternative in that case is LangChain with a local model server and a local vector database, because the orchestration primitives (chains, retrievers, tools) run entirely in your process. The difference in approach is concrete: LangChain asks you to supply the model and the store, while appbuilder-sdk asks you to supply a token and then selects from Baidu's catalogue. The README does state that the SDK interoperates with LangChain and OpenAI ecosystem pieces, so the two are not mutually exclusive, but the interoperability is an adapter around a cloud call, not a replacement for one. A second mismatch is provider neutrality: if you want to swap between vendors without changing pipeline code, a cloud-specific client is the wrong layer to build that abstraction on.
Maintenance, licensing and upgrade cost
The repository is not archived and the last push was on 2026-04-26. Releases are infrequent: 1.1.1 on 2025-09-21, 1.1.0 on 2025-06-20 and 1.0.6 on 2025-05-20. Note the README still advertises 1.1.0 as the latest version while the release list shows 1.1.1, so the documentation lags the package. Pin your version rather than tracking latest. The licence is Apache-2.0, declared in the repository LICENSE file and in the badge at the top of the README, with the setup.py header carrying the standard Apache boilerplate. Apache-2.0 is permissive and includes a patent grant, but it covers the SDK code only; the Qianfan services the SDK calls are governed by Baidu's cloud terms and by whatever quota you have claimed, which is a separate commercial relationship. Upgrading carries two risks visible in the repository: the pinned extras in setup.py mean a version bump can move Chainlit, Flask or LangChain under you, and the version string in setup.py must be changed in step with the release, which the truncated comment in that file explicitly warns about.
Editorial conclusion
Adopt it if your application already lives on Baidu Smart Cloud and you want the Qianfan AppBuilder console, knowledge base and component catalogue reachable from Python without hand-rolling HTTP calls. Do not adopt it if you need a self-hosted LLM stack, a non-Baidu cloud, or a provider-neutral abstraction layer, because every component routes through Baidu endpoints and a token. Before writing application code, verify your APPBUILDER_TOKEN works against a single Playground call, confirm which components your account has been granted quota for, and check the install.md page for whether you need the Java, Go or Docker path instead of pip.
Frequently asked questions
How do I install appbuilder-sdk?
The README gives one command for Python: python3 -m pip install --upgrade appbuilder-sdk, with Python 3.9 or later required. Java, Go and Docker installation paths are documented separately in install.md.
What is appbuilder-sdk used for?
It is the client SDK for Baidu Smart Cloud Qianfan AppBuilder. The README describes three uses: calling models and 40+ capability components, orchestrating knowledge bases and workflows, and deploying the result as an API service or Chainlit front end.
Does appbuilder-sdk work without a Baidu Cloud account?
No. Calls read a personal token from the APPBUILDER_TOKEN environment variable, and the README states that capability components require claiming a free trial quota before use. Both are tied to a Baidu Smart Cloud account.
How do I run a RAG pipeline with appbuilder-sdk?
The README lists the stages as parsing, chunking, embedding, index construction and retrieval, and answer generation, each mapped to a component class such as DocParser, DocSplitter, Embedding and BaiduVectorDBRetriever. A cookbook notebook at cookbooks/end2end_application/rag/rag.ipynb shows the stages combined.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/baidubce-app-builder)