Xberg: A Rust Document Intelligence Library with Fifteen Language Bindings
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
At a glance
- What is it?
- Xberg is a document and code intelligence engine with a Rust core that extracts text, tables, metadata, and structured data from 107 formats across 141 file extensions, with support for OCR, audio transcription, and code analysis across 371 programming languages. It is the successor to the Kreuzberg project.
- Who is it for?
- Xberg fits engineering teams that need to extract clean, structured content from a wide variety of document formats and cannot afford to assemble a pipeline from multiple per-format libraries. The 15 language bindings mean it integrates into most existing codebases without a language switch.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Problem Xberg Solves and Who It Targets
Extracting text from a PDF, parsing an Excel spreadsheet, transcribing an audio recording, and analyzing source code for symbols are four separate problems that typically require four separate libraries. Xberg addresses this by providing a single engine that handles all of them through a common API.
The README describes Xberg as a tool you point at anything: a PDF, a scanned image, a spreadsheet, an audio file, a URL, an archive, or a source tree, and it returns clean text, tables, metadata, and structured data. Format detection, reading, OCR, and extraction all happen inside one library call.
Xberg targets teams building document processing pipelines, RAG (retrieval-augmented generation) systems, code analysis tools, and data extraction workflows. The 15 language bindings (Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, and C FFI) make it accessible from most existing technology stacks without requiring a language rewrite or a separate service boundary.
The README notes that Xberg is the next iteration of Kreuzberg, a predecessor project at kreuzberg-dev/kreuzberg-v4-lts. The same document intelligence engine is present, rebuilt and rebranded under a fresh v1 version line.
The Rust Core and What It Provides
Xberg's core is written in Rust. The Cargo.toml workspace lists the main crate as xberg, with separate crates for specific OCR backends (xberg-candle-ocr, xberg-paddle-ocr, xberg-tesseract), language bindings (xberg-py, xberg-node, xberg-jni, xberg-php, xberg-ffi, xberg-wasm), and tooling.
The Rust version floor is 1.92. Capabilities that require large dependencies or optional runtimes are gated behind Cargo feature flags: url-ingestion for web crawling, transcription for audio-to-text, reranker for cross-encoder reranking, and flags for specific layout and OCR backends. Prebuilt language packages (pip, npm, etc.) and the Docker image bundle the common set of features; from-source builds enable only the features you select.
Six output formats are supported: plain text, Markdown, Djot, HTML, JSON tree, and Docling DocTags. Registered custom renderers can add additional formats. This flexibility allows the same extraction pipeline to feed different downstream consumers without repeated format conversion.
Content-hash caching, parallel batch processing, and per-file timeouts are built into the engine. The CLI exposes 14 commands. The REST API is launched with xberg serve. An MCP server mode is also available for integration with AI coding agents.
Installing Xberg in Python, Node.js, and Rust
For Python:
pip install xbergFor Node.js or TypeScript:
npm install @xberg-io/xbergFor Rust:
cargo add xbergThese three cover the most common integration paths. The Python package is built with maturin, which compiles the Rust core and exposes it as a Python module. The Node.js package uses N-API bindings. Additional bindings are available for Go (packages/go in the repository), Java (Maven at io.xberg/xberg), C# (NuGet at XbergIo.Xberg), Ruby (RubyGems), PHP (Packagist), Elixir (Hex), Dart (pub.dev), Swift, Zig, Kotlin, and WASM.
Docker and Helm are also available for deployment as a service. The Docker image is published at ghcr.io/xberg-io/xberg. Kubernetes deployment guidance is linked from the README.
The README notes that capabilities marked as requiring a feature flag (url-ingestion, transcription, reranker, and specific OCR backends) may not be present in all prebuilt packages. The documentation at docs.xberg.io covers which features are included in which distribution.
OCR, Layout Extraction, and Audio Transcription
OCR is one of Xberg's central capabilities. Four backends are supported: Tesseract (widely used open-source OCR), PaddleOCR, Candle (Rust-based inference), and VLM (vision-language model) backends. The engine supports fallback chains between backends, confidence scores, and language auto-detection. Plugins can extend the OCR backend set.
For documents with complex layouts, Xberg applies ML layout models to reconstruct reading order and identify structural elements. The README lists PP-DocLayout-V3 and RT-DETR as supported layout models, and TATR and SLANet for table structure reconstruction. This allows clean Markdown output from PDFs that contain multi-column text, figures, and embedded tables.
Audio and video transcription converts speech from MP3, M4A, WAV, WebM, and MP4 files to text via Whisper ONNX models ranging from the tiny model to large-v3. This requires the transcription feature flag.
Archive handling is recursive: .zip, .tar, .gz, and .7z files are traversed and the documents inside are extracted. The engine includes guards against zip-bomb attacks through compression ratio limits, nesting depth limits, and size bounds.
Code Intelligence, Embeddings, and Structured Extraction
Xberg extracts code intelligence from 371 programming languages: functions, classes, imports, symbols, and docstrings. This is designed for RAG pipelines that need syntax-aware chunking rather than arbitrary text splits. The README describes this as feeding clean, structured representations of code to downstream retrieval systems.
For embedding generation, Xberg supports both local ONNX models and 165 external embedding providers through liter-llm. Sparse and late-interaction embeddings are supported alongside dense embeddings, and cross-encoder reranking is available for search result refinement.
Structured extraction produces schema-driven JSON from any document. The README describes this as providing structured output straight from a document through local models (Ollama, LM Studio, vLLM) or hosted LLMs, without prompt engineering on the caller's side.
URL ingestion (requiring the url-ingestion feature) allows pointing Xberg at an HTTP or HTTPS URL. Three modes are available: single document extraction, document mode for a specific URL type, and crawl mode that follows links using the crawlberg engine. This makes Xberg usable as a web scraping and extraction backend.
Comparing Xberg to Docling and Its Predecessor Kreuzberg
Docling is an open-source document processing library from IBM Research. It converts PDFs and Office documents to Markdown or JSON and is designed for ingestion into AI pipelines. The distinction from Xberg is scope: Docling is primarily Python-focused and specializes in PDF and Office document formats. Xberg has a Rust core, 15 language bindings, and explicitly targets 107 document formats including audio, web URLs, archives, and source code in addition to PDFs and office files.
The related searches for Xberg include a direct comparison: xberg vs docling. The README positions Xberg as the answer to the problem of assembling a pipeline from multiple libraries, while Docling's own positioning is PDF-to-Markdown conversion quality. For teams that only process PDFs and Office files in Python, Docling is a simpler dependency. For teams that need broader format coverage or multiple language bindings, Xberg's wider scope is the relevant factor.
Xberg's predecessor, Kreuzberg, is documented in the README as kreuzberg-dev/kreuzberg-v4-lts. The README states that Xberg is the same engine rebuilt under a fresh v1 line. Teams migrating from Kreuzberg will find the same capabilities with a new package name and version numbering.
License, Maintenance, and Version History
Xberg is licensed under the MIT license. The Cargo.toml workspace sets the version at 1.3.0 and the minimum Rust version at 1.92. The last push to the repository was on 2026-09-17. The latest GitHub release at the time of writing is v1.2.9, published on 2026-09-24.
The release cadence appears active: v1.2.7, v1.2.8, and v1.2.9 were released on consecutive days in September 2026. This pace suggests active development with incremental improvements rather than large infrequent releases.
The repository is marked as auto-generated by alef, a readme generation tool referenced in the README comments. The README notes that the file should not be edited manually and provides commands for regenerating and verifying it. This is relevant for contributors: documentation changes flow through alef rather than direct edits to README.md.
Editorial conclusion
Xberg fits engineering teams that need to extract clean, structured content from a wide variety of document formats and cannot afford to assemble a pipeline from multiple per-format libraries. The 15 language bindings mean it integrates into most existing codebases without a language switch. The relevant trade-off is that several high-value capabilities (URL ingestion, transcription, reranking, and some OCR backends) are opt-in Cargo feature flags that must be enabled at build time; prebuilt packages for Python, Node.js, and other languages bundle a common set. The MIT license permits commercial use. For teams migrating from Kreuzberg, the README states that Xberg is the same engine rebuilt under a fresh v1 line.
Frequently asked questions
What happened to the Kreuzberg project?
The README states that Xberg is the next iteration of Kreuzberg (kreuzberg-dev/kreuzberg-v4-lts), the same document intelligence engine rebuilt and rebranded under a fresh v1 line. Teams using Kreuzberg should migrate to Xberg for continued updates.
Which OCR backends does Xberg support?
Xberg supports Tesseract, PaddleOCR, Candle, and VLM backends, with fallback chains between backends, confidence scores, and language auto-detection. OCR backends are extensible via plugins. Specific backends may require Cargo feature flags when building from source.
Can Xberg extract text from audio files?
Yes. Xberg transcribes speech from MP3, M4A, WAV, WebM, and MP4 files using Whisper ONNX models ranging from tiny to large-v3. This capability requires the transcription Cargo feature flag and may not be present in all prebuilt packages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xberg-io-xberg)