Model or dataset
drmingler/smart-llm-loader avatar
drmingler/smart-llm-loader

smart-llm-loader: LLM-Based Chunking for PDFs and Scanned Documents

smart-llm-loader is a lightweight yet powerful Python package that transforms any document into LLM-ready chunks. Spend less time on preprocessing headaches and more time building what matters. From RAG systems to chatbots to document Q&A, SmartLLMLoader handles the heavy lifting so you can focus on creating exceptional AI applications.

291 stars27 forksPythonMIT

At a glance

What is it?
A Python package that converts documents to markdown with OCR, then asks a multimodal model to split them into semantic chunks. Useful when table and header structure matters; expensive when it does not.
Who is it for?
Adopt smart-llm-loader if your RAG pipeline keeps failing on invoices, forms or scanned pages where headers and tables carry meaning, and you accept a paid multimodal model in the ingestion path. Do not adopt it for large plain-text corpora, air-gapped deployments, or any pipeline where per-page API cost is the binding constraint.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 58 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The preprocessing gap smart-llm-loader targets

Most document loaders stop at extraction. PyPDF pulls text runs out of a PDF, pdfplumber returns tables, OCR tools return a wall of characters. What none of them decide is where one idea ends and the next begins. For a contract or a report that is tolerable, because paragraph breaks survive the conversion. For an invoice, a lab result, or a scanned form, it is not: the seller block, the line-item table and the totals table are three different retrieval targets, and fixed-size character splitting will cut across all of them.

smart-llm-loader takes the position that the split is a semantic decision, and that a multimodal model is the right thing to make it. The README frames the package as handling "the entire document processing pipeline": conversion to markdown, OCR for scanned pages, chunking, and adapters for LangChain and LlamaIndex. The audience is developers building RAG systems, document Q&A, or chatbots over PDFs who would otherwise write that glue themselves.

The honest framing is that this is a convenience wrapper with an opinion, not a new extraction engine. The extraction underneath is pypdf and pdf2image; the model calls go through litellm. The value is in the prompt and the chunk boundary logic, and in the metadata the package attaches to each chunk.

How the pipeline works: markdown, then LLM segmentation

The data flow implied by the dependencies and the example output runs in stages. A file path or URL enters the loader. pypdf and pdf2image handle the PDF; pdf2image implies Poppler is doing the rasterisation, which is why the README lists Poppler as a system dependency rather than a pip package. Pages that are images rather than text go through OCR. The result is markdown, which is the intermediate representation the package works on.

Then the model call happens. The loader sends that markdown to whichever multimodal model you named, and the model returns chunks. The README's invoice example shows what comes back: an array of objects with a content field and a metadata object carrying page, semantic_theme and source. The themes in that example are invoice_header, seller_information, client_information, items_table and a summary table. The items table is preserved as a markdown table with its pipe delimiters intact, which is the part that matters for retrieval: a markdown table survives embedding better than a column of disconnected numbers.

That semantic_theme key is the package's real output, more than the text. It gives you a filterable field at query time. If a user asks about totals, you can restrict retrieval to summary-table chunks instead of hoping the vector similarity sorts it out.

The chunk_strategy parameter selects between this and cheaper modes. The signature lists 'page' and 'custom' alongside the default 'contextual', and custom_prompt lets you replace the segmentation instructions. Page-based chunking presumably skips the model's judgment about boundaries and just splits by page, which is the fallback when you want deterministic output.

Installing smart-llm-loader and running a first document

Poppler comes first. It is a system binary, not a Python package, so pip will not install it for you. On Ubuntu or Debian the README gives:

bash
sudo apt-get install poppler-utils

On macOS the equivalent is brew install poppler. On Windows the README points at the Poppler for Windows release archive and says to add its bin directory to your system PATH. If you skip this step, PDF rasterisation fails at runtime rather than at install time, which makes the error easy to misread as a code problem.

With Poppler in place, the package itself installs from PyPI:

bash
pip install smart-llm-loader

Poetry users get the same package with poetry add smart-llm-loader. Both are listed in the README.

A first run needs an API key for whichever provider you choose. The README sets keys through environment variables and names the model with a litellm-style provider prefix:

python
import os
from smart_llm_loader import SmartLLMLoader

os.environ["GEMINI_API_KEY"] = "YOUR_GEMINI_API_KEY"
model = "gemini/gemini-1.5-flash"

loader = SmartLLMLoader(
    file_path="your_document.pdf",
    chunk_strategy="contextual",
    model=model,
)
documents = loader.load_and_split()

After load_and_split returns, documents is a list. The README inspects documents[0].page_content and documents[0].metadata. What you should see, based on the invoice example, is a metadata dictionary containing page, semantic_theme and source. If semantic_theme is missing or the chunks look like fixed-size text blocks, the model call is not doing what you expect, and the first thing to check is whether the model string matches a provider litellm recognises.

The constructor also accepts url instead of file_path, api_key as an explicit argument, and save_output with output_dir if you want the chunks written to disk. The README does not document the output file format.

The cost and determinism problem with model-driven chunking

Every contextual chunking run is an API call to a paid multimodal model. The README acknowledges this indirectly by recommending Gemini Flash as the pairing, and links to an external article about Gemini Flash chunking performance. That link is the entire evidence base for the performance claim in the README; there is no benchmark in the repository itself.

The consequence is that ingestion cost scales with document count and page count, and it is not a one-time cost. Re-ingesting a corpus after a schema change means paying again. For a few hundred invoices that is fine. For a million-page archive it is a budget line, and page-based chunking becomes the pragmatic choice even though it throws away the semantic_theme metadata.

The second problem is determinism. A model asked to segment a document will not always draw the same boundaries twice. The README does not document any caching, retry or idempotency mechanism, and the pyproject dependency list contains nothing that would provide one. If your pipeline needs reproducible chunk IDs across runs, you are building that yourself on top.

There is also a version signal worth reading plainly. The repository's only release is v0.1.1, tagged "Initial Release", and pyproject carries the classifier Development Status :: 4 - Beta. The last push to the repository was on 2026-08-03. That is recent, but a single tagged release and a beta classifier describe a young package, and the README does not document a deprecation policy or a versioning scheme.

When a plain text splitter is the better answer

LangChain's own RecursiveCharacterTextSplitter is the obvious alternative, and the difference is not subtle. It splits on a character hierarchy with no model call, no API key, and identical output every time. For markdown notes, source code, HTML-stripped articles, or any corpus where paragraph structure already maps to meaning, it does the job at zero marginal cost.

The trade-off is exactly the invoice case. A recursive splitter does not know that a totals table is a unit. It will break a table across two chunks if the character count lands badly, and the numbers lose their column headers. smart-llm-loader's answer is to let a model see the whole page before deciding, which is why the README's comparison against PyMuPDF uses a formatted invoice rather than prose.

There is a middle path the README does not discuss: keep a deterministic splitter and add a separate table-extraction step for pages that need it. That is more code but keeps the ingestion path offline and reproducible. The choice comes down to whether your documents are structurally simple enough that a character splitter is adequate. If they are, smart-llm-loader adds a paid dependency and a source of run-to-run variance for no gain.

Licence, dependencies and what upgrades will cost you

The package is MIT licensed, and pyproject declares license = "MIT". MIT places essentially no conditions on use beyond preserving the copyright notice and permission text. That is the package's licence, not the licence of the models you call through it; Gemini, OpenAI and Anthropic each have their own terms, and nothing in the repository speaks to those. This is a description of the licence text, not legal advice.

The dependency surface is the practical upgrade concern. The package pins langchain ^0.1.0, langchain-community ^0.0.10 and langchain-core ^0.1.10. LangChain's ecosystem has moved well past those minor versions, and caret constraints on 0.x releases are narrow by design. A project already on a newer LangChain will hit resolution conflicts, and the README does not document a compatibility matrix or a tested range. litellm ^1.61.3 and tiktoken ^0.8.0 are similarly pinned.

Python support is declared as ^3.9, with a 3.12 classifier. There is no CI evidence in the repository listing beyond a .github directory, and the README does not state which Python versions are tested. If you are on 3.13, verify before assuming.

The upgrade path for a 0.x package with one release is effectively "read the diff". The README documents no migration guide, no changelog file appears in the top-level entries, and the release notes for v0.1.1 are titled Initial Release. Budget for reading source on each bump.

Editorial conclusion

Adopt smart-llm-loader if your RAG pipeline keeps failing on invoices, forms or scanned pages where headers and tables carry meaning, and you accept a paid multimodal model in the ingestion path. Do not adopt it for large plain-text corpora, air-gapped deployments, or any pipeline where per-page API cost is the binding constraint. Before committing, run one representative document through chunk_strategy='contextual' and inspect documents[0].metadata for the semantic_theme field, then repeat with 'page' to see how much structure the LLM pass is actually buying you.

Frequently asked questions

Does smart-llm-loader work with scanned PDFs and images?

Yes. The README lists built-in OCR for scanned documents and images as a feature, and the dependency list includes pdf2image. Poppler must be installed as a system dependency first, since pip does not provide it.

Which LLM providers can smart-llm-loader use?

Any multimodal model supported by litellm, because the package calls models through litellm. The README gives examples with gemini/gemini-1.5-flash, openai/gpt-4o and anthropic/claude-3-5-sonnet, each with its own API key environment variable.

What chunking strategies does smart-llm-loader support?

The constructor's chunk_strategy parameter accepts page, contextual and custom. Contextual is the default and uses the model to pick semantic boundaries; custom pairs with the custom_prompt argument to replace the segmentation instructions.

Is smart-llm-loader free to use?

The package itself is MIT licensed, so there is no licence fee. It calls a paid multimodal model for contextual chunking, so ingestion has an API cost that scales with document volume.

How do I install smart-llm-loader?

Install Poppler as a system dependency first, then run pip install smart-llm-loader or poetry add smart-llm-loader. On Ubuntu or Debian the README uses sudo apt-get install poppler-utils.

Official sources

  1. drmingler/smart-llm-loader on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/drmingler-smart-llm-loader.svg)](https://hysenlabs.com/projects/drmingler-smart-llm-loader)