Versatile OCR Program annotates exam figures for training data, and its roadmap outran its commits
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
At a glance
- What is it?
- raphael-seo/Versatile-OCR-Program is a config-driven, two-stage OCR pipeline that turns exam papers into JSON or Markdown with captions, educational rationales and related-topic lists for machine learning use. Its accuracy claim is an unexplained range, its sample tables keep Japanese headers, and the AI pipeline promised at the top of the README has not arrived.
- Who is it for?
- This project fits a specific job: building labelled training data from dense, exam-style scientific PDFs where a plain text extraction loses the figures and the tables. Judge it on the annotation quality rather than the OCR rate, because the image descriptions, educational rationales and topic lists are the actual deliverable.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 149 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The AI pipeline promised at the top never landed
The loudest thing in this README is a roadmap, and the roadmap did not arrive. A COMING SOON block announces a customizable AI pipeline with memory, tailored to a student's, researcher's or developer's field, released in less than one month. A second block then retracts the date: the author had planned to release before June but was simultaneously managing a critical exam on 15 June, decided that shipping something incomplete would be worse, and said development would resume afterwards with the public release to follow once the system was ready. Five months later the last commit to the repository is still 13 May 2026, and no pipeline code is in the tree. What is in the tree is the OCR system the author built to study with, described as having started as a personal tool that unexpectedly drew attention. That is worth keeping in mind when reading the feature list: the delivered project and the marketed project are not the same size.
v3.0 is a config-driven rewrite with the old versions parked
The current version is labelled v3.0_initial and dated 13 May 2026, and its banner describes a refactor into a modular, config-driven architecture. Three documents carry the load. setup_guide.md covers the path from clone to API keys to Docker to a first run, usage.md holds the CLI flags, scenarios, configuration reference and FAQ, and a dated release note under changes/ records what changed from v2. The earlier v1.0 and v2.0 code is preserved in a legacy directory and explicitly described as kept for reference but not maintained, so there are three codebases in the repository and one of them is the live one. The layout matches the two-stage design: a configuration file at the root, a second declarative YAML named auto_run, and two stage scripts, auto_run_stage1.py and auto_run_stage2.py, alongside a prompts directory and sample_images. The visible usage workflow starts at Step 1, initial OCR extraction, by running ocr_stage1.
Every figure gets a caption, a rationale and a topic list
The output shape is the point of the project, and it goes well beyond text. Extracted elements such as diagrams, tables and figures are semantically annotated with contextual explanations, including automatically generated natural language descriptions of visual content, with the stated example of a caption like this figure shows the process of mitosis in four stages. The biology sample shows the full template in practice: an image description, an Educational value paragraph explaining what a learner gains, a Related topics list, and an Exam relevance section enumerating the kinds of question that could use the image, seven of them in the example given. For the mathematics sample the same four sections appear around a solid geometry figure. That structure is aimed at training data rather than at reading, which is why the project frames itself as optimised for machine learning rather than as an OCR tool.
Accuracy is a range with no method attached
The feature list claims high accuracy and puts a number on it: over 90 to 95 percent on real-world academic datasets such as EJU Biology and UTokyo Math. Nothing else in the repository supports that figure. There is no sample size, no definition of what counts as correct, no per-element breakdown separating text from math from tables from figures, and no comparison against a baseline. Read next to the sample output it becomes easier to calibrate. In the biology example the system describes the four labelled cells with visible hedging, cell A appears to be in anaphase, cell B possibly telophase, cell C prophase or prometaphase, and cell D metaphase, on exactly the classification the exam asked a student to perform. That hedging is honest behaviour for a vision model, and it is also the kind of output you would want to review rather than ingest at scale without a human in the loop.
Extracted tables keep their original-language headers
The samples are taken from real material, a 2017 EJU Biology paper and a 2014 University of Tokyo Math paper, and the promise is that the semantic context comes back in English even when the source is Japanese. The table sample shows where that promise thins out. The extracted grid keeps its Japanese column headers for prophase, metaphase and anaphase, with eight option rows pairing letters A through D, and only the prose around it is translated: the Summary explains in English that each option corresponds to a specific mapping of the four cells to the three phases. So a downstream training pipeline receives a structure whose keys are Japanese and whose explanation is English. That is a detail a consumer of the JSON has to handle, and the visible documentation does not describe a normalisation step for headers, only a claim of English-translated outputs.
Seven services do the work and most of them are paid
The built-with line names DocLayout-YOLO, the Google Vision API, Gemini Pro Vision, MathPix OCR, the OpenAI API and OpenCV, and that list is the real architecture. Layout detection comes from DocLayout-YOLO, text and figure reading is split across Google's Vision API and MathPix, and the natural language descriptions that make the output valuable for training are generated by a vision-capable model from Gemini or OpenAI. The setup path in setup_guide.md starts with API keys for that reason. So the project is not a local model you can run on a laptop: the annotation layer, which is the part that distinguishes it from ordinary OCR, is a paid API call per figure. OpenCV appears to be the only component running locally, and the languages covered out of the box are Japanese, Korean and English, with more described as a customisation job rather than a config toggle.
A dependency manifest and a Dockerfile are missing at the top level
The repository root is short and shows the shape of the project: a gitignore, a LICENSE file, the README, the two automatic-run stage scripts, auto_run.yaml, config.yaml, changes, legacy, planned_features.md, prompts, sample_images, setup_guide.md, src and usage.md. Two absences are worth flagging for anyone trying to reproduce a run. There is no dependency manifest at the root, no requirements.txt or pyproject, even though the documented path is a Docker deployment, and there is no Dockerfile at the top level either, so the container step described in the setup guide is not defined by anything visible in this tree. The licence field recorded for the repository is empty while a LICENSE file exists, so what the code is licensed under has to be read from that file rather than from the repository metadata. There are no published releases either, and the version label lives in the README banner instead of a tag.
Editorial conclusion
This project fits a specific job: building labelled training data from dense, exam-style scientific PDFs where a plain text extraction loses the figures and the tables. Judge it on the annotation quality rather than the OCR rate, because the image descriptions, educational rationales and topic lists are the actual deliverable. Two things to verify before a run. The stated accuracy is a range with no method, sample size or per-element breakdown behind it, and several of its components are paid APIs. And the repository's last commit was 13 May 2026, so the AI pipeline advertised at the top of the README is not something you can install today.
Frequently asked questions
What is an OCR program?
Optical character recognition software extracts text and images from documents such as scans and PDFs. Versatile-OCR-Program does that and goes further: it also pulls out mathematical formulas, tables, diagrams and charts, annotates each visual element with a natural language description and an educational rationale, and emits JSON or Markdown shaped for machine learning datasets.
What does Versatile-OCR-Program extract from a PDF?
Multilingual text in Japanese, Korean and English, mathematical formulas, tables, diagrams and charts, including complex exam-style layouts with dense scientific content and formula-heavy paragraphs. Structured output comes as JSON or Markdown with human-readable descriptions of expressions, table summaries and figure captions.
How do I set up Versatile-OCR-Program?
The v3.0 setup guide covers clone, API keys, Docker and a first run, while usage.md documents the CLI flags, scenarios, configuration reference and FAQ. The v3.0 release notes are under changes. Earlier v1.0 and v2.0 code sits in legacy and is explicitly not maintained, and the usage workflow starts by running ocr_stage1.
Can Versatile-OCR-Program run without paid APIs?
Plan for provider credentials. It is built with DocLayout-YOLO, the Google Vision API, Gemini Pro Vision, MathPix OCR, the OpenAI API and OpenCV, and the documented setup path begins with adding API keys, since the descriptive annotations are what generate the per-figure captions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/raphael-seo-versatile-ocr-program)