pdf2docx has been handed over and its version comes from CI
Open source Python library for converting PDF to DOCX.
At a glance
- What is it?
- A PDF to DOCX converter that Artifex has relicensed under MIT and stopped maintaining, where the published metadata still lists Artifex as the author and a support address, the version falls back to 0.5.6a1 in a plain checkout, and the Makefile looks for a doc directory the repository does not have.
- Who is it for?
- pdf2docx is worth using if MIT licensing and a maintained upstream matter to you, because the licence change is explicit and the code is small enough to fork, and the dependency set is six named packages with visible floors. Two things to weigh before depending on it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 157 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The first section of the page is a handover notice
Before anything else the page says pdf2docx is no longer actively maintained by Artifex. The repository will remain available and has been relicensed under the MIT License so that the community can freely use, fork and maintain the project. Pull requests from community contributors are welcome, but Artifex no longer provides active development or maintenance. The last commit to the default branch is dated 2026-05-01. The same section then does something unusual for a handover: rather than pretending to be a recommendation, it redirects anyone who wants a full featured PDF processing library to PyMuPDF or MuPDF.NET, both of which are Artifex projects too. The company is pointing users at its own other tools rather than trying to keep this one.
The package metadata still lists Artifex for support
The build script fills in `author='Artifex'`, `author_email='[email protected]'` and `url='https://artifex.com/'`, with the license field reading MIT. So the artifact on PyPI is attributed to Artifex, offers a company support mailbox, and points at the company website, while the repository page states that the company no longer maintains it. Nobody updated the packaging fields when the handover happened, or chose not to. Either way, a user who has an install problem has a support address that the maintainers have said will not answer, and the metadata gives no indication of that. The licence is the one field that was updated, and the licence is the field nobody reads.
The version comes from a CI-written file with a 0.5.6a1 fallback
There is no version string in the source. The build script reads one from a file.
def get_version(fname):
'''Read version number from version.txt, created dynamically per Github Action.'''
if os.path.exists(fname):
with open(fname, "r", encoding="utf-8") as f:
version = f.readline().strip()
else:
version = '0.5.6a1'
return versionThe docstring says version.txt is created dynamically per Github Action, which means a release build gets its number from CI and a local checkout does not. The fallback is `0.5.6a1`, an alpha of a version seven releases behind the newest tag, so anyone running `python setup.py` or building a wheel locally produces an artifact numbered below anything currently published. The tag history is also uneven: v0.5.13 is titled 0.5.12 to 0.5.13, describing the change rather than the version, while v0.5.12 is titled with its own number.
Requirements are read as raw lines, newlines included
`install_requires` comes from `load_requirements("requirements.txt")`, and that function opens the file and appends each line it reads, with no stripping, no blank line filtering and no comment handling. Six packages end up in the dependency list.
PyMuPDF>=1.26.7
python-docx>=0.8.10
fonttools>=4.24.0
numpy>=1.17.2
opencv-python-headless>=4.5
fire>=0.3.0The floors are spread across years. numpy sits at 1.17.2 and fire at 0.3.0, both early, while PyMuPDF is at 1.26.7. `opencv-python-headless` is the one to watch, because it carries a floor of 4.5 and no ceiling, so a fresh install resolves to whatever OpenCV is current, and OpenCV is where Python ABI breaks usually surface first. PyMuPDF is not optional either: it is the PDF engine underneath the whole conversion path.
The command line entry point is a fire app
There is one console script, `pdf2docx=pdf2docx.main:main`, and the only argument parsing dependency in the list is `fire`, at 0.3.0. So the command line surface is a Python library entry point rather than a hand written argument parser, which means the flags it accepts are derived from the signature of `main` rather than declared anywhere. The documentation index names four ways in: a Convert PDF page, an Extract table page, a Command Line Interface page and a Graphic User Interface page. The GUI is documented but no GUI dependency appears in the six requirements, so either it is optional, or it is expected to be present already, or the page describes an entry point whose dependencies live elsewhere. The visible page does not say which.
No pyproject.toml, and the build calls setup.py directly
The tree has a setup.py, a MANIFEST.in, a Makefile and a requirements.txt, and no pyproject.toml. The build target is two legacy invocations.
@python setup.py sdist --formats=gztar,zip
python setup.py bdist_wheelDirect setup.py calls are the old path, and with no build-system table in the project, the build frontend has to guess which setuptools to use. Two other fields point the same way: `zip_safe=False`, and `include_package_data=True` with packages found by `find_packages` excluding `build`, `dist` and `test`. That last part is why MANIFEST.in exists, since the contents of the installed package are decided by that file rather than by anything declared in Python.
make doc looks for a doc directory the tree does not have
The Makefile sets `DOCSRC := $(TOPDIR)/doc`, singular. The repository's entry is `docs/`, with an s. The doc target is guarded, so it checks whether `$(DOCSRC)/Makefile` exists and does nothing at all if it does not, which means the mismatch produces no error and no build. The same guard wraps the test target, which runs `make test` inside `test/`, and the clean target, which recurses into both directories and swallows failures with `|| exit 0`. Clean then removes three hardcoded names, `.pytest_cache`, `pdf2docx.egg-info` and `dist`, plus the build directory. The dependency for the doc target is not declared anywhere either: it is a comment reading `# pip install sphinx_rtd_theme`, so the theme has to be installed by hand before `make doc` would work even if the directory matched.
The page is 101 words and its last heading has nothing under it
The whole visible page is the status notice and a documentation index. Eight links: Installation, Quickstart, Convert PDF, Extract table, Command Line Interface, Graphic User Interface, Technical Documentation, and API Documentation. Then a heading reading Sample, and nothing after it in this copy. The documentation itself is hosted on readthedocs, configured through a `.readthedocs.yaml` at the root, so the repository carries no tutorials, no examples directory and no sample input file. One asymmetry is worth noting before you plan around it: the technical documentation page is labelled as being in Chinese, while the quickstart pages are not, so the in-depth material and the entry level material are in different languages.
Editorial conclusion
pdf2docx is worth using if MIT licensing and a maintained upstream matter to you, because the licence change is explicit and the code is small enough to fork, and the dependency set is six named packages with visible floors. Two things to weigh before depending on it. Artifex states plainly that it no longer provides active development or maintenance, and the last commit is dated 2026-05-01, so nothing here will be patched by the company that wrote the PDF engine underneath it. And the packaging is old in a way that bites locally: no pyproject.toml, a build driven by direct setup.py calls, a version read from a file CI has to generate, and requirements loaded as raw lines.
Frequently asked questions
how to use pdf2docx
The repository carries no example, and the page sends you to its documentation, which names four entry points: Convert PDF, Extract table, a Command Line Interface and a Graphic User Interface. The library exposes one console script, pdf2docx=pdf2docx.main:main, built on the fire package rather than a hand written argument parser.
how to install pdf2docx
The page has no install command; installation is documented at pdf2docx.readthedocs.io. The packaging declares python_requires >=3.10, six dependencies read from requirements.txt, and one console script. Note that a checkout without a CI generated version.txt builds as version 0.5.6a1.
Is pdf2docx still maintained by Artifex?
No. The page states that pdf2docx is no longer actively maintained by Artifex, that the repository remains available and has been relicensed under the MIT License so the community can maintain it, and that Artifex no longer provides active development or maintenance. The last commit is dated 2026-05-01.
What does pdf2docx depend on?
Six packages: PyMuPDF, python-docx, fonttools, numpy, opencv-python-headless and fire. The floors range from numpy 1.17.2 and fire 0.3.0 up to PyMuPDF 1.26.7, and opencv-python-headless has a floor of 4.5 with no upper bound.
What is a good pdf2docx alternative?
The repository does not compare itself to other converters. It points at PyMuPDF and MuPDF.NET for anyone who wants a full featured PDF processing library, and otherwise leaves the choice open. Its own scope is PDF to DOCX.
Can pdf2docx convert other formats, such as SDOCX, to PDF?
No. The library is built for one direction, PDF to DOCX, and its documented entry points are Convert PDF, Extract table, a Command Line Interface and a Graphic User Interface. Nothing in the repository covers SDOCX or any DOCX to PDF conversion.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/artifexsoftware-pdf2docx)