# pycdc: a C++ disassembler and decompiler that targets any Python bytecode version

> Decompyle++ ships two binaries, pycdas and pycdc, built with CMake from a C++ codebase under GPL-3.0. It is aimed at engineers who need to read compiled Python that no longer has matching source, and its main constraint is that decompilation output is not guaranteed to be clean.

**zrax/pycdc** — C++ python bytecode disassembler and decompiler

- Repository: https://github.com/zrax/pycdc
- Stars: 4,633 · Forks: 879
- Language: C++
- License: GPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/zrax-pycdc

## What pycdc solves for people who only have a .pyc file

Python ships source as .py files and compiles them to .pyc bytecode for the interpreter. When the source is lost, a deployed artifact, a vendored dependency without a repository, or a build directory cleaned years ago, the .pyc is all that remains. Decompyle++ exists for that situation. It provides two programs: pycdas, which disassembles bytecode into a readable listing, and pycdc, which attempts to translate the same bytecode back into valid Python source.

The audience is narrow but real. Reverse engineers inspecting a shipped application, maintainers recovering a module from an old deployment, and people studying how a particular Python version compiles a construct. The project states its distinguishing goal directly in the README: it seeks to support bytecode from any version of Python, where other projects have achieved decompilation with varied success. That ambition is the reason the codebase carries its own bytecode tables and object model rather than leaning on the host interpreter.

## How the C++ pipeline turns bytecode into an AST

The repository layout shows the mechanism. pyc_module.cpp and pyc_code.cpp handle the container: a .pyc file is a header plus a marshalled code object, and the code walks that structure. pyc_object.cpp, pyc_numeric.cpp, pyc_string.cpp and pyc_sequence.cpp model the marshalled values that appear as constants. bytecode.cpp with bytecode_ops.inl holds the opcode definitions, which is where version differences get absorbed. ASTNode.cpp, ASTree.cpp and data.cpp build and hold the syntax tree that the decompiler emits as source. FastStack.h backs the evaluation stack the decompiler tracks while walking instructions.

The two front ends sit on top of that shared core. pycdas.cpp prints the disassembly, pycdc.cpp prints reconstructed source and sends errors to stderr. Because the parsing is done in C++ against its own tables, the tool does not need the matching Python interpreter installed to read an older or newer .pyc, which is the practical payoff of the design. The README also notes that both tools accept marshalled code objects as produced by marshal.dumps(compile(...)), which means the input does not have to be a file on disk in the usual .pyc container form.

## Building pycdc from source with CMake

There is no package on PyPI and the README does not describe a prebuilt installer, so the documented path is a source build. CMake generates the makefile or project file, then make builds the binaries. The README lists three CMake options for debugging: -DCMAKE_BUILD_TYPE=Debug for symbols, -DENABLE_BLOCK_DEBUG=ON for block debugging output, and -DENABLE_STACK_DEBUG=ON for stack debugging output. Those are useful when you are chasing a decompilation failure rather than just reading output.

```bash
cmake .
make
```

After the build, the two executables sit in the tree. To disassemble a file, pass the path as the only argument; the listing goes to stdout.

```bash
./pycdas [PATH TO PYC FILE]
```

To attempt source reconstruction, run the decompiler the same way. Source goes to stdout and any errors go to stderr, so redirecting the two streams separately is worth doing on a first run.

```bash
./pycdc [PATH TO PYC FILE]
```

For a marshalled code object rather than a .pyc container, both tools take -c -v <version>. The README is explicit about why: the objects themselves do not carry version metadata, so you must supply the version yourself.

```bash
./pycdc -c -v <version>
```

The repository also ships a test suite. On a Unix-like system or MSYS, make check JOBS=4 runs it, and FILTER=xxxx narrows the run to matching tests. That is the fastest way to confirm your build behaves as expected before you trust it on real input.

## Where pycdc fails and when it is the wrong tool

The README frames decompilation as an aim rather than a promise, and that wording matters. Bytecode loses information that source had: comments, original names in some cases, and the exact shape of expressions before the compiler flattened them. A decompiler reconstructs a plausible source file, not the original one. Expecting pycdc to return a file that is byte-identical to the lost source, or even one that always runs, is the wrong expectation to bring.

The version claim is also a goal, not a guarantee. Supporting any version of Python means the opcode tables and the AST logic have to cover constructs from many releases, and the README does not publish a compatibility matrix saying which versions decompile cleanly. If your .pyc comes from a very recent or very old interpreter, treat the first run as an experiment. When pycdc errors out on a file, pycdas on the same file is still useful: a disassembly listing is a direct rendering of the instructions and does not depend on reconstructing control flow into source.

One more boundary: this is not a Python library. There is no importable module and no pip install path in the documentation. If your pipeline is written in Python and you wanted to call a decompiler as a function, you would be shelling out to a compiled binary, and you would be carrying GPL-3.0 obligations into whatever you ship alongside it.

## pycdc compared with uncompyle6 and decompyle3

The obvious alternative family is the Python-based decompilers, uncompyle6 and its successor decompyle3. The difference in approach is not cosmetic. Those projects are written in Python and run under a specific interpreter version, using grammar rules tied to the bytecode of the versions they support. They install through pip, which makes them far quicker to try, and they historically produced clean output for the versions they targeted.

Decompyle++ inverts the trade. It is a C++ program built with CMake, it does not need a Python runtime to parse the bytecode, and it carries its own opcode tables so it can attempt versions its authors never ran. The cost is a build step, a compiled binary in your toolchain, and a decompiler whose README does not claim complete coverage. If you have a modern .pyc and a matching Python installation, a pip-installed decompiler is the lower-friction first attempt. If you have an old artifact, no matching interpreter, or you want disassembly and decompilation from the same binary, pycdc is the more direct fit.

## Licence and the cost of keeping a build current

Decompyle++ is released under the GNU General Public License, version 3, with the LICENSE file in the repository root as the reference. That is a copyleft licence, and it is worth reading before you link the code into anything you distribute. Using the pycdc and pycdas binaries as standalone tools for internal analysis is a different situation from embedding the C++ sources in a product you ship. This is a description of the licence identifier, not legal advice; the LICENSE file and your own counsel settle the question.

The upgrade cost is a source rebuild, not a package bump. There are no retrieved releases for this repository, so there is no published version number to pin and no changelog to scan before upgrading. You track the master branch, pull, re-run cmake and make, and re-run make check JOBS=4 to see whether the bundled tests still pass. The last push was on 2026-04-07, so the tree is not abandoned, but the absence of tagged releases means you should record the commit you built from if you need to reproduce a result later.

## Recovering a module from a .pyc: a first real run

A realistic first task is recovering one module from a compiled artifact. Start by disassembling rather than decompiling, because the listing tells you immediately whether the file is readable at all and which Python version compiled it. The output is printed to stdout, so redirect it to a file you can search.

```bash
./pycdas [PATH TO PYC FILE]
```

Read the top of the listing for the code object's metadata and the constant table. If the opcodes look sane and the constants are the strings and numbers you expected, the file parsed correctly and the decompiler has a fair chance. Then run the decompiler and keep the two streams apart, since the README says errors go to stderr.

```bash
./pycdc [PATH TO PYC FILE]
```

Check stderr first. An empty error stream is a good sign, not a guarantee. Then read the source output as a draft: verify the imports, the function signatures and any control flow that involves exception handling or comprehensions, because those are the places where a reconstructed tree is most likely to differ from what you remember writing. If the decompiler fails outright, fall back to the disassembly listing and reconstruct the logic by hand; that is slower but it does not depend on the AST stage succeeding.

## Conclusion

Adopt pycdc when you have a .pyc file whose source is gone and you need either a readable disassembly or a best-effort reconstruction, and you accept GPL-3.0 terms and a build step. Do not adopt it as a general-purpose Python toolchain component or expect the decompiler to always emit runnable code; the README presents decompilation as a goal, not a guarantee. Before relying on it, build the tree with CMake, run make check JOBS=4 to see how the bundled tests behave on your machine, and try pycdas first on your target file, since disassembly is the more deterministic of the two outputs.

## FAQ

### How can I convert a .pyc file to Python with pycdc?

Build the project with CMake and make, then run ./pycdc [PATH TO PYC FILE]. The decompiled source is printed to stdout and any errors are printed to stderr, so redirect the two streams separately if you want to inspect failures.

### How do I use pycdc?

Run ./pycdc [PATH TO PYC FILE] to attempt decompilation, or ./pycdas [PATH TO PYC FILE] to print a bytecode disassembly instead. For a marshalled code object rather than a .pyc container, pass -c -v <version> because the object carries no version metadata.

### How do I install pycdc?

The README documents a source build rather than a package install: generate a makefile or project with CMake, then run make. On a Unix-like system or MSYS you can run make check JOBS=4 to execute the bundled tests.

### How do I install pycdc on Windows?

The README describes generating a project file with CMake and opening it in a tool such as MSVC, then building from there. It also notes that the make-based test command works on MSYS, which is the documented route for running the tests on Windows.

## Sources

- [Issues](https://github.com/zrax/pycdc/issues)
- [License: GPL-3.0](https://github.com/zrax/pycdc/blob/master/LICENSE)
- [README](https://github.com/zrax/pycdc/blob/master/README.md)
- [zrax/pycdc on GitHub](https://github.com/zrax/pycdc)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zrax-pycdc
