jonaslejon/malicious-pdf: a generator for 67 malicious PDF test files
đź’€ Generate malicious PDF test files for testing phone-home callbacks, SSRF, XSS, NTLM credential theft, and data exfiltration in PDF viewers, converters, and web applications. Can be used with Burp Collaborator or Interact.sh
At a glance
- What is it?
- The project produces a directory of PDFs that phone home, fetch remote resources or leak credentials, so you can point them at a Burp Collaborator or Interactsh listener and see what your upload pipeline actually does. It is a testing tool, not a scanner, and the README is honest about that.
- Who is it for?
- Adopt it if you run a file-upload endpoint, a PDF converter or a document-processing pipeline and you already have a Collaborator or Interactsh listener to receive callbacks. Do not adopt it if you are looking for a scanner that tells you whether a PDF is malicious; this project only produces the malicious side.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 34 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem jonaslejon/malicious-pdf solves, and who needs it
Upload endpoints are usually tested with a benign PDF and a wrong-content-type PDF. Neither tells you whether the server-side renderer resolves external references, whether a converter shells out, or whether a viewer leaks an NTLM hash when it opens a file. The README lists the intended targets plainly: web pages and services that accept PDFs, security products, PDF readers, PDF converters, server-side processing libraries such as PDFBox and iText, and PDF static analysis tools. The author's stated reason for building it was needing a batch of PDFs with different links rather than one hand-crafted sample. That framing matters. This is a corpus generator for red teamers, bug bounty hunters and pentesters who already know what a callback means, not a defensive tool. The README says it is for educational and professional purposes only, and the repository ships a SECURITY.md and a CONTRIBUTING.md alongside the script.
How the generator builds each PDF and where the callbacks go
The entry point is a single Python file, malicious-pdf.py, with one function per test case. The complete test matrix in the README maps each output file to a function name, a CVE or research reference, an attack vector, a method and an impact. test1.pdf, for example, is produced by create_malpdf() and uses a /GoToE action with a UNC path for external file access, tied to CVE-2018-4993; test1_1.pdf uses the same function with an HTTPS URL instead. test2.pdf comes from create_malpdf2() and builds an XDP form whose submit event fires automatically, which is how form data leaves the document. test3.pdf uses /OpenAction with app.openDoc() to load an external document. test4.pdf injects an external XSLT stylesheet into XFA for a UNC path callback, referencing CVE-2019-7089. test5.pdf is a plain /URI action that produces a DNS prefetch. The pattern is consistent: the URL you pass on the command line is substituted into each template, and the file is written to disk. Nothing is sent by the generator itself. The callback only happens when a viewer, converter or scanner opens the file and follows the reference, which is why the tool pairs with an out-of-band listener.
Installing it and running your first batch against Interactsh
The README gives a two-command install. requirements.txt contains a single dependency, validators, so the install is quick and there is no build step.
pip install -r requirements.txt
python3 malicious-pdf.py burp-collaborator-urlReplace burp-collaborator-url with the hostname your listener gives you, for example an Interactsh URL. The output goes to the output/ directory as test1.pdf, test2.pdf, test3.pdf and so on. You should see the files appear there, and you can confirm the count against the test matrix in the README.
Two options change where the files land and what is embedded in them. The --output-dir flag sets the destination directory, and --no-credit suppresses the attribution metadata that is otherwise written into each PDF.
python3 malicious-pdf.py https://your-interact-sh-url --output-dir DIR --no-creditThe obfuscation flag takes a level from 0 to 3 according to the options block, with a level 4 described separately in the README. Level 1 hex-encodes PDF names and octal or hex-encodes strings. Level 2 adds JavaScript bracket notation and case and whitespace variation in javascript: URIs. Level 3 adds FlateDecode stream compression. Level 4 wraps JavaScript payloads in a base64 decoder stub so the original API calls never appear as literal substrings, which is the setting aimed at defeating naive /JS regex scanners.
python3 malicious-pdf.py https://your-interact-sh-url --obfuscate 2
python3 malicious-pdf.py https://your-interact-sh-url --obfuscate 4After generating, upload the files one at a time to the target and watch the listener. A DNS lookup or HTTP request arriving at your listener is the finding; which file produced it tells you which vector the target honours.
The obfuscation levels are for testing your detection, not for evasion in production
It is worth being precise about what the obfuscation flag is for. The README states that staged JavaScript payloads such as form-field /V values and the base64 decoder defeat naive /JS regex scanners, and the level 4 description says the original API calls never appear as literal substrings. That is a statement about static analysis tools, and it means the generated corpus doubles as a test set for whatever scanner sits in front of your upload endpoint. If your scanner only greps for /JS, level 4 files will pass it and the callback will still fire when the file reaches a renderer. If your scanner does structural parsing or executes in a sandbox, the result will differ. The project does not ship a scanner and does not claim to, so the only way to know where your own tooling lands is to run your tooling against the output. Note the options block lists levels 0 through 3 while the README documents level 4 in a separate paragraph; the two descriptions are not in one table, so read both before scripting a loop over levels.
What the README does not tell you, and where the tool is the wrong choice
The README does not document per-viewer or per-converter results. It lists what each file is designed to do, not which versions of Adobe Reader, Foxit, PDF.js, PDFBox or iText actually execute it. Some of the referenced CVEs are old, and a patched viewer will simply ignore the action. So a silent listener after uploading test4.pdf is not proof the target is safe; it may mean the renderer never processed the XFA stylesheet. Treat every negative result as inconclusive until you have a positive control, for example a plain /URI file such as test5.pdf that should produce a DNS prefetch on almost anything.
The second limitation is scope. This is a generator, not an analyser. If your actual question is whether a PDF you received is malicious, this project cannot answer it. The related searches around checking a PDF for a virus, repairing a malicious PDF or whether previewing a PDF in Outlook is safe describe defensive problems, and the repository contains nothing aimed at them. There is also no rollback or cleanup path documented, which matters because you are deliberately writing files that trigger network requests; keep them in a directory you control and do not leave them in a shared upload folder. Finally, the tool assumes you have an external listener. Without Burp Collaborator or Interactsh, the output is a folder of PDFs and no way to observe whether anything happened.
How it differs from Bad-Pdf and from a general upload scanner
The README credits Bad-Pdf among its sources, and the difference in approach is the breadth of the corpus rather than the technique. Bad-Pdf is a single-purpose NTLM hash capture tool: it builds one PDF that triggers an SMB request so a responder can collect a hash. jonaslejon/malicious-pdf covers that ground as one vector among many, and adds SSRF, XXE, XSS, XFA form submission, XSLT stylesheet callbacks, URI prefetch and data exfiltration, spread across 67 files with a documented matrix. The other comparison is Burp Suite UploadScanner, also credited. UploadScanner is a Burp extension that mutates uploads and observes responses inside a proxy session; malicious-pdf is a standalone script that writes files you then feed to the target by whatever means you prefer, including a proxy. If your workflow is already inside Burp and you want automated mutation of every upload field, UploadScanner fits better. If you want a fixed, named set of PDFs you can hand to a colleague, attach to a report, or run against a converter outside a browser, the generated corpus is the more direct artifact.
Maintenance, licence and the cost of keeping up
The repository is not archived, and the last push was on 2026-08-27, which is recent enough that the test matrix reflects current research, including the ExpMon Adobe Reader analysis from April 2026 that the README credits as the inspiration for test33_13, test33_14, test33_15 and obfuscation level 4. Releases are sparse: v1.0.0 and v1.0.1 both landed on 2026-04-20. The upgrade cost is therefore low in dependency terms, since requirements.txt pins nothing and lists only validators, but it is non-trivial in knowledge terms. Each new CVE in a PDF renderer is a candidate for a new test function, and the value of the corpus decays as viewers patch. Pulling a new commit may add files to output/ without warning, so script against a named list rather than a glob if your test harness counts files. The licence is BSD-2-Clause, which is permissive and permits commercial use and modification with the copyright notice retained; the repository includes the LICENSE file, and if you redistribute generated PDFs you should check how the --no-credit flag interacts with attribution expectations. That is a description of the licence text, not legal advice.
Editorial conclusion
Adopt it if you run a file-upload endpoint, a PDF converter or a document-processing pipeline and you already have a Collaborator or Interactsh listener to receive callbacks. Do not adopt it if you are looking for a scanner that tells you whether a PDF is malicious; this project only produces the malicious side. Before you use it, verify two things on your own setup: which of the 67 files your target viewer actually renders (the README does not document per-viewer results), and whether the /JS and /OpenAction tokens survive your own static analysis tooling at the obfuscation level you chose. The repository's last push was on 2026-08-27, so the test matrix is current as of that date.
Frequently asked questions
Can a PDF be a malicious file?
Yes. jonaslejon/malicious-pdf generates 67 PDF test files that use mechanisms such as /GoToE actions, /URI actions, XFA form submission, external XSLT stylesheets and /OpenAction JavaScript to make a viewer, converter or scanner contact a remote host. The project exists specifically because PDFs can carry those actions.
How do I check for a malicious PDF with this project?
You do not. jonaslejon/malicious-pdf only generates malicious PDFs; it contains no scanning or analysis capability. To check a PDF you received, you need a separate analysis tool, and the project's own README frames itself as the offensive side of that test.
Can a virus come from a PDF?
The README does not make that claim and the repository does not ship a virus. What it demonstrates is that a PDF can carry actions that trigger network callbacks, credential leaks and data exfiltration when opened or processed, which is the class of behaviour the test files are built to exercise.
What should I do if I open a malicious PDF?
The repository does not document incident response steps, so it cannot answer this. What it does tell you is that the callbacks are triggered by the viewer following references in the file, which is why the README recommends using the generated files with a Burp Collaborator or Interactsh listener so you can observe whether a request was made.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jonaslejon-malicious-pdf)