s0md3v/Photon: an OSINT crawler that saves what it finds
Incredibly fast crawler designed for OSINT.
At a glance
- What is it?
- Photon is a Python crawler built for reconnaissance rather than indexing. It walks a target, extracts emails, keys, files and JavaScript endpoints, and writes them into a per-domain loot folder. Here is how it installs, what the extraction actually covers, and where the design gets in the way.
- Who is it for?
- Adopt Photon when you want a single pass over a site that ends with a directory of emails, files, parameterised URLs and JavaScript endpoints you can grep, and when the target is static enough that a requests-based crawler can see it. Skip it if the site renders client side, if you need a crawl to resume after a crash, or if you need a compliance story the tool itself does not provide.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 26 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Photon is for, and who actually needs it
Photon is not a search engine crawler and not a sitemap generator. The README describes it as an "incredibly fast crawler designed for OSINT", and the feature list makes the intent concrete: while walking a site it collects URLs in and out of scope, URLs that carry parameters such as example.com/gallery.php?id=2, emails and social media accounts, files like pdf and png, secret keys including auth and API keys and hashes, JavaScript files together with the endpoints referenced inside them, strings matching a regex you supply, and subdomains with related DNS data.
The audience follows from that list. A penetration tester mapping a target before writing a report, a bug bounty hunter looking for an exposed key or an undocumented endpoint, or a security team auditing its own external surface. Someone who wants a searchable index of page text should look elsewhere. The output is a folder of findings, not a corpus.
One framing note matters. The repository is not archived and the last push was on 2026-09-04, but the most recent tagged release is v1.3.0 from 2019-04-05. The README's claim that Photon "is under heavy development" with updates "being rolled regularly" sits awkwardly next to a release history that stops in 2019. Treat the code as stable and the release notes as stale.
How the crawl works and where the data lands
The repository layout is small: photon.py at the top level, a core/ package, a plugins/ directory, requirements.txt, a Dockerfile and a CHANGELOG.md. The dependency list is four entries, requests, requests[socks], urllib3 and tld, which tells you the fetching layer is plain HTTP with optional SOCKS proxying and that domain parsing is handled by the tld package rather than a full browser engine.
That choice defines the mechanism. Photon issues HTTP requests, parses the responses, and follows links it finds. Anything a site renders with client-side JavaScript after the initial response is outside what the crawler can read. The tld dependency is what lets the crawler decide whether a discovered link is in scope or out of scope, which is why the README distinguishes the two categories in its output.
Results are written per domain. The Docker instructions mount a folder at /Photon and the README's example mounts the current directory as /Photon/google.com, so the target name becomes the loot folder. Inside that folder the categories from the feature list are separated: URLs, files, JavaScript, keys, and so on. The Exporter plugin, documented under "export formatted result", converts the saved output into JSON for downstream tooling.
Two plugins extend the seed set and the DNS side. The wayback plugin pulls URLs archived by archive.org and uses them as seeds, which is the documented answer to the problem that a live crawl only sees what is reachable now. The dnsdumpster plugin handles DNS data dumping. Both are documented on the project wiki rather than in the README body.
Installing Photon and running a first crawl
The README gives a Docker path and the repository ships a Dockerfile, so that is the least ambiguous way to start. The image is built on python:3-alpine, clones the repository during the build, installs requirements.txt, declares /Photon as a volume, and sets the entrypoint to python photon.py with --help as the default command. The README describes the resulting image as lightweight at 103 MB.
Build the image from a clone of the repository:
git clone https://github.com/s0md3v/Photon.git
cd Photon
docker build -t photon .Run it against a target. The -u flag is the one the README uses for the URL, and the container name photon is what you inspect later to find the volume:
docker run -it --name photon photon:latest -u google.comIf you would rather have the results in a directory you choose, the README mounts the current directory at the target's loot path. Replace the target in both places consistently:
docker run -it --name photon -v "$PWD:/Photon/google.com" photon:latest -u google.comAfter the run finishes, the loot folder holds the extracted categories. The README points to docker inspect photon for locating the volume if you did not mount one. A non-Docker install is implied by requirements.txt and photon.py but the README does not spell out the pip steps, so follow the wiki's compatibility and dependencies page if you go that route. The README also documents an --update option for installing and checking for updates, and states that updating does not lose saved data.
Where Photon stops being the right tool
The HTTP-only design is the first boundary. Single-page applications that assemble their DOM in the browser will hand Photon an empty shell. The README's JavaScript extraction reads endpoints out of script files it can fetch, not out of a rendered page, so a site whose routes only exist after hydration will look far smaller than it is. If your target is one of those, a crawler with a real browser engine is the correct instrument.
The second boundary is state. Nothing in the README or the repository layout describes resuming an interrupted crawl, checkpointing visited URLs, or exporting the frontier. A long run that dies has to start over. For a small site this is irrelevant. For a large one it is the difference between a tool you use and a tool you abandon.
The third is politeness and load. The README lists timeout and delay among the configurable options and admits that "crawling can be resource intensive". Those two facts belong together: the defaults are not described in the README, so you should set delay deliberately before pointing Photon at anything you do not control, and you should know that concurrency is a setting you own rather than a safety net.
Finally, the release cadence. The last tagged release is v1.3.0 from 2019. The repository is not archived and the last push was on 2026-09-04, so commits happen, but a user who needs a documented changelog for a recent version will not find one after 2019. Plan for reading the source when behaviour surprises you.
Photon against a general-purpose crawler
The obvious alternative is a scraping framework such as Scrapy, and the difference is not speed. It is what the tool considers a result.
Scrapy gives you a spider abstraction, an item pipeline, a scheduler with a persistent queue, and a settings surface for concurrency, delays and retries. You write the extraction rules. Photon gives you the extraction rules already written and a directory structure to receive them, at the cost of the scheduler and the pipeline being whatever the project implemented rather than something you configure.
So the trade is inverted. With Scrapy you spend the first hour writing a spider and the next month tuning it. With Photon you spend five minutes on a command line and then discover that the categories it extracts are the categories you get. If your reconnaissance question is exactly "what emails, keys, files, parameters and JS endpoints does this site expose", Photon answers it before you have finished reading Scrapy's tutorial. If your question is anything else, you are fighting the tool.
There is a middle position worth naming. The README links a Photon Library page on the wiki, which implies the crawler can be driven as a library rather than only through photon.py. That is the path to take if you want Photon's extractors inside a pipeline you control, and it is the path the README gives the least detail about.
Licence, maintenance and the cost of keeping it running
Photon is GPL-3.0. If you run it as a command line tool, the licence obligations are the ordinary ones. If you take the code, modify it, and distribute the result, or embed it in something you ship, the copyleft terms apply to the derived work. That matters for anyone planning to fold Photon's extractors into a commercial product, and it is worth reading LICENSE.md in full rather than assuming the boundary. This is a description of the licence, not legal advice.
Upgrade cost is low and slightly strange. The README documents --update as both the install and the check, and states that updating preserves saved data, so the mechanism exists. But with the newest tag dating to 2019, an update mostly moves you along the master branch rather than between releases. You are tracking a branch, which means you should keep your own record of which commit you ran. The CHANGELOG.md file in the repository root is the place to look for what changed.
Operationally, the recurring costs are the ones the tool cannot absorb: proxy or SOCKS configuration if you need it, since requests[socks] is in the dependency list, and the storage for the loot folders, which grow with the number of files and JavaScript assets the target serves.
Editorial conclusion
Adopt Photon when you want a single pass over a site that ends with a directory of emails, files, parameterised URLs and JavaScript endpoints you can grep, and when the target is static enough that a requests-based crawler can see it. Skip it if the site renders client side, if you need a crawl to resume after a crash, or if you need a compliance story the tool itself does not provide. Verify first that your Python environment satisfies the four entries in requirements.txt, that you can write to the loot directory, and that you have authorisation for the target, because nothing in the tool stops you from pointing it at a domain you do not own.
Frequently asked questions
How do I install s0md3v/Photon with Docker?
Clone the repository, change into the directory, and run docker build -t photon . to build the image. The README then shows docker run -it --name photon photon:latest -u google.com to crawl a target.
How do I use s0md3v/Photon to crawl a site?
Pass the target with the -u flag, for example docker run -it --name photon -v "$PWD:/Photon/google.com" photon:latest -u google.com. The README states that results are saved to a per-domain loot folder, which the volume mount lets you place where you want.
What kind of data does s0md3v/Photon extract?
The README lists URLs in and out of scope, URLs with parameters, intel such as emails and social media accounts, files like pdf and png, secret keys including auth and API keys and hashes, JavaScript files and the endpoints inside them, strings matching a custom regex, and subdomains with DNS data.
Can s0md3v/Photon crawl sites that render with JavaScript?
The dependency list is requests, requests[socks], urllib3 and tld, so the fetching layer is plain HTTP. Content that only appears after client-side rendering is not part of what the crawler reads, though it can still pull endpoints out of JavaScript files it fetches.
How do I update s0md3v/Photon without losing results?
The README documents an --update option for installing and checking for updates, and states that updating does not lose saved data. The most recent tagged release listed is v1.3.0 from 2019-04-05, so an update largely follows the master branch.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/s0md3v-photon)