ParamSpider filters out the URLs it calls boring, and never says which
Mining URLs from dark corners of Web Archives for bug hunting/fuzzing/further probing
At a glance
- What is it?
- A small Python tool that pulls archived URLs for a domain you name and drops the ones it considers uninteresting, with five documented flags. Its packaging metadata says version 0.1.0 while its tags say 1.0.x, the README is declared as Markdown but written as HTML, and the container clones from GitHub instead of copying a local tree.
- Who is it for?
- This fits someone doing authorised reconnaissance against a domain they own or have permission to test, who wants the parameters an application has historically accepted without sitting through a crawler. Three things to know first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 7 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The scope is yours to set, and the filter is not documented at all
The tool does one thing. Given a domain, or a file of domains, it fetches related URLs for them from a web archive and prints the results.
The domain always comes from you. There is a flag for a single domain and a flag for a file containing a list of them. There is no seeding from anywhere else, no discovery of adjacent domains, and no automatic widening of scope, so the set of hosts a run touches is the set you typed.
Then there is the filter, and it is the most consequential thing in the tool and the least explained. The description says it filters out URLs it considers boring, so that you can focus on the ones that matter most.
Nothing after that sentence explains it. There is no list of what gets dropped, no threshold, no output of what was removed, and no flag to turn filtering off and print everything. So the tool's central claim, that it leaves you with the interesting URLs, cannot be checked from the documentation, and you cannot get the unfiltered set without changing the code.
The archive is the other half of the mechanism, and it is what makes this safe to run broadly: the data is historical and already public, so a run against a domain you are assessing does not touch that host.
Five flags, and one of them substitutes your string into parameter positions
The usage documentation is a single command plus four flags, which is short enough to read in full.
The base command takes a domain. The list flag takes a file of domains instead. A streaming flag prints results as they arrive rather than at the end. A proxy flag takes an address in host and port form, and the example uses a port commonly associated with a local proxy. And a placeholder flag sets the string that is substituted for URL parameter values, with a stated default of a single uppercase token.
That last one is the interesting flag, and the example is what makes it clear. The command shown passes an HTML fragment containing a heading as the placeholder, which is not a value you would ever send to a server. It is a probe: the idea is that the archive knows a parameter once existed at that position, the tool rebuilds the URL, and the substituted string makes an injection attempt recognisable in a log or distinguishable from ordinary traffic.
So the default output already contains the marker token in every parameter position rather than a real value, and the flag exists so you can swap in something whose presence tells you where to look. That is a reasonable design for bug hunting, and it is also the reason the output should only be pointed at endpoints you are allowed to test.
Two smaller things. The streaming flag's own description in the README spells the word terminal incorrectly, which suggests the examples section was written quickly.
The packaging metadata says 0.1.0 and the tags say 1.0
This is the clearest defect in the repository, and it is checkable in three places.
The packaging script declares version 0.1.0. The two release tags are 1.0.0 and 1.0.1, the second dated seven days after the first, both in August 2023.
So the version a user installs from the README and the version in the release history are different numbers. And the README's own install route produces that mismatched number every time:
git clone https://github.com/devanshbatham/paramspider
cd paramspider
pip install .Following those three lines gives you a distribution labelled 0.1.0 regardless of what the tags say.
The timeline makes it worse. The last commit to the default branch is dated 2026-03-07. The newest release tag is from 2023-08-31. So there are roughly two and a half years of commits with no release, and the metadata was never brought up to the tag line.
For a tool this small the practical consequence is small too, since two dependencies is the whole install. But it does mean you cannot say which version you are running from the package metadata alone, and if you find the number in a log it will not correspond to any release.
The README is HTML, and the packaging declares it as Markdown
The long description for the package index is read from the README at build time and declared as Markdown content.
The README is not Markdown in any meaningful sense. Its visible body is raw HTML: a heading block with an alignment attribute, a subheading, a paragraph of centred links with HTML entities between them, and a star history badge. There is very little Markdown in it at all.
So the declaration and the file disagree, and the effect is that a package index renders the HTML as literal text. A reader looking at the package page sees angle brackets and entity references where they expected a formatted description.
The same script has one other omission worth noting: it declares no minimum Python version. The classifiers are not given in the visible text either, so the install tooling has no stated floor and will attempt the install on an interpreter the author never tested.
The dependency list is two packages, one HTTP client and one for coloured terminal output. That is the entire runtime surface, which for a tool that makes archive queries is all it needs, and it is the one part of the packaging that is unambiguously correct.
The container clones from GitHub and pins a retired Python
The container definition is seven commented steps and one base image, and two of its choices work against you.
The first is the base image, a slim Python image pinned to version 3.8. That release is retired upstream, so the image receives no security updates, and since the packaging declares no Python floor, nothing in the metadata prevents someone installing the same tool on a current interpreter and getting a different result.
The second is how the source gets in. The build installs git and then clones the repository from the forge, working inside the clone and installing from there. It never copies a local tree.
So the container is not built from the checkout you are looking at. If you have local changes they are not in the image. If you want a reproducible image you cannot achieve one, because the clone resolves to whatever the default branch holds at build time. And the build needs network access to the forge before it can build anything.
The entry point is simply the command name, so the container runs the tool with your arguments and nothing else, which is the right shape for this kind of image. The seven comments describing each step are unusually chatty for a Dockerfile, and they are accurate.
A navigation bar links to a contributing section that is not in the file
The README is a hundred and fifty words and a handful of examples, and two things about its shape are worth naming.
The first is the navigation bar. It is a row of centred links, each labelled with an emoji and separated by HTML entities, pointing at the about, installation, usage, examples and contributing sections.
Four of those five targets exist. The fifth does not. There is no contributing section in the file, so that link lands on nothing. For a project inviting outside contributions that is the wrong link to leave broken, and it is broken in a way that is invisible unless you click it.
The second is the root directory. Alongside the package directory and the packaging script there is a folder named static, with no explanation anywhere in the README or the packaging of what it holds. For a command line tool with two dependencies, a static asset directory is either leftover from a web interface that no longer exists or a placeholder, and neither is recorded.
There is one more piece of community plumbing worth noting, which is the star history badge at the bottom. For a project whose last commit was in March 2026 and whose newest tag is from 2023, that badge is the only element in the file suggesting anyone has touched it lately.
Editorial conclusion
This fits someone doing authorised reconnaissance against a domain they own or have permission to test, who wants the parameters an application has historically accepted without sitting through a crawler. Three things to know first. The filter is the tool's central mechanism and it is entirely undocumented, so run it and compare against raw archive results before you trust what it dropped. The version in the packaging metadata does not match the tags, so a local install and a release install are not the same artefact. And the placeholder option inserts your string into parameter positions, which is what makes the output useful for testing but also means the default output already contains a marker rather than a real value.
Frequently asked questions
how to install paramspider
Clone the repository, change into the directory and install it locally with pip. The packaging script declares two runtime dependencies, a requests client and a terminal colour library, and registers the command through a console script entry point. Note that the declared version in that script is 0.1.0 while the repository's release tags are 1.0.0 and 1.0.1.
how to use paramspider
Pass a domain with the single-domain flag, or a file of domains with the list flag. Other documented options stream results as they arrive, route requests through an HTTP proxy given as host and port, and set the string substituted into URL parameter values, which defaults to a single uppercase token.
how to install paramspider in kali linux
The README does not mention that Linux distribution or any other distribution by name. What it documents is a clone followed by a local install with pip, which is distribution independent and works anywhere the two declared dependencies resolve. The container definition is also distribution independent, built from a slim Python base.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/devanshbatham-paramspider)