Open-source project
stamparm/ipsum avatar
stamparm/ipsum

stamparm/ipsum: a daily bad-IP feed with occurrence scores, and how to wire it into ipset

Daily feed of bad IPs (with blacklist hit scores)

2,376 stars178 forksUnknownUnlicense

At a glance

What is it?
IPsum aggregates 30+ public blocklists into one daily file where each address carries a count of how many lists it appeared on. It is a plain-text feed, not a service, and the count is the whole point.
Who is it for?
Adopt IPsum if you already run iptables or ipset and want a scored, plain-text blocklist you can filter by occurrence count before it touches your ruleset. Do not adopt it if you need per-IP context, attribution or a hosted lookup API; the repository is a file, not a service, and the README documents no rollback path.
Can I use it commercially?
Yes. Unlicense is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What IPsum actually solves, and who it is for

Individual blocklists disagree. One list flags an address, another does not, and an operator who subscribes to a dozen of them ends up maintaining a dozen parsers, a dozen formats and no shared notion of confidence. IPsum collapses that work: it retrieves and parses more than 30 publicly available lists of suspicious or malicious IP addresses, including IPnoise, on a daily basis, and publishes the merged result into this repository. The output is not a verdict on any single address. It is an occurrence count, meaning how many of the source lists contained that IP. The README states that higher counts generally mean higher confidence and fewer false positives when blocking inbound traffic, and the file is sorted by that count from highest to lowest. The audience is narrow and practical: someone running a server or a network edge who wants to drop inbound packets from addresses that several independent sources already distrust. That is a firewall consumer, not a SOC analyst. If you need to know why an address is listed, IPsum will not tell you.

The mechanism: 30+ lists in, one scored file out

The pipeline is a scheduled aggregation job, not a query service. According to the README, all source lists are automatically retrieved and parsed every 24 hours and the final result is pushed to this repository. Nothing about that description implies a running API, a database or a daemon on your side. You consume a file over HTTPS from raw.githubusercontent.com, or you consume a preprocessed subset from the levels/ directory. The repository layout confirms the shape of the output: LICENSE, README.md, ipsum.txt and levels/. The main file carries a header (the README's own examples strip lines beginning with #), then one IP per line followed by its occurrence count as a whitespace-separated field. The levels/ directory holds raw IP lists bucketed by that count, so levels/3.txt contains addresses found on three or more blacklists. That bucketing is the design decision worth noticing: instead of asking consumers to parse and filter, the project precomputes the common thresholds. The trade-off is that you now depend on the project's choice of buckets. If you want a threshold of five, you either take levels/5.txt or you filter ipsum.txt yourself. The README does not document how the bucket files are generated or how often they are rewritten relative to ipsum.txt.

Installing IPsum: there is nothing to install

IPsum has no package, no binary and no service to start. The README's examples fetch the feed directly from the repository, which means the setup cost is the cost of the tool you already use to block traffic. The README gives a one-liner for producing a deployable list of bad IPs that appear on at least three blacklists. It strips comment lines, drops entries whose trailing count is 1 or 2, and keeps the first field:

bash
curl -fsSL https://raw.githubusercontent.com/stamparm/ipsum/master/ipsum.txt 2>/dev/null | grep -v "^#" | grep -Ev '[[:space:]]([12])$' | cut -f 1

What you should see is a stream of bare IP addresses, one per line, with no counts and no comments. That stream is what you feed to a blocking tool. The README's ipset example is the more realistic first use, and it is written as a root shell session. It installs the two packages, flushes and recreates a set named ipsum, adds every filtered address to it, then removes any stale INPUT rule and inserts the drop rule at the top of the chain:

bash
sudo -i
apt-get update && apt-get install -y iptables ipset
ipset -q flush ipsum
ipset -q create ipsum hash:ip
for ip in $(curl https://raw.githubusercontent.com/stamparm/ipsum/master/ipsum.txt 2>/dev/null | grep -v "#" | grep -Ev '[[:space:]]([12])$' | cut -f 1); do ipset add ipsum $ip; done
iptables -D INPUT -m set --match-set ipsum src -j DROP 2>/dev/null
iptables -I INPUT -m set --match-set ipsum src -j DROP

After this runs, inbound packets whose source is in the set are dropped before they reach your services. Note what the snippet does not do: it does not persist the set across reboots, it does not schedule the refresh, and it does not tell you how to undo the rule. The README stops at the point where the rule is live.

The occurrence count is a heuristic, and the Wall of Shame shows its limits

The scoring model assumes that independent lists agreeing on an address is evidence. That assumption breaks in a specific and visible way. The README's Wall of Shame table for 2026-09-28 lists the highest-scoring addresses with their DNS lookups, and a large share of the eight- and nine-count entries resolve to scanner infrastructure: multiple 66.132.x.x addresses under censys-scanner.com, einstein.census.shodan.io and sky.census.shodan.io, and a long run of o0NN.scanner.modat.io hosts. Those are research and scanning services, not compromised hosts. If your threat model is opportunistic scanning noise, blocking them is exactly what you want. If your threat model is targeted intrusion, an eleven-count address that turns out to be a census scanner is a false positive with a high confidence score attached. The count measures how many lists include an address, not how malicious it is, and the README says as much by describing the count as a confidence signal rather than a severity one. There is a second limitation with no workaround in the documentation: the feed is inbound-oriented. The README frames the count in terms of blocking inbound traffic, and the README describes no outbound filtering, no per-IP metadata, no ASN or geolocation enrichment, and no historical record of when an address entered and left the list. You get today's file. Yesterday's is not archived in the repository.

How IPsum differs from pulling a single curated blocklist

The obvious alternative is to subscribe to one well-maintained blocklist and skip the aggregation entirely. The difference is in what each gives you when you are wrong. A single-list feed gives you a binary answer and one operator's judgement; if that list is aggressive you block legitimate traffic, and if it is conservative you miss addresses its peers already flagged. IPsum gives you a number that lets you set your own tolerance, which is why the levels/ directory exists and why the README's own examples filter at three. That flexibility has a cost the single-list approach does not have: you must decide the threshold yourself, and the README offers no guidance beyond the statement that higher counts mean fewer false positives. A second alternative is a commercial or hosted threat-intelligence API, which typically adds context such as first-seen timestamps, tags and attribution. IPsum has none of that, and in exchange it has no account, no key, no rate limit described in the README, and no vendor relationship. The honest comparison is that IPsum is a data file with a scoring convention, and the alternatives are products. Choose based on whether you need context or just a filter.

Maintenance cost, refresh cadence and the Unlicense

The upstream side is maintained: the last push to the repository was on 2026-09-28, and the README describes the aggregation running every 24 hours. That tells you the feed is current, not that your deployment is. The refresh loop is yours to build. The README shows the fetch command and the ipset population command but documents no timer, no cron entry and no systemd unit, so the interval between updates is whatever you write. The same gap applies to rollback: there is no documented way to remove the ipsum set and restore the previous ruleset, which matters because the snippet inserts a DROP rule at the head of INPUT. If you are fetching from raw.githubusercontent.com in a loop, you are also depending on that host being reachable and on the file's line format staying stable; the README does not describe a format version or a change policy. On licensing, the repository is under the Unlicense, which the README links to unlicense.org. That is a public-domain-style dedication rather than a permissive licence with attribution conditions, which means redistributing or embedding the feed does not carry the notice requirements that, for example, an MIT-licensed dataset would. This is a description of what the licence identifier says, not legal advice; if you are repackaging the data commercially, read the licence text and your own counsel's view of it.

Editorial conclusion

Adopt IPsum if you already run iptables or ipset and want a scored, plain-text blocklist you can filter by occurrence count before it touches your ruleset. Do not adopt it if you need per-IP context, attribution or a hosted lookup API; the repository is a file, not a service, and the README documents no rollback path. Before switching it on, verify the current format of ipsum.txt, confirm that the levels/ directory matches the threshold you intend to use, and stage the iptables rule so you can remove the ipset set without locking yourself out.

Frequently asked questions

What does IPsum do?

It retrieves and parses more than 30 publicly available lists of suspicious or malicious IP addresses every 24 hours and publishes the merged result into this repository. Each line carries an IP address plus a count of how many source lists contained it.

How do I use IPsum to block bad IPs?

The README's example pipes ipsum.txt through grep and cut to keep only addresses appearing on at least three lists, then adds them to an ipset set named ipsum and inserts an iptables INPUT rule that drops matching sources. Nothing is installed on the host beyond iptables and ipset.

What does the number next to each IP in IPsum mean?

It is the occurrence count: how many of the source lists that address appears on. The README states that higher counts generally mean higher confidence and fewer false positives when blocking inbound traffic, and the file is sorted from highest count to lowest.

Does IPsum provide pre-filtered lists?

Yes. The levels/ directory holds raw IP lists bucketed by number of blacklist occurrences, and the README gives levels/3.txt as the example, containing addresses found on three or more blacklists.

Is IPsum a hosted service or an API?

Neither. The README describes a daily aggregation job that pushes the result to this repository, and its examples fetch the file over HTTPS from raw.githubusercontent.com. There is no documented API, account or key.

Official sources

  1. Issues
  2. License: Unlicense
  3. Project website
  4. README
  5. stamparm/ipsum on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stamparm-ipsum.svg)](https://hysenlabs.com/projects/stamparm-ipsum)