Open-source project
cure53/HTTPLeaks avatar
cure53/HTTPLeaks

HTTPLeaks: one HTML file enumerating everything that phones home

HTTPLeaks - All possible ways, a website can leak HTTP requests

2,126 stars206 forksHTMLBSD-2-Clause

At a glance

What is it?
HTTPLeaks is a single hand-maintained HTML file, plus a README and a licence, that contains every construct known to make a browser, mail client or server-side DOM send a request it should not. Every entry points at a domain that deliberately does not resolve, entries are labelled live, legacy or probe, and the CSS section shadows itself so a naive run under-reports.
Who is it for?
Use HTTPLeaks as a test input if you maintain an HTML sanitizer, a web mail client, a reader mode, a web proxy or a Content Security Policy, because it is the only public corpus that is maintained as a living document rather than generated from a specification, and its legacy entries cover the mail-client rendering paths no browser test will find.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 21 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The domain does not resolve, and that is the design

The first thing to understand about this repository is that its entire contents are an attempt to make software send requests, and that the requests all go nowhere on purpose.

Every entry in the file points at a URL of the form https://leaking.via/<element>-<attribute>. For a construct that fires, the browser resolves that host, and if nothing answers, the request fails. The benefit is that the failure is silent to the user and completely visible to you, because you control the DNS or the listener.

The README is explicit about why. The domain is not real; the file is a corpus, not a beacon.

That sentence is the most important line in the project. Consider the alternative. If the entries pointed at a domain the project controlled, then every person who opened the file would send a request to the project author, the file would function as a web bug with a known reader list, and anyone who found a new leak construct would be reporting it by email to whoever owned the domain. Every security researcher's least favourite artefact is one that phones home when demonstrated.

Pointing at a non-resolving domain in a real but unowned TLD gives you the same diagnostic property with none of that. Requests still leave the browser, still hit your listener, still show up in a packet capture or a wildcard DNS log, and reach nobody. A TLD was chosen rather than an example.com subdomain precisely because browsers and mail clients treat the schemes and hosts differently, and a plausible-looking host exercises the same code paths as a real one.

The naming scheme does double duty for the same reason. Because the path is derived from the construct, the URL alone tells you what fired. https://leaking.via/img-srcset means a srcset attribute caused it; leaking.via/style-import means a CSS @import did. There is no need to correlate a request against a list, no instrumentation in the page, and no mapping table to maintain. A packet capture and the file are enough.

The contributing rules make uniqueness a requirement, so a path identifies exactly one entry. The same convention applies to CSS, with <context>-<property> in place of <element>-<attribute>. Two projects could merge these files and a single capture would still be unambiguous.

DOMPurify's stated non-goal is this file's reason to exist

The Related section points at two projects, and the first one reframes everything.

DOMPurify, by the same author, sanitizes HTML, SVG and MathML. Its demos include a hooks-proxy-demo showing how to route or drop resource-loading attributes with a hook. And its threat model explains why stopping HTTP leaks is explicitly a non-goal of a sanitizer.

That is the key sentence, and it is not a criticism. A sanitizer's job is to stop script execution. If a sanitizer also stripped every src, href, srcset, ping and image-set from allowed markup, it would break most of the web, because those attributes are how images and links work. So the author drew the line at execution and left resource loading to the policy layer, which is Content Security Policy. That is a defensible division of labour, and it is a common one.

But it means there is a category of attack the sanitizer does not address, and HTTPLeaks is the inventory of that category. A page with no script in it can still tell a server that it was opened, which IP opened it, and roughly when. In the mail-client case that is a read receipt by another name, and the README's framing is exactly that: think of the body of an HTML mail, where a leak tells someone that you just opened it. Not always bad, almost never good.

The other listed use cases follow the same logic. A proxy that rewrites src and href but not srcset, ping, @import or image-set has no anonymity to offer, which is a claim about a specific implementation gap rather than a general one. And for CSP testing, the question is which of these your policy actually blocks, and which requests leave the browser process without ever appearing in devtools.

That last clause is where DOMFortify comes in. It retrofits Trusted Types sanitization onto legacy pages, and the README's note is that this file is a good input for testing what a sanitizer-backed policy still lets through. So the two projects compose: DOMPurify removes execution, DOMFortify adds a types layer, and HTTPLeaks measures what both of them still let out of the door.

The relationship is worth stating plainly because it explains the file's shape. This is not a general web security checklist. It is the boundary case list for one specific architectural decision, maintained by the person who made that decision.

The CSS block shadows itself, so a naive run under-reports

The How to test with it section has four numbered steps, and the fourth one is a defect in the corpus itself that anyone measuring results will hit.

Later CSS declarations of the same property win, so where the file lists several url() variants on one selector you will only see the last one fire. Split them up if you need each individually.

CSS cascade works that way, and it means the file contains entries that can never be observed in a single run. If the stylesheet block has one selector with, say, a background-image set three times to demonstrate three URL spellings, the first two are dead on arrival: the browser applies the third and issues one request. The other two paths are still in the file, still counted by anyone reading it, and never seen by anyone testing it.

That is a design tension rather than an oversight, and the file is a document, not a test runner. A single HTML file that must be readable top to bottom in document order cannot both group related constructs and make them independently observable. Grouping is what makes the corpus useful to a human; independence is what makes it useful to a measurement.

The stated workaround is to split the entries into separate selectors or separate files, which a tester can do by hand. The cost is that the result stops being the shipped artefact. Once you have edited the file, you are no longer measuring what other people will download, and your numbers describe your fork.

The section header is worth noting too. It is described as CSS, and broken into stylesheet, inline and exotic. The exotic bucket is where image-set() and similar constructs live, and those are exactly the ones a naive rewriting proxy misses, which is the use case the README names. So the part of the corpus that under-reports on a single run overlaps with the part that matters most.

The other three testing steps are the part worth adopting wholesale, and they are the reason this project is more useful than a static checklist. Serve the file over HTTPS from an origin that is not leaking.via, because opening it from disk changes the rules for several loaders including mixed content, the file: origin and CSP meta handling. Watch the network at the DNS or wildcard-listener level, not only in devtools, because several requests are made by the browser process such as OpenSearch, after the page is idle such as compression dictionaries, or by a prefetch and prerender pipeline. And run it more than once, with scripting on and off, because the noscript block only fires with scripting off, in each engine you care about, and in the actual client you are protecting, because browser results do not transfer to mail clients and vice versa.

Legacy entries are kept by policy, because a mail client is not a browser

The file contains entries that cannot fire in any current browser, and keeping them is an explicit rule rather than an oversight.

The README groups entries into three kinds. Live leaks in current engines, described as the bulk of the file. Legacy leaks that only fire in engines that are gone from the desktop, in IE, old Firefox and the Flash and Java plugin era, but which live on in mail clients and embedded webviews. And probes that are expected not to fire and are there as regression inputs, with comments saying expected: no request.

The justification for keeping the legacy block is one sentence: Outlook still renders VML and [if mso] blocks. These stay in on purpose: the corpus is more useful complete.

And the contributing guidelines make it a hard rule: do not remove legacy entries, annotate them instead.

That is the correct call and it is worth understanding why. A modern browser is a moving target in one direction, and the constructs it dropped are, for the most part, gone. A mail client is a different situation. Outlook's rendering engine is not a browser, its support for VML and conditional comments is a compatibility feature that will not be removed while anyone sends HTML mail, and an attacker who knows that can construct a payload that a browser-based sanitizer will happily pass and that Outlook will happily render. The corpus sections confirm the scope: MSIE data islands, VML, and a separate block of URL spellings that sanitizers tend to misclassify.

So the legacy entries are not nostalgia. They are the only public documentation of a threat surface that exists today in the most widely deployed HTML renderer in the world, which is an email client that most people never open in a browser at all. And the testing instructions reflect that: run it in the actual client you are protecting, Outlook or a webmail sandbox or your proxy, because browser results do not transfer to mail clients and vice versa.

The `is_executable`-style helper is worth flagging too, though it is a small thing. A few entries need user interaction, specifically ping, formaction, MathML href and longdesc, and the visible text says click me or hover me where that is the case. A request that only fires after a click is a different risk from one that fires on open, so the corpus distinguishes them in the rendered text rather than only in a comment. That is a small design decision that makes the file honest about what it found when someone reports a result from it.

There are no releases, so you have to pin a commit yourself

The repository has no GitHub releases, which is stated as a fact in the project metadata rather than as a gap, and for a security corpus it is the most consequential thing in the file listing.

The whole repository is four entries: .gitignore, LICENSE, README.md and leak.html. There is no package.json, no build script, no CI configuration, no changelog and no version file. The primary language is recorded as HTML, which is what you would expect when the deliverable is one HTML file.

That is not a criticism of the packaging, because there is nothing to package. A browser loads a file. The raw URL is the distribution mechanism, and it points at the main branch, so downloading the file today gives you the current state of the corpus rather than a tagged snapshot.

The problem is for anyone who wants to use it as a test input. Suppose you run HTTPLeaks against your sanitizer every night in CI. A new entry lands in leak.html, your test starts failing, and nothing in your artefact tells you which version of the corpus produced it. You cannot diff against a tag, because there are no tags, and you cannot cite a version in a bug report. The only durable identifier is the commit hash, and the README does not tell you to record it.

This is a real gap and it is the kind of thing that is easy to fix and never gets fixed, because the project's own contributors are working in the file and not in a release process. The comparison is worth drawing. The three Related projects, DOMPurify, DOMFortify and the hooks demo, are all versioned npm packages with release histories, because they are code. HTTPLeaks is a document, and documents are conventionally served from a branch.

The second consequence of no build step is the contribution mechanics. A new leak is a pull request against one HTML file with section comments used as anchors. Merging two PRs that both add entries to the same section is a merge conflict in a large text file, and the reviewer's job is to check that the new entry uses the naming scheme, sits under the matching %Section, has a comment saying where it fires and whether interaction is needed, and is either a live leak, a legacy entry or a probe marked as expected: no request. That is a workable process for a small number of concurrent PRs. It is not a scalable one, and it is the constraint that eventually pushes a project like this towards either generated data or a split file.

The README anticipates that, actually. Ideas for other presentations of this data, JSON, per-engine tables, a scripted runner, are also welcome; the single file is the source of truth. So the maintainer has the right answer in mind and has not needed it yet.

Probes that must not fire are what make this a corpus and not a demo

The third category of entry is the one that turns this from a list into something you can regression-test against, and it is the least obvious.

Probes are expected not to fire and are there as regression inputs, with comments saying expected: no request. The example given is the attr() URL restriction in CSS Values 5.

A live entry is a construct that currently causes a request. Its presence in the file is a fact about the present. A probe is a construct that ought to cause no request, and its presence is a claim about the future: if the engine implements the restriction, the probe stays quiet, and if the engine regresses or the restriction is removed, the probe starts firing.

That is the difference between a demonstration page and a test corpus, and it is the difference that makes the file useful in CI. A demonstration page tells you what is broken today. A corpus with negative cases tells you what stayed fixed.

The naming convention makes the two indistinguishable on the wire, and that is deliberate. A request to leaking.via is a failure regardless of whether it came from a live entry or a probe. There is no category in the URL, so an automated checker does not need to know which is which, and a human reading a capture does not have to consult a table. What distinguishes them is a comment in the file, which is the right place for the distinction because the distinction is about expectation rather than about behaviour.

The sections give a sense of the coverage, and the order is roughly document order rather than alphabetical. head metadata, link relations, images, forms, media, object and embed, script and speculation rules, frames, Declarative Shadow DOM, noscript, model, CSS in three flavours, SVG, XSLT, MSIE data islands, VML, MathML, and finally the block of URL spellings that sanitizers tend to misclassify.

Two of those entries are worth pausing on for anyone building a filter. Speculation rules are a browser feature where a link is speculatively prefetched or prerendered, which is a request triggered by markup that looks inert. Declarative Shadow DOM attaches a shadow tree from parsed HTML rather than script, which is relevant to any policy that reasoned about script-created shadow roots. Both are the kind of thing that arrives in a browser years after your sanitizer was written, and both are the reason the file has to be maintained by hand.

XSLT and MathML are in the same category. MathML href is one of the interaction-requiring entries, and an XML transformation instruction set in a document is a considerably larger surface than a stylesheet.

The acknowledgements list twenty-one named people, including several who are well known for browser internals work. That is a signal about where the entries came from: this file is the accumulated output of people who read engine source, which is why it exists at all, in the form the README gives: nobody really knows anymore which elements and attributes can request external resources, which is why this project exists, to keep track.

The testing protocol is the deliverable, and it is the part to copy

A single HTML file with 1500 constructs in it is worth a certain amount. The four-step testing protocol in the README is worth more, because it is the part that generalises.

Step one is to serve the file over HTTPS from an origin that is not leaking.via, and the reason given is that opening it from disk changes the rules for several loaders, specifically mixed content, the file: origin and CSP meta handling. That is three separate mechanisms broken by one shortcut. A file: page has an opaque origin, so anything comparing origins behaves differently. A file: page is not a secure context in the way an HTTPS page is, so mixed-content rules and the availability of certain APIs change. And a CSP delivered in a meta element is not applied in a file: context. Each of those would change which entries fire, and a tester who opened the file locally and reported a clean result would be reporting a measurement of a configuration nobody deploys.

Step two is to watch at the DNS or wildcard-listener level rather than only in devtools, and the reasons are enumerated: several requests are made by the browser process, OpenSearch being the example; after the page is idle, compression dictionaries being the example; or by a prefetch and prerender pipeline. And the crucial qualification: they do not all show in the Network panel. So the default tool a web developer reaches for is structurally incapable of seeing part of what this file is designed to detect. That single sentence is the most practically important line in the project, and it generalises well beyond HTTP leaks, because it is a statement about the difference between what a browser does and what the devtools Network panel is instrumented to report.

Step three is to run it more than once: with scripting on and off, because the noscript block only fires with scripting off; in each engine you care about; and in the actual client you are protecting. The last clause is the one that catches people: browser results do not transfer to mail clients and vice versa. Given that a third of the file is legacy constructs that only fire in mail clients, a browser-only test is measuring less than half the corpus.

Step four is the CSS shadowing caveat discussed earlier, and it is the one piece of advice that requires you to modify the artefact.

Put together, the protocol is a statement about what a security test has to be to be worth running. Serve it the way it will be deployed. Observe it at a level below the user interface. Test it in the thing you actually care about, not in the thing that is convenient. And know which of your assertions the test itself cannot support.

That protocol transfers to any leak-testing problem, which is why a repository with four files and no releases is worth reading carefully.

Editorial conclusion

Use HTTPLeaks as a test input if you maintain an HTML sanitizer, a web mail client, a reader mode, a web proxy or a Content Security Policy, because it is the only public corpus that is maintained as a living document rather than generated from a specification, and its legacy entries cover the mail-client rendering paths no browser test will find. Do not treat it as a scanner: it is a file you load, not something you point at a target, and it will not tell you anything about your own pages. Do not wire it into an automated regression suite without solving the release problem, because the repository has no GitHub releases at all, so there is no version to pin and you have to record the commit hash in your own test metadata. Verify four things. Serve the file over HTTPS from an origin other than leaking.via rather than opening it from disk, because the file: origin changes mixed-content and CSP meta handling. Watch the network at the DNS or wildcard-listener level rather than in devtools, because OpenSearch, compression dictionaries and prefetch traffic do not all appear in the Network panel. Split the CSS entries apart before measuring, because later declarations of the same property win and you will otherwise under-count. And read DOMPurify's threat model, which states that stopping HTTP leaks is a deliberate non-goal, so you know this corpus is testing the half of the problem your sanitizer does not cover. The deciding fact is that a non-resolving domain makes the file safe to open and a real leak surface impossible to weaponise, which is why it is the right kind of artefact to keep in your test fixtures.

Frequently asked questions

What is HTTPLeaks and what is it for?

It is a single HTML file that enumerates every way a document can cause a browser, mail client, proxy or server-side DOM to send an HTTP request it should not. The listed uses are testing HTML sanitizers and filters, testing web mailers and reader modes, checking what a web proxy or anonymizer fails to rewrite, and checking which of these a Content Security Policy actually blocks.

How do I test with HTTPLeaks safely?

Serve leak.html over HTTPS from an origin that is not leaking.via, rather than opening it from disk, because the file: origin changes mixed-content, origin and CSP meta handling. Then watch the network at the DNS or wildcard-listener level rather than only in devtools, since OpenSearch, compression dictionary and prefetch requests do not all appear in the Network panel. Run it with scripting on and off and in the client you actually protect.

Do the requests in HTTPLeaks go anywhere?

No. Every entry points at https://leaking.via/<element>-<attribute>, and the README states that the domain is not real and the file is a corpus, not a beacon. The naming also means the path alone identifies which construct fired, so a packet capture and the file are enough to interpret a result without any instrumentation in the page.

Why does HTTPLeaks keep entries for old browsers and IE?

Because those constructs still fire where the desktop no longer renders them. Outlook still renders VML and [if mso] blocks, and embedded webviews and mail clients keep the legacy path alive. The contributing guidelines say not to remove legacy entries and to annotate them instead, on the grounds that the corpus is more useful complete and that browser results do not transfer to mail clients.

Is HTTPLeaks related to DOMPurify?

Yes, same author. DOMPurify sanitizes HTML, SVG and MathML, and its threat model states that stopping HTTP leaks is explicitly a non-goal of a sanitizer, because stripping resource-loading attributes from allowed markup would break the web. HTTPLeaks covers the half DOMPurify deliberately leaves to Content Security Policy, and the README also points at DOMFortify for Trusted Types coverage on legacy pages.

Does HTTPLeaks have versioned releases?

No. The repository has no GitHub releases, and the raw file is served from the main branch. If you use it as a CI input you have to record the commit hash yourself, because nothing in the artefact identifies which revision produced a given result.

Official sources

  1. cure53/HTTPLeaks on GitHub
  2. Issues
  3. License: BSD-2-Clause
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cure53-httpleaks.svg)](https://hysenlabs.com/projects/cure53-httpleaks)