# HTML-Pipeline: a Ruby chain for turning user text into safe HTML

> A small framework of chainable filters where text filters run first, a convert filter turns Markdown into HTML, sanitization happens automatically, and node filters do the last pass. Started at GitHub, now standalone.

**gjtorikian/html-pipeline** — HTML processing filters and utilities

- Repository: https://github.com/gjtorikian/html-pipeline
- Stars: 2,329 · Forks: 385
- Language: Ruby
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/gjtorikian-html-pipeline

## Four filter kinds and a fixed order between them

The whole library is one idea: user content goes in as a UTF-8 string, a sequence of filters transforms it, and HTML comes out. The README describes it as a small framework for defining CSS-based content filters and applying them to user provided content, and the filters fall into four kinds.

`TextFilter`s operate on the plain UTF-8 string before any HTML exists. `ConvertFilter` turns text into HTML, with the shipped implementation being `MarkdownFilter`. A `SanitizationFilter` removes dangerous or unwanted elements and attributes. `NodeFilter`s operate on the resulting UTF-8 HTML document.

The order is fixed and this matters more than the individual filters. Text filters run first, then the convert filter, then sanitization, then the node filters. Each filter hands its output to the next filter's input, so a text filter cannot see HTML and a node filter cannot see the original Markdown. The README also notes you can assemble the sequences into a single pipeline or call each filter individually, which is the escape hatch when you want one stage without the rest.

The classes named in the text and convert stages are `ImageFilter`, which turns an image `url` into an `<img>` tag, `PlainTextInputFilter`, which HTML escapes text and wraps the result in a `<div>`, and `MarkdownFilter`, which builds HTML through Commonmarker. Every filter defines a method called `call`, with `@text`, `@config` and `@result` available, and changes to those instance variables pass on downstream.

## Installing the gem and running a single filter

It is a RubyGem, so the normal route is a line in your Gemfile followed by a bundle:

```ruby
gem 'html-pipeline'
```

```sh
$ bundle
```

Or install it directly:

```sh
$ gem install html-pipeline
```

If you only want the Markdown conversion and nothing else, the README shows that you can instantiate one filter and call it, skipping the pipeline object entirely:

```ruby
filter = HTMLPipeline::ConvertFilter::MarkdownFilter.new
filter.call(text)
```

That single filter path is worth noting because it is also the path that skips sanitization. A filter called on its own does not pass through the automatic sanitization stage described below, so if the input is user supplied you want the pipeline form rather than the direct call. The README's own comment on that example says the filter object is reusable, and that the pipeline is the recommended way to call it repeatedly.

One detail to watch when you copy from the README: the pipeline example requires `html_pipeline` with an underscore, while the gem name itself is hyphenated. If you are following the examples literally, that require line is the place to check first when nothing loads.

## Writing your own filter as a subclass

Custom filters are the point of the framework, and the README's worked example builds one that changes every instance of `Hey` to `Hello`, then runs it as a text filter alongside the Markdown conversion and mention linking:

```ruby
require 'html_pipeline'

class HelloJohnnyFilter < HTMLPipelineFilter
  def call
    text.gsub("Hey", "Hello")
  end
end
```

The contract is small: subclass the base filter class and define `call`. Inside it you read `@text`, `@config` and `@result`, and whatever you leave in those variables is what the next filter receives.

The example pipeline that follows is worth reading as the shape of a real configuration rather than as copy-paste code, because it shows the four stages together with a sanitization config passed in:

```ruby
pipeline = HTMLPipeline.new(
  text_filters: [HelloJohnnyFilter.new]
  convert_filter: HTMLPipeline::ConvertFilter::MarkdownFilter.new,
  sanitization_config: HTMLPipeline::SanitizationFilter::DEFAULT_CONFIG,
  node_filters: [HTMLPipeline::NodeFilter::MentionFilter.new]
)
```

There is a missing comma between the text filters and the convert filter in that snippet as printed, and the surrounding text notes that the sanitization config line is not needed because sanitization occurs by default. Both are reasons to read the documentation for the current version rather than paste from a blog post, and the presence of an `UPGRADING.md` at the repository root suggests the API has moved between versions.

## Sanitization sits between conversion and the node filters

This is the part of the design that separates the library from a Markdown converter. The README states that HTML is automatically sanitized after the `ConvertFilter` runs and before the `NodeFilter`s are processed, and gives the reason plainly: to prevent malicious or unexpected input from entering the pipeline.

The positioning is deliberate. Converting Markdown produces HTML, so sanitizing before conversion would be sanitizing the wrong thing, and sanitizing only at the end would leave a window in which node filters operate on unsanitized markup. Placing it in the middle means your node filters, which are the ones most likely to inject markup of their own, run against already-filtered HTML.

Sanitization is configured by hash, and the README defers the details to the Selma documentation, warning that users must configure it correctly to get correct behaviour in combination with handlers which manipulate HTML. A default config ships as `HTMLPipeline::SanitizationFilter::DEFAULT_CONFIG`, and the sample custom allowlist in the README is deliberately minimal:

```ruby
ALLOWLIST = {
  elements: ["p", "pre", "code"]
}
```

An allowlist of three elements is a good illustration of the default posture. If a filter needs to emit an `img` or an anchor and sanitization removes it, the fix belongs in this configuration rather than in a bypass, which is what the release history suggests the maintainers have been enforcing.

## Context and result carry state between filters

Some filters accept optional `context` and `result` hashes, and the README uses them to pass arguments and metadata between the filters in one pipeline. The documented case is disabling footnotes in the Markdown filter, which is done through the context hash rather than through a filter-specific setter:

```ruby
context = { markdown: { extensions: { footnotes: false } } }
filter = HTMLPipeline::ConvertFilter::MarkdownFilter.new(context: context)
filter.call("Hi **world**!")
```

The alternative is to build the pipeline first and pass the context at call time, which is the pattern to prefer when the same pipeline serves requests with different options:

```ruby
pipeline.call(user_supplied_text, context: { markdown: { extensions: { footnotes: false } } })
```

The README's further examples show a context hash shared across named pipelines, carrying an `asset_root` and a `base_url`, which is how you tell a mention filter where links should point. Two things stand out in those examples. First, that different parts of an application can have different pipelines, so a web comment pipeline and an email pipeline can share filters while differing in their configuration. Second, that the README notes pipelines are not limited to the web, and gives an email example using `PlainTextInputFilter` with an `ImageFilter`.

The `base_url` key in that shared context is the one that interacts with the mention filters, and the release history shows it being adjusted rather than settled: v3.2.2 added support for an `@` prefix on the `MentionFilter` base_url and fixed a possible bug when using `MentionFilter` together with `TeamMentionFilter`.

## Standalone since GitHub stopped using it

The second line of the README is the most consequential sentence in the project. It was started at GitHub, GitHub no longer uses it, and the gem must be considered standalone and independent from GitHub. Any assumption that this library tracks what github.com does internally is wrong, and any expectation that GitHub will maintain it is also wrong.

What remains is a conventional gem with a `Gemfile`, an `html-pipeline.gemspec`, a `Rakefile`, a `CHANGELOG.md`, an `UPGRADING.md`, a `.rubocop.yml`, a `.ruby-version`, and separate `lib/` and `test/` directories. It is MIT licensed, and the last push to the repository was on 2026-06-02.

The release cadence is slow but not abandoned. v3.2.2 shipped on 2024-12-15 with the mention filter changes, a bugfix making sanitization-only filters work, and a dependency bump that moved commonmarker from the 1.1 series to 2.0. v3.2.3 on 2025-04-24 added support for just-text pipelines. v3.2.4 on 2026-01-06 added Ruby 4.0 support and allowed completely nil sanitization, along with two dependabot bumps of the checkout action.

Read as a sequence, those releases describe a library in low-churn maintenance that fixes real edge cases as they surface. Nil sanitization being added in v3.2.4 and just-text pipelines in v3.2.3 both point the same way: making the framework flexible about which stages a caller actually needs, rather than adding filters.

## Conclusion

HTML-Pipeline earns its place when you render user-supplied text in more than one place, because the same sanitized chain can be pointed at a web comment and an email body without duplicating the escaping logic. The design decision to sanitize between conversion and node filters is the reason to adopt it rather than call a Markdown library directly, and it is also the thing you must configure deliberately if your handlers touch the HTML. Install it as a gem, keep your own filters to the text stage where possible, read `UPGRADING.md` before jumping a major version, and check the commonmarker version your bundle resolved, since 3.2.2 moved it to the 2.x line.

## FAQ

### How do I install HTML-Pipeline in a Ruby project?

Add `gem 'html-pipeline'` to your Gemfile and run `bundle`, or install it directly with `gem install html-pipeline`. The README frames installation as something that should typically happen inside a virtual environment, and points at the gem for the actual dependency set.

### Does HTML-Pipeline sanitize user input automatically?

Yes. The README states that HTML is automatically sanitized after the convert filter runs and before the node filters are processed, specifically to prevent malicious or unexpected input from entering the pipeline. The settings come from a hash you configure, with a default config provided as `HTMLPipeline::SanitizationFilter::DEFAULT_CONFIG`.

### How do I write a custom filter?

Subclass the base filter class and define a `call` method, using `@text`, `@config` and `@result` inside it. The README's example subclasses `HTMLPipelineFilter` and rewrites text with a gsub, then passes an instance of it into the `text_filters` array when building an `HTMLPipeline`.

### Is HTML-Pipeline still maintained by GitHub?

No, and the README says so directly: the project was started at GitHub, GitHub no longer uses it, and the gem must be considered standalone and independent from GitHub. It is MIT licensed and receives occasional releases, with v3.2.4 published on 2026-01-06.

## Sources

- [gjtorikian/html-pipeline on GitHub](https://github.com/gjtorikian/html-pipeline)
- [Issues](https://github.com/gjtorikian/html-pipeline/issues)
- [License: MIT](https://github.com/gjtorikian/html-pipeline/blob/main/LICENSE)
- [README](https://github.com/gjtorikian/html-pipeline/blob/main/README.md)
- [Releases](https://github.com/gjtorikian/html-pipeline/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gjtorikian-html-pipeline
