Open-source project
puma/puma avatar
puma/puma

Puma: The Ruby Web Server That Trades Threads for Throughput

A Ruby/Rack web server built for parallelism

7,917 stars1,502 forksRubyBSD-3-Clause

At a glance

What is it?
Puma is a multi-threaded, multi-process HTTP 1.1 server for Rack applications. It is the default server for Rails, and its cluster mode is the part most teams get wrong.
Who is it for?
Adopt Puma if you run a Rack application on MRI, JRuby or TruffleRuby and want a server that ships with Rails and needs no reverse proxy for SSL or rolling restarts. Do not adopt it expecting parallel Ruby execution on MRI: the GVL still serializes CPU-bound work, and a thread count raised to hide that will cause contention instead.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Ruby, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Puma solves for Rack applications on Ruby

A Rack application is a Ruby object that takes an environment hash and returns a status, headers and body. Something has to accept the socket, parse HTTP, and call that object. Puma is that something. The README describes it as a "simple, fast, multi-threaded, and highly parallel HTTP 1.1 server for Ruby/Rack applications," and it is the default server for Rails, included in the generated Gemfile.

The design targets two costs at once. Threads reduce memory per concurrent request compared with a process-per-request model. Processes, in cluster mode, reduce memory across the whole application through copy-on-write. The audience is anyone deploying a Rack app who cares about requests per second per megabyte of RAM: Rails and Sinatra shops, and teams running on JRuby or TruffleRuby where threads execute Ruby in parallel.

On MRI the picture is narrower than the marketing implies. A Global VM Lock allows only one thread to run Ruby code at a time. Puma still helps when the application spends its time waiting on external HTTP calls, because that waiting happens in parallel. It does not help a CPU-bound endpoint, and no thread setting changes that.

Threads, workers and the GVL: how the machinery fits together

Puma maintains a thread pool and scales the number of threads between a minimum and a maximum based on traffic. The default is 0:16, and on MRI it is 0:5. The README warns against setting the maximum very high, since that can exhaust system resources or cause contention for the Global VM Lock.

Cluster mode forks workers from a master process, and each child keeps its own thread pool. The thread setting is per worker, so -w 2 -t 16:16 produces 32 threads across two processes. Preloading loads application code before forking, which lets the operating system share pages between master and workers through copy-on-write. The README states that when the worker count is greater than 1 and --prune-bundler has not been specified, preloading is enabled by default.

The configuration DSL exposes before_fork, before_worker_boot and after_worker_shutdown hooks for code that must run at specific points in the fork lifecycle. The README recommends pairing these with preload_app!, because otherwise constants loaded by the application will not be available inside the hooks. That is a real ordering constraint, not a style preference.

One number is easy to miss: Puma creates its own threads for internal purposes such as handling slow clients. Even with -t 1:1, the README says to expect around 7 threads created in the application.

Installing Puma and running a first Rack app

The README gives a two-line quick start. Install the gem, then run the executable with no arguments:

bash
$ gem install puma
$ puma

Without arguments, Puma looks in the working directory for a rackup file named config.ru. If one is present, the server boots against it. SSL support is compiled in only when OpenSSL development files are installed on the system; without them, Puma still installs and runs, but it will not allow SSL connections.

Inside a Rails application you do not need the gem install step, since Puma is already in the generated Gemfile. The README notes that many configuration options and Puma features are not available through rails server, and recommends the executable instead:

bash
$ bundle exec puma

For Sinatra, the command-line shortcut is ruby app.rb -s Puma, but configuring Puma through a config file requires the puma executable and a rackup file:

ruby
# config.ru
require './app'
run Sinatra::Application

With that file in place, bundle exec puma starts the application. To set the thread pool explicitly, pass the -t flag with minimum and maximum, as in puma -t 8:32. To add cluster mode, combine it with -w, as in puma -t 8:32 -w 3, or set the WEB_CONCURRENCY environment variable. A config file can set workers :auto to match the worker count to available processors, which requires the concurrent-ruby gem. For a first run, start with the defaults and change one value at a time; the README points to the test suite's config directory and to docs/deployment.md for worked examples and for the tradeoffs behind thread and process counts.

Where Puma stops being the right tool

The GVL is the first boundary. An endpoint that spends its time in Ruby computation will not go faster because you added threads; on MRI it will go slower once contention sets in. The README is explicit that truly parallel Ruby implementations, TruffleRuby and JRuby, do not have this limitation, which is another way of saying MRI does.

Preloading and phased restart are mutually exclusive. Phased restart kills and restarts workers one at a time, while preloading copies master code into workers, and the README states plainly that the two cannot be used together. A team that wants preloading for memory savings and phased restart for zero-downtime deploys has to pick one.

Cluster mode also raises the cost of application code that is not fork-safe. The hook documentation exists because connections, background threads and library state created before forking do not survive cleanly into a worker. This is not a Puma defect, but it is work that a single-process server would not ask of you.

Finally, Puma is an HTTP 1.1 server. It is not a reverse proxy, a load balancer or a static file CDN. The README describes standalone deployment with SSL, rolling restarts and a request bufferer as a supported mode, but that is a statement about not needing a proxy in front, not a claim to replace the rest of the edge tier.

Puma against Unicorn and Passenger

Unicorn is the closest comparison and the cleanest one, because Puma's cluster mode is a deliberate answer to it. Unicorn forks workers and serves one request per worker process at a time. Puma keeps the fork model and adds a thread pool inside each worker, so a single process can handle several requests concurrently, mostly while they wait on IO. That is the whole difference in approach: Unicorn buys isolation and predictable memory at the cost of concurrency per process, while Puma buys concurrency per process and asks you to reason about thread safety in your application and its gems.

Passenger takes a different route again. It is a multi-language application server that manages process spawning, and it can sit in front of Puma rather than replace it. Choosing between them is less about raw serving speed and more about how much of the deploy pipeline you want the server to own.

The honest summary is that Puma's advantage on MRI is IO concurrency and memory, not CPU parallelism. If your workload is CPU-bound Ruby, the thread pool is not the lever, and the README's own pointer to JRuby and TruffleRuby is the more useful signal.

Maintenance, releases and licence cost

The repository is not archived, and the last push was on 2026-09-16. Recent releases listed for the project are v7.2.1 and v8.0.2, both dated 2026-05-27, and v8.0.1 dated 2026-04-26. Two maintained release lines exist in that list, which matters if you pin a major version: the 7.x and 8.x branches are both receiving releases rather than one being frozen at the moment 8.0 shipped.

Upgrade cost is concentrated in configuration rather than in the HTTP layer. The README points to puma -h, Puma::DSL and lib/puma/dsl.rb as the authoritative option list, and notes that workers :auto has gotchas documented in that file. A major-version bump is the moment to re-read the DSL rather than assume your puma.rb still maps cleanly onto current options. The PUMA_LOG_CONFIG environment variable prints the loaded configuration as part of boot, which is the fastest way to confirm what a config file actually resolved to after an upgrade.

Puma is licensed under BSD-3-Clause, a permissive licence that allows use in closed-source products and imposes attribution and disclaimer conditions. This is a description of the licence text, not legal advice; if you redistribute Puma or a modified version, read LICENSE in the repository and get your own counsel.

Editorial conclusion

Adopt Puma if you run a Rack application on MRI, JRuby or TruffleRuby and want a server that ships with Rails and needs no reverse proxy for SSL or rolling restarts. Do not adopt it expecting parallel Ruby execution on MRI: the GVL still serializes CPU-bound work, and a thread count raised to hide that will cause contention instead. Do not use it with preload_app! if your deploy depends on phased restart, because the README states the two cannot be combined. Before committing, verify three things against your own application: whether OpenSSL development files are present at install time, whether concurrent-ruby is in your bundle if you plan to set workers :auto, and whether your config file uses phased restart.

Frequently asked questions

Does Puma need a reverse proxy in front of it?

Not necessarily. The README describes Puma as standalone, with SSL support, zero-downtime rolling restarts and a built-in request bufferer, which it presents as enough to deploy without any reverse proxy.

Can Puma run more than one thread per worker process?

Yes. Cluster mode forks workers from a master process and each child process still has its own thread pool, so -w 2 -t 16:16 spawns 32 threads in total, 16 in each worker.

Why does Puma create more threads than the -t setting asks for?

Puma creates additional threads for internal purposes such as handling slow clients. The README states that even with -t 1:1 you should expect around 7 threads created in your application.

Can preloading be combined with phased restart in Puma?

No. The README states that preloading cannot be used with phased restart, because phased restart kills and restarts workers one by one while preloading copies the master's code into the workers.

Does Puma serve SSL connections without extra setup?

Only if OpenSSL development files are installed on the system when Puma is installed or compiled. Without them, Puma still installs and runs, but it will not allow SSL connections.

How do I see which configuration Puma actually loaded?

Set the PUMA_LOG_CONFIG environment variable to a value. The README states that the loaded configuration will then be printed as part of the boot process, which is intended for debugging.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. puma/puma on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/puma-puma.svg)](https://hysenlabs.com/projects/puma-puma)