Library / SDK
ankane/production_rails avatar
ankane/production_rails

ankane/production_rails: a production checklist that is mostly a dependency list

Best practices for running Rails in production

2,324 stars138 forksUnknownCC-BY-4.0

At a glance

What is it?
Four markdown files and no code, organised as a checklist of things to add to a Rails app: a logging config, two database timeout blocks, a job adapter line, and a set of named event categories. The concrete numbers are useful, a good half of the recommendations are the author's own gems, and the last push was on 2026-03-31.
Who is it for?
Use production_rails as a review checklist, not as an architecture. It earns its place in two places specifically: the database timeout blocks, which are the only concrete numbers in the guide and are a genuine trap to copy by hand, and the Notable event list, which names failure modes most teams never count.
Can I use it commercially?
Yes, with credit. CC-BY-4.0 allows commercial use as long as you credit the authors and indicate what you changed. It is written for creative content, so check how it applies to any code.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Four markdown files, and the checklist format is the design

The entire repository is Development.md, Scaling.md, README.md and LICENSE.txt. There is no gem, no plugin, no code to install and no configuration to import, and the README says what it is in its first line: best practices for running Rails in production. So the unit of this project is a recommendation, and the format is a list of them grouped by concern: security, errors, logging, audits, migrations, web requests, background jobs, email, caching and performance, monitoring, timeouts, analytics, and new features. Almost every entry follows the same shape. Name a problem in one line, then name a library that solves it, and for the entries where a line of configuration is enough, show the line. The last section points to two companion documents, Development Rails and Scaling Rails, which means the scope of this file is deliberately the production slice and the reading continues elsewhere. What the format buys you is speed. What it costs you is depth, because there is no argument anywhere for why one choice is better than another, and a reader who already knows the answer gets the same list as a reader who does not.

Reading the recommendations as one maintainer's dependency list

The README discloses its own bias in the second paragraph, and it is worth taking the disclosure at face value rather than skimming past it. It says recommendations come from personal experience and from work at Instacart, and that a number of the open source projects linked are ones the author created. Counting what that means in practice: Strong Migrations for catching unsafe migrations at development time, Slowpoke for request timeouts, Notable for tracking notable events, PgHero for Postgres issues, Ahoy and Ahoy Email for analytics and message history, and Field Test for A/B testing features. That is six of the named tools, plus a companion guide on sensitive data and another on secure Rails, all from one person. Other entries point elsewhere, with Sidekiq, Memcached, Dalli, Papertrail, Rollbar, Postmark, CloudFront, Lograge and Rollout named as outside recommendations. The distinction is not that the author's tools are worse. It is that following this guide as written puts a single maintainer on the critical path of your deploys, your migrations, your database monitoring and your feature flags, and that is a concentration decision you should make once, deliberately, rather than by reading a list.

The timeouts section is the only part with numbers, and the units differ

If you take one thing from this guide, take the database timeout configuration, because it is specific, it is easy to get wrong, and the guide links out to a separate document on Ruby timeouts for the reasoning. For Postgres in config/database.yml it gives:

yml
production:
  connect_timeout: 2
  checkout_timeout: 5
  variables:
    statement_timeout: 5000 # ms

For MySQL and MariaDB the shape is different and so is the unit convention:

yml
production:
  connect_timeout: 1
  read_timeout: 1
  write_timeout: 1
  checkout_timeout: 5
  variables:
    max_execution_time: 5000 # ms
    max_statement_time: 5 # sec

Five seconds of statement time is written as 5000 with an ms comment in the Postgres block and as 5 with a sec comment in the MySQL block, and the connect timeout differs too, two seconds against one. None of that is wrong, and all of it is a trap for anyone copying values between engines. A query that runs long is the classic Rails failure mode, because a slow query holds a connection, a held connection blocks a checkout, and a blocked checkout is what your request timeout eventually fires on. Setting these numbers in the database rather than in the application is the point of the section, and the checkout timeout of five seconds in both blocks is the value that actually protects the pool.

Notable events: the failures the guide thinks you should be counting

The monitoring section is split into tracing, uptime, the database, and a category called notable events, and that last category is the most opinionated thing in the repository. The list is: errors, slow requests, jobs and timeouts, 404s, validation failures, CSRF failures, unpermitted parameters, and blocked and throttled requests. Consider what is in that list. Validation failures and unpermitted parameters are not exceptions, they are your application working correctly and telling you something. A 404 can be a typo or a probe. A CSRF failure is close to a security signal by definition. Grouping them with errors says these events belong in one counting system with one alert policy, and it implies you will be looking at rates rather than individual occurrences. That is a real design opinion, not a formatting choice, and it is the kind of thing this checklist is best at: one line that changes how you set up monitoring. The database gets its own attention for the same reason, with PgHero named for Postgres and Active Record Query Logs named for tracking the origin of SQL queries, which answers the question of which code path produced the slow statement.

The logging config reduces volume and captures params, which is a decision

Lograge is recommended to reduce log volume, and Papertrail as the centralised service to put it in. The configuration shown is where the guide stops being a list and starts being opinionated, because it is specific about what to keep and what to drop:

ruby
config.lograge.enabled = true
config.lograge.custom_options = lambda do |event|
  options = event.payload.slice(:request_id, :user_id)
  options[:params] = event.payload[:params].except("controller", "action")
  options
end

Two things are happening. The lambda keeps request_id and user_id, which are added in the companion append_info_to_payload override using request.uuid and current_user.id when there is a current user, so every log line can be tied to a request and an account. And it captures the request params, stripping out controller and action only. So the guide's position on volume, which is to cut it down, coexists with a position on detail, which is to keep the parameters that describe what the user did. That combination means request bodies are shipped to whatever centralised service you pick. The guide does address sensitive data elsewhere, linking a separate page on securing it, so the concern is not ignored, but the snippet as written captures params by default and it is up to you to decide that is right for the data your endpoints accept.

Background jobs: one adapter line and three control levers

The background job section is the shortest and the most immediately actionable. The adapter is one line:

ruby
config.active_job.queue_adapter = :sidekiq

The rest is about what to do when the queue misbehaves, and the guide names ActiveJob::TrafficControl for three things: quickly disable jobs, throttle, and limit concurrency. One line shows the blunt instrument:

ruby
BadJob.disable!

Those three levers are the whole reason to put a framework behind your queue. Disable is the emergency stop for a job class that is misfiring, throttle reduces how much load it adds, and concurrency limiting stops one slow job from consuming every worker slot. The first real limitation of this approach is that the guide names Sidekiq in the adapter line and BadJob in the disable example, which are different backends, so those two snippets are not a pair you can paste together unchanged. The second is that the section says nothing about idempotency, retries or the poison-job case, which are the three questions that actually decide whether your background work is safe to retry. For a checklist that is otherwise complete on the operational side, that omission is notable, and the Scaling Rails document is the place to look for it.

What the guide cannot tell you, and how current it is

Three limits are structural rather than fixable. The guide is Rails-only and assumes you are already running Rails, so it says nothing about choosing a framework, provisioning machines, or deploying. It is a list of additions, which means it has no opinion on architecture and no failure analysis: the entries tell you to use a request timeout, not what happens when the process behind it dies. And it is a document with a version history of zero, since the repository has no GitHub releases at all and no tags, so there is nothing to diff when a recommendation changes. On currency, the last push was on 2026-03-31 and the repository is not archived, but that means nothing has changed for roughly half a year, and several of the named tools move faster than that. The licence is CC-BY-4.0, which is the right choice for a document you might want to translate or republish, and it carries the usual attribution obligation if you do. One small friction point that will confuse a first-time reader: the closing section sends suggestions to the issues of ankane/rails-best-practices, a differently named repository from this one, so the feedback loop for this particular file does not land where you would expect.

Editorial conclusion

Use production_rails as a review checklist, not as an architecture. It earns its place in two places specifically: the database timeout blocks, which are the only concrete numbers in the guide and are a genuine trap to copy by hand, and the Notable event list, which names failure modes most teams never count. Judge the rest of it on the dependency concentration it implies, because Strong Migrations, Slowpoke, Notable, PgHero, Ahoy and Field Test are all one maintainer's projects and the guide says so up front. Two limits to accept before you rely on it. It has never published a release, and its last push was 2026-03-31, so there is no version to pin and nothing has changed in roughly half a year. And it says nothing about deployment, capacity or architecture, so the questions it cannot answer are the ones that decide whether your Rails app survives real traffic.

Frequently asked questions

What is ankane/production_rails and what does it contain?

It is a written guide, not a library, and the repository holds four files: README.md, Development.md, Scaling.md and LICENSE.txt. README.md is the production checklist, with the other two documents covering development and scaling.

What database timeouts does production_rails recommend?

For Postgres it shows connect_timeout 2, checkout_timeout 5 and statement_timeout 5000 in milliseconds. For MySQL and MariaDB it shows connect_timeout 1, read_timeout 1, write_timeout 1, checkout_timeout 5, max_execution_time 5000 in milliseconds and max_statement_time 5 in seconds. Both blocks are for the production entry in config/database.yml.

Which background job backend does production_rails recommend?

Sidekiq, configured with config.active_job.queue_adapter set to :sidekiq. For controlling misbehaving jobs it names ActiveJob::TrafficControl for disabling, throttling and concurrency limits, and shows BadJob.disable! as an example of a disable call.

Are the gems recommended by production_rails maintained by the guide's author?

Several are, and the README says so, describing the recommendations as coming from personal experience and work at Instacart and noting that a number of the linked projects were created by the author. Those include Strong Migrations, Slowpoke, Notable, PgHero, Ahoy, Ahoy Email and Field Test.

Does production_rails have versioned releases?

No GitHub releases are published and there are no tags. The last push to the repository was on 2026-03-31, so there is no version to pin and nothing has changed in roughly half a year.

Official sources

  1. ankane/production_rails on GitHub
  2. Issues
  3. License: CC-BY-4.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ankane-production-rails.svg)](https://hysenlabs.com/projects/ankane-production-rails)