# github/scientist: A Ruby Library for Refactoring Critical Paths in Production

> github/scientist runs old and new code side by side in a live Ruby application, returning the old result while comparing the two. It is for teams that cannot afford to let a refactor break production, and it costs an experiment class plus a publish hook to get value from it.

**github/scientist** — :microscope: A Ruby library for carefully refactoring critical paths.

- Repository: https://github.com/github/scientist
- Stars: 7,759 · Forks: 504
- Language: Ruby
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/github-scientist

## What github/scientist is for, and who it is not for

github/scientist is a Ruby library for carefully refactoring critical paths. The README frames the problem directly: tests can guide a refactoring, but you want to compare the current and refactored behaviors under load. That is the gap it fills. A test suite tells you the new code passes the cases you thought of. An experiment tells you whether the new code agrees with the old code on the traffic you actually receive.

The intended user is a team running a Ruby service with a code path it is nervous about: permission checks, pricing, serialization, anything where a silent behavior change costs money or trust. The README's running example is a permissions check, `model.check_user(user).valid?` replaced by `user.can?(:read, model)`.

It is not a test framework, not a feature flag system, and not a load testing tool. If your code path is cheap to get wrong, or covered by tests you trust, the machinery here is overhead. It also assumes you can add a comparison and a publish path to production code, which is a real organizational cost, not just a gem install.

## How the control and candidate blocks run

The mechanism is two blocks and a run call. The `use` block holds the original behavior, called the control. The `try` block holds the new behavior, called the candidate. `experiment.run` always returns whatever the control returns.

Behind that return value the README lists what happens: the library decides whether to run the candidate, randomizes the order in which the two blocks run, measures wall time and cpu time in seconds, compares the candidate result to the control result, swallows and records exceptions raised in the candidate when `raised` is overridden, and publishes the information.

The randomization matters more than it looks. Running the control first every time would let caching or warm-up effects bias the comparison, so the order is shuffled. Measurement is split into wall time and cpu time, which is the difference between a slow dependency and slow local work.

Comparison defaults to `==`. When both blocks raise, the default is to compare the two errors' classes and messages with `==`. Both defaults are overridable, which is where most of the real work happens.

One detail worth knowing before you wire it in: if you declare no `try` blocks, none of the machinery is invoked and the control value is always returned. An experiment with no candidate is a no-op, not a failure.

## Installing the gem and running a first experiment

The library is a Ruby gem named `scientist`, distributed from the repository at github/scientist. The README's examples require it with `require "scientist"`. Add it to your Gemfile and install it with Bundler.

With the gem installed, the shortest working experiment uses the `science` helper after including the `Scientist` module. The helper instantiates an experiment and calls `run` for you:

```ruby
require "scientist"

class MyWidget
  include Scientist

  def allows?(user)
    science "widget-permissions" do |experiment|
      experiment.use { model.check_user(user).valid? } # old way
      experiment.try { user.can?(:read, model) } # new way
    end # returns the control value
  end
end
```

The method returns the control value, so callers see no change in behavior. What you should see depends on your experiment implementation: with the default, the README says the examples run but the `try` blocks do not run yet and nothing is published. To change that, define your own experiment class and include `Scientist::Experiment`:

```ruby
require "scientist/experiment"

class MyExperiment
  include Scientist::Experiment

  attr_accessor :name

  def initialize(name)
    @name = name
  end

  def enabled?
    true
  end

  def publish(result)
    p result
  end
end
```

Including `Scientist::Experiment` in a class automatically sets it as the default implementation via `Scientist::Experiment.set_default`. That call is skipped if you include the module in a module rather than a class, which is a quiet way to end up with experiments that never publish.

## Comparison, context and cleaning are where the work lives

The default `==` comparison is rarely enough. When the old code returns `User` instances and the new code returns `UserService::User` instances, `==` will report a mismatch on every call. The README's answer is `compare`, which takes a block receiving the control and candidate values. Its example maps both to logins and compares those arrays.

Error comparison has the same shape. If either block raises, the default compares classes and messages with `==`. `compare_errors` replaces that with your own block. The README shows a case where a message was reworded during the refactor, so the default would flag every raised error as a mismatch even though the behavior is equivalent.

Context is the second piece. Results are not useful without a way to identify them, so `context` takes a Symbol-keyed Hash of extra data, available in `Experiment#publish` via the `context` method. A class that runs many experiments can define `default_scientist_context` to merge shared data into every experiment's context.

The third piece is `clean`. When a value is large or contains objects you do not want to store, `clean` defines how to reduce it before publication. The README's example keeps only sorted logins from a list of `User` instances. The cleaned value is what appears in the published observations.

There is also `before_run`, for expensive setup that should only happen when the experiment is actually going to run. The README's example deep-copies a large object for the candidate, because the code under test modifies it in place. Without that, the candidate would see an object the control already mutated.

## Where github/scientist does not fit

The candidate runs in the same process as the control, on real requests. If the new code writes to a database, sends an email, charges a card, or mutates shared state, running it alongside the control duplicates that side effect. The library compares return values and errors; it does not sandbox anything. The README's own example of an in-place mutation, handled with a deep copy in `before_run`, is a small version of this problem. A write path is a much larger one.

The second limit is that the default comparison is `==`, and a mismatch is only as meaningful as the comparison you wrote. A loose `compare` block can hide real differences just as easily as a strict one can flood you with false positives.

The third limit is operational. Nothing here decides which requests run the candidate. That is `enabled?`, and the README points to a separate section on ramping up experiments rather than specifying a mechanism. There is no documented built-in rate limiting, sampling or kill switch beyond whatever you put in `enabled?`. The README does not document rollback. If you cannot cheaply turn an experiment off, you have added a production dependency without an off ramp.

Finally, if the behavior you are changing is already covered by a test suite you trust, or if the path is not critical, an experiment adds a publish pipeline and a comparison function for little return.

## How it compares to feature flags and test doubles

The closest common alternative is a feature flag that routes a percentage of traffic to the new code and lets you watch error rates. The difference in approach is what gets executed. A flag runs either the old path or the new path for a given request, so you observe the new code's real outcomes in production and compare aggregate metrics across the two groups. github/scientist runs both paths for the same request, returns the old one, and compares the two results directly. A flag gives you production behavior on a subset; an experiment gives you a same-input diff on every enabled request, at the cost of running the candidate code you do not yet trust.

Against test doubles, the difference is the input. A mock or stub fixes the inputs you thought to write down. An experiment uses the inputs production actually produces, including the ones nobody imagined. Tests are also repeatable and fast; experiments depend on live traffic and on a publish path you have to build.

The two are not exclusive. A common sequence is to use tests to get the new code close, then run an experiment to find the cases tests missed, then use a flag to shift traffic once the comparison comes back clean.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-18. The most recent release listed is v1.6.5, dated 2024-12-16. So the release cadence and the commit activity are not in step: the tree has moved since the last tagged version. If you pin to a release, you are pinning to something older than the repository state, and the README's examples may describe behavior that has not shipped in a gem.

The licence is MIT, per the repository's LICENSE.txt. MIT is permissive: it allows use, modification and redistribution with the licence and copyright notice retained. That is a summary of the licence text, not legal advice; read LICENSE.txt and your own organization's policy before relying on it.

Upgrade cost is low by design. The public surface is small: `science`, `use`, `try`, `run`, `compare`, `compare_errors`, `context`, `clean`, `before_run`, `enabled?`, `raised` and `publish`. Experiments live in your application code, so a gem upgrade does not rewrite them. The real maintenance burden is on your side: an experiment that nobody removes keeps running the candidate forever. Nothing in the README describes an expiry mechanism, so deleting the `try` block and the experiment when the comparison is clean is a manual step you have to schedule.

## Conclusion

Adopt github/scientist when you are changing a code path whose failure would be expensive and you have a way to publish and read comparison results. Do not adopt it as a general test replacement, and do not enable candidate execution without an enabled? gate you control. Before wiring it into a live path, verify three things in your own environment: that your experiment class sets itself as the default via Scientist::Experiment.set_default, that your compare or compare_errors block matches the shapes your control and candidate actually return, and that publish writes somewhere you will actually look. The README does not document rollback, so plan for disabling an experiment by returning false from enabled? rather than expecting a built-in kill switch.

## FAQ

### How do I use github/scientist in a Ruby method?

Include the Scientist module in the class, then wrap the old behavior in `experiment.use` and the new behavior in `experiment.try` inside a `science` block. The method returns the control value, so callers see no change. Without a custom experiment class, the README states the try blocks do not run and nothing is published.

### What is github/scientist?

It is a Ruby library for carefully refactoring critical paths, published as the `scientist` gem under the MIT licence. It runs a control block and a candidate block, returns the control result, and compares the two.

### Does github/scientist publish its comparison results automatically?

No. The README states that with the default experiment the examples run but nothing gets published. You define a class that includes Scientist::Experiment and implement `publish(result)` to send results wherever you want them.

### Can github/scientist compare values that are not equal with ==?

Yes. The default comparison uses `==`, and you can override it with `compare`, which receives the control and candidate values. For raised errors, `compare_errors` replaces the default class-and-message comparison.

### Which value does github/scientist return from an experiment?

It always returns whatever the `use` block returns, which the README calls the control. The candidate result is recorded and compared, never returned to the caller.

## Sources

- [github/scientist on GitHub](https://github.com/github/scientist)
- [Issues](https://github.com/github/scientist/issues)
- [License: MIT](https://github.com/github/scientist/blob/main/LICENSE)
- [README](https://github.com/github/scientist/blob/main/README.md)
- [Releases](https://github.com/github/scientist/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/github-scientist
