grosser/parallel: Ruby map-reduce across processes, threads and Ractors
Ruby: parallel processing made simple and fast
At a glance
- What is it?
- The parallel gem runs a block over a collection in forked processes, threads or Ractors, with one API for all three. It is aimed at Ruby developers doing map-reduce style work such as parallel downloads and uploads, and its main trade-off is that each backend has different memory and safety behaviour.
- Who is it for?
- Use grosser/parallel when you have a collection of independent items and a block that is slow enough to be worth spreading across CPUs or waiting threads. Do not use it for work that must run in a fixed order, for code that mutates shared state under in_processes, or for Ractor mode in production, since the README calls Ractors experimental and unstable.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 24 days ago.
- What is it written in?
- Mainly Ruby, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What grosser/parallel is for
The gem targets map-reduce shaped work: you have a collection, you have a block that is slow, and the items do not depend on each other. The README names parallel downloads and uploads as the example use case, and the API is built around that shape. Parallel.map, Parallel.each, each_with_index, map_with_index and flat_map all take a collection and a block. There is also Parallel.any? and Parallel.all?, which answer short-circuit questions over a collection without you writing the coordination yourself. The intended reader is a Ruby developer who has already noticed that a loop is spending its time waiting on IO or burning one core, and who wants the concurrency model to be a keyword argument rather than a separate library per backend.
Three backends behind one map call
The backend is chosen per call. in_processes: 3 forks workers, in_threads: 3 spawns threads, and in_ractors: 3 uses Ruby 3.0+ Ractors. The README describes the workers as a pool that grabs the next piece of work when it finishes, which is why the results come back in the original order even though completion order differs. The three backends are not interchangeable. Processes give speedup through multiple CPUs and protect variables from change, at the cost of extra memory. Threads give speedup for blocking operations and allow shared mutable variables, with no extra memory. Ractors give CPU speedup with no extra memory and are very fast to spawn, but the README calls them experimental and unstable, and variables must be passed explicitly, for example by mapping to pairs that include the values you need. Because the choice is per call, the same block can be benchmarked under each mode by setting in_threads: 0 or in_processes: 0, which the README lists as a tip for running the same code with different setups.
Installing grosser/parallel and running a first map
The README gives one install command and no other setup step. It is a plain Ruby gem, so there is nothing to configure before the first call.
gem install parallelAfter that, the README's first example maps a three-element array with no options, which runs in two processes on a two-CPU machine. The block receives one element at a time and the call returns the results in input order.
# 2 CPUs -> work in 2 processes (a,b + c)
results = Parallel.map(['a','b','c']) do |one_letter|
SomeClass.expensive_calculation(one_letter)
endTo pin the worker count instead of letting the gem detect it, pass in_processes or in_threads. The README shows both forms side by side, and the same collection finishes in one run when the count matches the number of items.
# 3 Processes -> finished after 1 run
results = Parallel.map(['a','b','c'], in_processes: 3) { |one_letter| SomeClass.expensive_calculation(one_letter) }
# 3 Threads -> finished after 1 run
results = Parallel.map(['a','b','c'], in_threads: 3) { |one_letter| SomeClass.expensive_calculation(one_letter) }Controlling flow: Break, Kill, worker numbers and hooks
Stopping early is explicit. Raising Parallel::Break inside the block stops after all current items are finished, and the README shows that the raised value becomes the return value of the map call. Parallel::Kill is the harder stop: it kills all sub-processes instantly, and the README warns to only use it when whatever is executing is safe to kill at any point. For progress reporting there are :start and :finish hooks, called on the main process and protected with a mutex. :start receives the item and index, :finish receives the item, index and result. Setting finish_in_order: true makes the :finish hook fire in input order, which the README notes will take longer to see initial output. There is also Parallel.worker_number, which tells a task which worker slot it is running in, and the README's example output shows items landing on workers 0 and 1 in an interleaved order rather than in sequence. If you need to know the index and nothing else, the README points at Parallel.each_with_index as the more performant option.
Where the gem bites: ActiveRecord, autoloading and the fork boundary
The README spends more space on ActiveRecord than on any other integration, which is a fair signal about where users get stuck. With threads you need connection pooling and should adjust the pool size in config/database.yml; with forks you need reconnects. The README offers three fixes of decreasing confidence. The one it calls reproducible is calling User.connection.reconnect! after the Parallel.each block. The other two are marked "maybe helps": wrapping the work in ActiveRecord::Base.connection_pool.with_connection, or reconnecting once inside every fork with a memoized flag. A separate failure mode is NameError: uninitialized constant, described as a race when models are autoloaded inside parallel threads in lazy-loading environments such as development, test or migrations. The stated fix is to load the classes before the parallel block, either with require '<modelname>' or by referencing ModelName.class. Neither of these is a bug in the gem so much as the cost of forking or threading underneath a framework that assumes a single long-lived connection.
The Marshal pipe and the HMAC serializer
Worker processes talk to the parent over an anonymous pipe, and by default the payload is serialized with Marshal. The README is direct about the consequence: a same-UID attacker who can reopen the pipe could inject a forged Marshal payload into the parent, which would be remote code execution. The mitigation offered is Parallel::Serializer::Hmac.new, which length-prefixes and HMAC-SHA256 signs each message with a per-worker secret generated before fork. On mismatch it raises SecurityError, which the README notes is not a StandardError, so a bare rescue will not catch it. This option matters if you have hardened the host against ptrace and /proc/<pid>/mem access, for example with ptrace_scope >= 2, and want to close the /proc/<pid>/fd/<n> pipe-reopen vector as well. If you are not in that threat model, the default is the default.
How it compares with concurrent-ruby
concurrent-ruby is the obvious alternative, and the difference is in what each one owns. concurrent-ruby is a collection of concurrency primitives: thread pools, futures, promises, atomic references and higher-level abstractions you compose yourself. grosser/parallel owns one narrow thing: applying a block to a collection across workers, with the backend selected by keyword. You do not get a futures API or a promise graph from parallel, and you do not get a map-reduce loop from concurrent-ruby without assembling it. The README itself points at the boundary: it notes that if you want the process count to account for CPU quota rather than the count the OS reports, you add concurrent-ruby to your Gemfile. That is a small but telling admission that the CPU detection is the OS view, and the quota-aware path lives in a different library.
Configuration, maintenance and licence
Two settings are worth knowing before deployment. PARALLEL_PROCESSOR_COUNT=16 overrides the detected processor count, which the README frames as a way to reconfigure a tool that uses parallel without inserting custom logic. isolation: true prevents reuse of previous worker processes. The interrupt signal is configurable too: INT from Ctrl+c is caught by default, and interrupt_signal: 'TERM' catches TERM from kill. On maintenance, the last push to the default branch was on 2026-09-06, and the repository is not archived. The README still carries a TODO section about replacing signal trapping, which suggests the signal handling is known to be the rough part. The licence is MIT, per MIT-LICENSE.txt in the repository root, which is permissive and imposes no copyleft obligation on your application; that is a description of the licence file, not legal advice. There are no retrieved release notes, so upgrade cost cannot be judged from release history here.
Editorial conclusion
Use grosser/parallel when you have a collection of independent items and a block that is slow enough to be worth spreading across CPUs or waiting threads. Do not use it for work that must run in a fixed order, for code that mutates shared state under in_processes, or for Ractor mode in production, since the README calls Ractors experimental and unstable. Before adopting it, check whether your process count should come from PARALLEL_PROCESSOR_COUNT rather than the OS CPU count, and confirm that your ActiveRecord setup either reconnects after forking or wraps work in ActiveRecord::Base.connection_pool.with_connection.
Frequently asked questions
How do I install grosser/parallel?
The README gives a single command, gem install parallel. There is no other setup step documented before the first call.
Does grosser/parallel work on Ruby 3.0 and later for Ractor mode?
The README lists Ractors as Ruby 3.0+ only, and describes them as experimental and unstable. It also notes that variables must be passed in explicitly and that start and finish hooks are called on the main thread.
Why does grosser/parallel lose the ActiveRecord connection?
The README states that multithreading needs connection pooling and forks need reconnects, and that the connection pool size in config/database.yml should be adjusted when multithreading. It gives User.connection.reconnect! after the parallel block as the reproducible fix.
How do I stop a grosser/parallel map early?
Raise Parallel::Break to stop after all current items are finished, or Parallel::Kill to stop all sub-processes by killing them instantly. The README warns to use Kill only when the executing work is safe to kill at any point.
How do I set the number of workers in grosser/parallel without changing code?
Set the PARALLEL_PROCESSOR_COUNT environment variable, for example PARALLEL_PROCESSOR_COUNT=16, and the gem uses that value instead of the number of processors detected. The README describes this as a way to reconfigure a tool that uses parallel without inserting custom logic.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/grosser-parallel)