mrjob: writing Hadoop Streaming jobs as Python classes
Run MapReduce jobs on Hadoop or Amazon Web Services
At a glance
- What is it?
- Yelp's mrjob lets you write a MapReduce job as a Python class and run it on EMR, Dataproc, your own Hadoop cluster, or locally for testing. The README is honest about which of those paths it supports fully and which it only partly does.
- Who is it for?
- mrjob is the right tool when the target really is Hadoop Streaming and you want the same script to run on a laptop and on EMR, and the wrong tool when the target is a modern Spark or Kubernetes batch job, where the ecosystem has moved on. GitHub reports the last push on 2026-04-02 and the README points at v0.7.4 as the stable documentation, so judge the API against that version.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One class, three mapper and reducer methods
The README's example is the classic word count, written as a Python class that subclasses `MRJob` from `mrjob.job`. Three methods carry the whole job: a `mapper` that takes a key and a line and yields word and count tuples, a `combiner` that sums counts per word before the shuffle, and a `reducer` that sums what survived the combiner. The class ends with a `MRWordFreqCount.run()` call under a `__main__` guard, which is what makes the file both importable and runnable.
from mrjob.job import MRJob
import re
WORD_RE = re.compile(r"[\w']+")
class MRWordFreqCount(MRJob):
def mapper(self, _, line):
for word in WORD_RE.findall(line):
yield (word.lower(), 1)The interesting structural choice is that there is no driver code, no argument parsing and no Hadoop configuration in the job itself. The same class runs everywhere because the runner is chosen at invocation time with a `-r` flag. That is the whole design of the library in one line: separate the job logic from the cluster.
The same script runs four places with one flag
The README's Try It Out section runs one file four ways, and the differences are only the runner name:
python mrjob/examples/mr_word_freq_count.py README.rst > counts
python mrjob/examples/mr_word_freq_count.py README.rst -r emr > counts
python mrjob/examples/mr_word_freq_count.py README.rst -r dataproc > counts
python mrjob/examples/mr_word_freq_count.py README.rst -r hadoop > countsThe first line is the one you will run most, because running the job locally with no cluster is how you iterate on mapper and reducer logic. It still runs through Hadoop Streaming locally, so the local run is a real exercise of the pipeline rather than a simulation of one.
The feature list is where mrjob stops being a job writer and starts being a deployment tool. The README claims multi-step jobs where one map-reduce step feeds the next, launching Spark jobs on EMR or your own cluster, and something it calls duplicating your production environment inside Hadoop: uploading your source tree into the job's `$PYTHONPATH`, running make and other setup scripts, setting environment variables such as `$TZ`, and installing Python packages from tarballs on EMR. All of that is configured through an `mrjob.conf` file rather than in code.
Credentials come from the environment, not the config file
Setup for the two cloud targets is deliberately small. For EMR you create an AWS account, get an access key and secret from the account page, and set `$AWS_ACCESS_KEY_ID` and `$AWS_SECRET_ACCESS_KEY`. For Dataproc you enable Cloud Storage, the Storage JSON API and the Dataproc API in the API manager, create a service account key of type JSON, install the Google Cloud SDK, and set `$GOOGLE_APPLICATION_CREDENTIALS` to the key file. For your own Hadoop cluster the README says no setup is needed at all.
The dependency story follows the same shape. As of v0.7.0 the README states that AWS and Google Cloud support are optional dependencies, installed through extras such as `mrjob[aws]`, and the base install is the plain form:
pip install mrjob
pip install mrjob[aws]`setup.py` confirms the shape. The `install_requires` list contains a single package, `PyYAML`, because that is what reads the config. The `aws` extra pulls in `boto3` and `botocore`, and the `google` extra pulls `google-cloud-dataproc`, `google-cloud-logging` and `google-cloud-storage`. There are also `rapidjson`, `simplejson` and `ujson` extras, which are faster JSON parsers, so the JSON layer is swappable.
Configuration search order and what it lets you change
For anything beyond the basics, mrjob reads a config file, and the README is specific about where from. It looks in three places, and the order matters when you are debugging why a setting did not apply:
MRJOB_CONF=./mrjob.conf
~/.mrjob.conf
/etc/mrjob.confThe first entry is really the contents of the `$MRJOB_CONF` environment variable, so any of those three lines can hold a path or, in the first case, inline contents. A file dropped in `~/.mrjob.conf` therefore applies to every job you run as that user, which is the right place for cluster defaults and the wrong place for anything experiment specific.
Two cloud-specific features in the README's feature list have no equivalent in the local or self-hosted path. The SSH tunnel to the Hadoop job tracker is marked EMR only, which makes sense because the job tracker UI is not exposed publicly. Spark launching is described for EMR and your own cluster, not for Dataproc. Both of those are the kind of detail that decides whether the library fits your target, and both are easier to find here than in most project documentation.
The honest limitation: Hadoop Streaming is the target, not the framework
The name of the thing mrjob generates is Hadoop Streaming, and that fact bounds everything the library can do. The README links the Apache Hadoop Streaming documentation as its reference, which means your job is a Python script piped through the JVM. Mapper and reducer code runs in a Python process on each task, so per-record Python overhead is the performance ceiling, and the combiner exists precisely because that overhead is the thing to minimise.
This is not a criticism of the library; it is the reason the library is small and readable. mrjob has no GitHub releases, and its repository tree is modest: `mrjob/`, `docs/`, `tests/`, `setup.py`, `setup.cfg`, `CHANGES.txt` and a `Makefile` whose default target runs the test suite. There is no plugin architecture and no job graph layer to learn.
What you give up by choosing it is the newer ecosystem. If your batch work would otherwise be a Spark job written in PySpark, mrjob will not help, and if you have outgrown single-pass map-reduce you will find yourself implementing multi-step logic by hand in config. The README's own comparison is with Hadoop Streaming and Elastic MapReduce, which tells you the intended audience was the 2011 era of Hadoop and has stayed there.
One more caveat about the framing. The README calls EMR support full and Dataproc support basic, which is a distinction worth taking at face value. If you are deploying to Google Cloud, the parts of the documentation that assume Amazon are not describing your situation.
Reading the repository as a map of the package
The layout tells you more than the feature list does. `mrjob/` is the package itself, `docs/` is the documentation that the README links to at mrjob.readthedocs.org, `tests/` is the suite the `Makefile` invokes through `setup.py test`, and `CHANGES.txt` is where behaviour changes get recorded. `MANIFEST.in` and `setup.cfg` are there for packaging, and `.travis.yml` shows the continuous integration the project used.
`setup.py` also reveals that parts of the project have changed hands over time. The copyright header at the top of the file lists Yelp and Contributors for 2009 to 2015 and 2016 to 2017, then Google Inc for 2018, Yelp again for 2019, and Affirm, Inc for 2020. That is a maintenance trail in four lines, and it lines up with the feature list: AWS support first, then Google Cloud support added later by a different set of hands.
The Makefile is the shortest useful summary of the workflow. `PYTHONPATH` is set to the current directory, `all` depends on `test`, and `install` shells out to `setup.py install`. Nothing about running a job appears there, because running a job is the user's business, not the project's.
GitHub reports the repository as Python, not archived, last pushed 2026-04-02, on the `master` branch. GitHub reports no license identifier, though `LICENSE.txt` exists in the tree and `setup.py` carries an Apache License 2.0 header. A library without a machine-readable licence is a small friction point worth raising if you intend to depend on it inside a company.
Editorial conclusion
mrjob is the right tool when the target really is Hadoop Streaming and you want the same script to run on a laptop and on EMR, and the wrong tool when the target is a modern Spark or Kubernetes batch job, where the ecosystem has moved on. GitHub reports the last push on 2026-04-02 and the README points at v0.7.4 as the stable documentation, so judge the API against that version. Start with the word count example in `mrjob/examples` and run it locally before touching a cluster.
Frequently asked questions
How do I install mrjob and its cloud extras?
The base install is `pip install mrjob`, which pulls in only `PyYAML`. AWS support is an optional extra installed as `pip install mrjob[aws]`, and `setup.py` shows it adding `boto3` and `botocore`. The `google` extra exists for Dataproc and brings in the Google Cloud Dataproc, logging and storage clients.
Which runners can I target from a single mrjob script?
The README's example runs the same file with `-r emr`, `-r dataproc` and `-r hadoop`, or with no runner flag for a local run. EMR support is described as full, Dataproc support as basic, and your own Hadoop cluster needs no credentials.
Where does mrjob look for its configuration file?
It checks the contents of `$MRJOB_CONF`, then `~/.mrjob.conf`, then `/etc/mrjob.conf`. Anything beyond the standard setup, such as running in a non-default AWS region or uploading a source tree, is configured in that file rather than in code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yelp-mrjob)