# Ray reduces to three abstractions, and that is either the appeal or the tax

> Ray is a distributed runtime for Python with two layers: a core built on tasks, actors and objects, and five AI libraries on top of it. The one-line install hides a Bazel monorepo spanning Python, C++ and Java, and the general-purpose claim is both why people adopt it and what they later regret.

**ray-project/ray** — Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

- Repository: https://github.com/ray-project/ray
- Website: https://ray.io
- Stars: 43,949 · Forks: 8,098
- Language: Python
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/ray-project-ray

## Three abstractions, and everything else builds on them

The core is small enough to state in one line, which is the most useful thing about it. Tasks are stateless functions executed in the cluster, Actors are stateful worker processes created in the cluster, and Objects are immutable values accessible across the cluster. Everything in the AI libraries is a pattern over those three, so the learning curve is really one curve. The choice between them is a real design decision rather than a preference: a task is the right shape for work with no shared state, an actor is for anything that accumulates state or serialises access, and an object is the unit of distribution that both produce. A reader coming from a batch system will find the absence of a dataset abstraction at this level notable, because there is not one, and the Data library supplies it above the core rather than inside it.

## Five libraries that behave like five products

Above the core sit separate libraries, and the list is short enough to be worth reading as a menu rather than a bundle. Data provides scalable datasets for machine learning, Train handles distributed training, Tune does scalable hyperparameter tuning, RLlib is for reinforcement learning, and Serve is for scalable and programmable serving, with a separate Workflows library also linked. What the menu hides is that these are not interchangeable and each carries its own operational shape. Serve has its own generated protobuf surface, Tune has its own trial model, Train has its own placement semantics. Choosing the runtime and choosing the library are two decisions, and the second one is where most of the difficulty lives, because the documentation treats all five as siblings rather than as a progression.

## The install is one command and the floor is Python 3.10

Getting Ray in is not the hard part:

```bash
pip install ray
```

That is the whole instruction given for a stable install, with nightly builds pointed at a separate installation page rather than offered here. The interpreter requirement is stated in the project configuration rather than the README, and it is requires-python of 3.10 or newer, which is a genuine constraint for anyone still on 3.9 in an older environment. The README also makes a portability claim worth testing before you rely on it, that Ray runs on any machine, cluster, cloud provider and Kubernetes, and mentions a growing ecosystem of community integrations. What the install does not do is tell you what to put on the cluster, which is the decision that determines whether this is a five-minute setup or a week of capacity planning.

## The Python wheel is a slice of a Bazel monorepo that also holds C++ and Java

The repository makes clear that pip gives you a distribution of a much larger thing. The tree contains cpp/, java/, python/, src/, rllib/, bazel/, thirdparty/, ci/, release/, doc/ and docker/ directories alongside BUILD.bazel, WORKSPACE and a .bazelversion, and the top level has build-docker.sh, build-image.sh and build-wheel.sh along with three code generators, gen_py_proto.py, gen_ray_pkg.py and gen_redis_pkg.py. There is also a .rayciversion file, so the CI build is versioned separately from the package. For a Python developer the practical consequence is that the runtime you install has a C++ core with a Java surface beside it, and that debugging a Ray problem can mean reading code you did not know was involved. The linting and type-checking setup reflects the same sprawl, with pylintrc, pyrefly.toml, pytest.ini and a clang-format alongside the Python tooling.

## Generated protobuf modules are exempt from type checking on purpose

The type-checking configuration contains an exemption with a stated reason, which is a better signal of engineering practice than the absence of one would be. Two module globs, ray.serve.generated.* and ray.core.generated.*, are set to skip following imports and to ignore missing imports, and the comment explains that generated protobuf modules have no static names and should be treated as Any whether or not the generated files exist in the checkout, with a note about parity with a replace-imports-with-any setting in the pyrefly configuration. The lint configuration is similarly candid about its debt. Ruff runs at a line length of 88 with a long ignore list, and the comment on it carries a TODO crediting a contributor, explaining that some entries were carried over from flake8 and others from ruff and that they should eventually be removed. Flattening these into one tool is a real, acknowledged work item rather than a hidden one.

## One support channel carries a team commitment, the rest are community

The getting-involved table publishes expected response times per channel, and the differences are sharper than most projects admit. GitHub issues are the only row answered by the Ray OSS Team, with an estimated response under two days. Everything else is community: the Discourse forum for development discussion and usage questions at under a day, Slack for collaborating with other users at under two days, StackOverflow at three to five days, a Bay Area meetup monthly, and a social account monitored daily by developer relations staff. So the routing is explicit. Bugs and feature requests have a team behind them, and usage questions do not, which is worth knowing before you choose where to spend your effort on an unfamiliar problem. It is also the clearest signal in the README about what the project considers its own responsibility.

## General-purpose is the promise, and it is also what you pay for

The stated rationale is that machine learning workloads are increasingly compute-intensive and that a single-node development environment such as a laptop cannot scale to meet them, and the design goal is that the same code scales from a laptop to a cluster, with no other infrastructure required if your application is written in Python. Read as a promise that is attractive. Read as an operating model it is the risk, because a runtime designed to performantly run any kind of workload makes no recommendation about whether your workload suits it, and the repository contains no comparison with Spark or with any other system, so the decision is left entirely to the reader. The research record behind the design is published rather than asserted, with the Ray paper, the HotOS paper, the Exoshuffle paper on large-scale data shuffle, an Ownership paper on a distributed futures system for fine-grained tasks, and separate Tune and RLlib papers. The branch is receiving work, with the last push on 2026-09-29 and the newest release ray-2.58.0 published on 2026-08-23.

## Conclusion

Adopt Ray when your workload is Python, the parallelism is fine-grained rather than a batch pipeline, and you want one runtime for training, tuning, serving and reinforcement learning instead of four tools. Do not adopt it for a dataflow job that a batch system already does well, because the general-purpose design gives you placement and failure decisions nobody makes for you. Verify first which library you actually need, since Data, Train, Tune, RLlib and Serve are different products with different failure modes, and confirm the Python version you have is 3.10 or newer before anything else.

## FAQ

### What is Ray AI used for?

Ray is a distributed framework for scaling Python and AI applications, made up of a core runtime plus AI libraries covering datasets, distributed training, hyperparameter tuning, reinforcement learning and serving. It is aimed at workloads too large for a single machine.

### What does Ray do?

It runs Python code across a cluster using three abstractions: tasks as stateless functions, actors as stateful worker processes, and objects as immutable values shared across the cluster. The AI libraries are patterns built on top of those three.

### How is Ray different from Spark?

The repository does not make that comparison anywhere, so it does not answer it. What it does state is that Ray is Python-first, general-purpose and built on tasks, actors and objects, with a claim that the same code scales from a laptop to a cluster.

## Sources

- [Official documentation](https://ray.io)
- [Official README](https://github.com/ray-project/ray#readme)
- [Project repository](https://github.com/ray-project/ray)
- [Release notes](https://github.com/ray-project/ray/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ray-project-ray
