TensorFlowOnSpark: Running TensorFlow Training Jobs Inside a Spark Cluster
GitHub describes it as TensorFlowOnSpark brings TensorFlow programs to Apache Spark clusters.. The repository metadata lists Python as its primary language. The metadata lists the Apache-2.0 license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- TensorFlowOnSpark is a pip-installable Python library that launches TensorFlow workers and parameter servers on Spark executors. It fits teams with existing Spark and HDFS pipelines; its last push was on 2022-04-21, so treat it as a frozen dependency.
- Who is it for?
- Adopt TensorFlowOnSpark if your training data already lives in HDFS and your cluster runs Spark, and you want your existing TensorFlow script started on executors with a Spark-managed lifecycle. Do not adopt it for a new greenfield project on a single GPU box, for Windows, or on a cluster where you cannot control the Python environment on every executor.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Probably not. The repository last received commits 39 months ago, on July 10, 2023.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem TensorFlowOnSpark targets: TensorFlow jobs that need a Spark cluster, not a new cluster
TensorFlow has its own distribution story: a chief worker, parameter servers, and workers that talk to each other over the network. Spark has a different one: a driver, executors, and a scheduler that decides where those executors run. If a team already operates a Hadoop or Spark cluster and stores training data in HDFS, standing up a second, separate TensorFlow cluster means duplicating scheduling, resource accounting, and access control.
TensorFlowOnSpark addresses that overlap. The README describes it as bringing scalable deep learning to Apache Hadoop and Apache Spark clusters, and the stated goal is to minimize the amount of code changes required to run existing TensorFlow programs on a shared grid. The audience is therefore specific: data platform teams with a Spark cluster already in production, who want training jobs to be scheduled by the same machinery as their ETL, and who want the data to stay where it already is.
It is not a framework for designing models. There is no new layer of abstraction over Keras or tf.estimator. The library is plumbing between the Spark scheduler and the TensorFlow runtime, and the model code stays yours.
How the Spark-compatible API starts, feeds and stops a TensorFlow cluster
The README lays out three phases. Startup launches the TensorFlow main function on the executors, along with listeners for data and control messages. Shutdown shuts down the TensorFlow workers and PS nodes on the executors. Between those two, data ingestion has two modes.
InputMode.TENSORFLOW has TensorFlow read data files directly from HDFS using its own input APIs. Nothing flows through Spark at that point; Spark is only the launcher and the resource manager. InputMode.SPARK instead sends Spark RDD data to the TensorFlow nodes through a TFNode.DataFeed class, and the README notes that the project uses the Hadoop Input/Output Format to read TFRecords from HDFS.
The distinction matters operationally. In SPARK mode the data path goes through Spark, so the RDD partitioning and the number of TensorFlow workers have to line up sensibly. In TENSORFLOW mode each worker reads its own shard from HDFS, which removes Spark from the data path but means the file layout, not the RDD, determines how work is split. The README also mentions server-to-server direct communication for faster learning when the network allows it, which is a topology constraint rather than a feature toggle: executors that cannot reach each other directly will not get that benefit.
TensorBoard is listed among the supported TensorFlow functionality, so the usual summary-writing workflow still applies, but the README does not describe how event files are collected from executors. That is a gap worth knowing about before you plan dashboards.
Installing TensorFlowOnSpark and running a first job from the examples directory
The README gives the install as a pip package, with the version choice driven by your TensorFlow major version. Note the asymmetry: TensorFlow 2.x uses the plain package name, while TensorFlow 1.x requires pinning to 1.4.4.
# for tensorflow>=2.0.0
pip install tensorflowonspark
# for tensorflow<2.0.0
pip install tensorflowonspark==1.4.4After installation you should have the tensorflowonspark module importable in the same Python environment as your TensorFlow install. The repository's requirements.txt lists h5py, numpy, packaging, py4j, pyspark, scipy, setuptools and tensorflow, so a bare environment will pull Spark's Python bindings as well.
For a first real run, the repository ships examples under examples/mnist/, examples/resnet/, examples/segmentation/ and examples/utils/. The README points at the wiki for environment-specific getting-started guides: a single-node Spark Standalone guide, a YARN guide, and an AWS EC2 guide. Those guides, not the README, are where the actual submission command lives.
The README is explicit about version pairing in the other direction too: since TensorFlow 2.x breaks API compatibility with TensorFlow 1.x, the examples have been updated accordingly, and TensorFlow 1.x users must check out the v1.4.4 tag for compatible examples and instructions. If you mix a 2.x library install with 1.x example code, nothing in the README says the result is supported.
One deployment note is stated plainly: the Windows operating system is not currently supported, with the README linking to issue 36 as the reason.
Where TensorFlowOnSpark is the wrong tool
The clearest limitation is the platform boundary. Windows is unsupported per the README, so a Windows-based development or training host is out before you evaluate anything else.
The second limitation is version coupling. The library is installed as a pip package whose behavior depends on the TensorFlow major version, and the README's own instructions fork at TensorFlow 2.0. That means an upgrade of TensorFlow is an upgrade of TensorFlowOnSpark, and the examples you copy must come from the matching tag. Teams that track TensorFlow releases quickly will find themselves managing two release trains at once.
The third is the environment on the executors. Because the library launches your TensorFlow main function inside Spark executors, the Python environment on every executor has to satisfy the same requirements as the driver. The README does not document a mechanism for shipping a virtual environment to executors, so this is a cluster-configuration problem you solve yourself.
Finally, the project is not a general-purpose distributed training framework. If your data is not in HDFS and not in a Spark RDD, and you do not already run Spark, the integration work buys you nothing. A single-machine multi-GPU setup, or a managed training service, will get you to a first result with fewer moving parts.
How it differs from Horovod on Spark and from plain distributed TensorFlow
The obvious comparison is Horovod's Spark integration, since both let a Spark job start distributed TensorFlow work. The difference is in the communication model. Horovod is built around ring allreduce over MPI or Gloo, and its Spark runner exists to place ranks on executors. TensorFlowOnSpark instead keeps TensorFlow's native worker and parameter-server roles and adds a Spark-compatible API to start and stop them, with the README calling out server-to-server direct communication as the fast path when the network permits.
That difference shows up in how you think about the job. With a parameter-server layout, the number of PS nodes and workers is a modeling decision you make explicitly, and the SPARK input mode gives you a DataFeed path from RDDs into those nodes. With an allreduce layout, there is no PS tier to size, but the collective has to complete on every step.
The second comparison is against running distributed TensorFlow directly with tf.distribute and a cluster spec you maintain yourself. That approach has no Spark dependency at all, which is an advantage if you do not have a Spark cluster, and a disadvantage if you do, because you then have to solve scheduling, resource isolation and data locality separately. TensorFlowOnSpark's value is entirely in reusing the Spark cluster you already run; remove that premise and the library has little to offer.
Maintenance status, upgrade cost and the Apache-2.0 licence
The repository is not archived, but the last push was on 2022-04-21, and the most recent release listed is v2.2.5 on the same date. Treat the project as stable but not moving. Any bug you hit in the Spark or TensorFlow integration is one you will be fixing locally, and the README's pointers to the wiki and the user group are the support channels it names.
The practical consequence is dependency pinning. Because the install instructions fork on the TensorFlow major version, and because requirements.txt leaves tensorflow and pyspark unpinned, a fresh install today resolves to whatever the current versions are, which is not the combination the 2022 releases were built against. Pin tensorflow, pyspark and py4j explicitly in your own environment rather than relying on the package metadata.
On licensing, the README states that the use and distribution terms are covered by the Apache 2.0 licence, and the repository carries a LICENSE file plus a Code-of-Conduct.md and Contributing.md. Apache-2.0 includes a patent grant and permits commercial use, but it also carries attribution and notice requirements, and this project's source headers carry a Yahoo copyright line. If you redistribute a modified copy or bundle it into a product, have your own counsel read the LICENSE and NOTICE handling rather than treating this paragraph as advice.
Editorial conclusion
Adopt TensorFlowOnSpark if your training data already lives in HDFS and your cluster runs Spark, and you want your existing TensorFlow script started on executors with a Spark-managed lifecycle. Do not adopt it for a new greenfield project on a single GPU box, for Windows, or on a cluster where you cannot control the Python environment on every executor. Before committing, verify three things on your own cluster: that the pip-installable tensorflowonspark version matches your TensorFlow major version, that Spark's executor Python matches the driver Python, and that your Spark deployment mode supports the network topology the training job needs. The repository's last push was on 2022-04-21, so budget for pinning every dependency rather than expecting upstream fixes.
Frequently asked questions
Is TensorFlow still used?
The repository does not contain usage statistics for TensorFlow, so this cannot be answered from it. What the README does show is that TensorFlowOnSpark publishes separate install instructions for tensorflow>=2.0.0 and for tensorflow<2.0.0, which means both major lines were still being supported when those instructions were written.
Is TensorFlow free to use?
The repository does not state TensorFlow's licence. It does state that TensorFlowOnSpark itself is covered by the Apache 2.0 licence, with the terms in the LICENSE file in the repository.
Is Spark replacing Hadoop?
The README treats them as complementary rather than competing: TensorFlowOnSpark is described as bringing scalable deep learning to Apache Hadoop and Apache Spark clusters, and the InputMode.TENSORFLOW path reads data files directly from HDFS.
Is Spark 100x faster than MapReduce?
The README makes no performance comparison between Spark and MapReduce, so no figure can be given. Its only speed-related claim is about server-to-server direct communication achieving faster learning when available.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yahoo-tensorflowonspark)