Framework
google-parfait/tensorflow-federated avatar
google-parfait/tensorflow-federated

TensorFlow Federated: two APIs for machine learning on data that stays put

An open-source framework for machine learning and other computations on decentralized data.

2,454 stars605 forksPythonApache-2.0

At a glance

What is it?
TFF ships a high-level federated learning API and a lower-level functional core for writing new federated algorithms, with a single-machine simulation runtime included so you can experiment before deploying anywhere.
Who is it for?
TensorFlow Federated is a research framework for federated learning and other computations over decentralized data, and it suits you if you intend to try an algorithm rather than just train a model. The `tff.learning` layer covers the common case of applying federated averaging to a model you already have, and Federated Core is where the real flexibility lives, and also the real complexity.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Federated learning as a research problem, not just a technique

The README's opening definition is the one to keep in mind: TFF is a framework for machine learning and other computations on decentralized data, developed to support open research and experimentation with federated learning. Federated learning is described there as training a shared global model across many participating clients that keep their training data locally. The example given is prediction models for mobile keyboards trained without uploading sensitive typing data to servers.

That framing matters because it explains what kind of tool this is. TFF is not a training library you point at a dataset. It is a place to build and evaluate a training algorithm, including the aggregation step and the communication pattern, in a way that can later run in more than one environment. The README also draws a boundary the name often gets wrong: the same building blocks support non-learning computations such as aggregated analytics over decentralized data, which is where the Federated Core layer earns its keep.

The central design commitment is that federated computations are expressed declaratively, so they can be deployed to diverse runtime environments. Combined with a single-machine simulation runtime shipped alongside, that gives you a workable loop: write the computation once, run it in simulation against synthetic or sampled clients, then move the same definition to a real deployment.

This repository is `google-parfait/tensorflow-federated`, a mirror rather than the upstream `tensorflow/federated` location, and the package metadata points documentation at `tensorflow.org/federated`. The license is Apache 2.0.

The FL API and the Federated Core API

The README organizes TFF's interfaces into exactly two layers, and understanding the split tells you which one you need.

The `tff.learning` layer is the Federated Learning API. It offers high-level interfaces for applying the included federated training and evaluation implementations to existing TensorFlow models. The intended user is someone with a working model and data who wants federated averaging rather than a different aggregation strategy. The contributing guidance reinforces the layering: if you are interested in developer experience, the place to start is the implementations behind `tff.learning`.

Federated Core, the lower layer, is for expressing novel federated algorithms. The README describes it as combining TensorFlow with distributed communication operators inside a strongly-typed functional programming environment, and notes that this layer is the foundation `tff.learning` itself was built on. Strong typing here is not decoration: a federated computation has a type describing both its inputs and the grouping of values across clients, so a mistake in how you aggregate shows up as a type error rather than a subtly wrong training run.

Practically, you should not start at Federated Core unless you are doing something the built-in algorithms do not cover. It is the layer to read once you understand what `tff.learning` does for you, and the layer you extend when you need a different aggregation rule, a custom clipping or privacy mechanism, or a computation that is not learning at all.

Both layers are documented in separate files under `docs/`: `federated_learning.md` and `federated_core.md`.

Building blocks and the shape of the repository

The `examples/` directory is a better map of what the framework can do than the README's feature summary. It contains `simple_fedavg`, which is the canonical starting point, plus directories for `datasets`, `personalization`, `stateful_clients` and `custom_data_backend`.

Each of those names corresponds to a distinct problem. `simple_fedavg` is plain federated averaging. `datasets` is where prepared federated datasets live. `stateful_clients` addresses the case where a client keeps state across rounds, which changes the semantics of what a client returns. `custom_data_backend` is for federations whose data does not fit the built-in client data interfaces. `personalization` covers the case where a single global model is the wrong shape entirely and each client needs its own adaptation.

The build system is Bazel rather than setuptools alone. The tree has `WORKSPACE`, `BUILD`, `.bazelrc` and `.bazelversion`, and `pyproject.toml` exists alongside `requirements.in` with lock files for Python 3.12 and 3.13. The package metadata classifies the project as `Development Status :: 5 - Production/Stable` and pins `requires-python` to `>=3.12,<4`.

The runtime dependencies are worth reading as a description of scope. They include `tensorflow>=2.21.0,==2.21.*`, `numpy~=2.0`, `scipy~=1.16`, `grpcio~=1.74` for the gRPC communication path, `dp-accounting==0.6.0` for differential privacy accounting, `federated_language~=0.5.4`, `absl-py`, `dm-tree`, `ml_dtypes`, `portpicker` for finding free ports during simulation, and `tqdm`. That list says this is a framework with a full communication stack and privacy accounting built in, not a lightweight simulator.

What the release notes reveal about API churn

The release notes are unusually informative, and the honest summary is that TFF's API surface is still moving. The most recent tagged release is `v0.88.0`, TensorFlow Federated 0.88.0, published 2024-09-26.

That release added `tff.tensorflow.to_type` and `pack_args_into_struct` and `unpack_args_from_struct` to the public API under `framework`, while deprecating `tff.types.tensorflow_to_type` in favor of the new location. It also moved `tff.types.structure_from_tensor_type_tree` and `tff.types.type_to_tf_tensor_specs` into the `tff.tensorflow` package, and removed a cluster of older framework symbols including `merge_cardinalities`, `CardinalityCarrying`, `CardinalityFreeDataDescriptor`, `CreateDataDescriptor`, `DataDescriptor` and `Ingestable`. The removal of that whole data-descriptor family is the kind of change that breaks code written against a previous version.

`v0.87.0` from 2024-09-17 added an AdamW implementation to `tff.learning.optimizers` and made `None` gradients skip updates the way `tf.keras.optimizers` does. It also changed the behavior of `DPGroupingFederatedSum::Clamp` to set negatives to zero, with the stated reason that sensitivity calculation for differential privacy noise was calibrated for non-negative values. That is a correctness fix inside the privacy path rather than a cosmetic change, and it is the kind of detail worth reading before relying on a clipping routine.

The same release fixed a bug where `build_adafactor` incremented its step counter twice per `next()` call and a bug where tensor learning rates in `build_sgdm` failed with mixed dtype gradients. Version 0.86.0, published 2024-08-20, added `tff.tensorflow.transform_args` and continued the same pattern of additions, deprecations and removals.

Meanwhile commits continue into 2026, with the last push on 2026-09-26. Development is active while the release tags lag, so if you install a tagged version you may be months behind the branch you are reading about.

Where federated learning stops being the right tool

Federated learning solves a specific problem and carries specific costs, and TFF makes both visible.

The first cost is communication. A federated computation distributes a model to clients, collects updates and aggregates them, and the volume of that traffic is the dominant cost in many deployments. A simulation runtime on one machine does not model that cost honestly, because the round trips are free locally. That gap between a simulation and a real cross-network deployment is where federated projects most often discover their actual performance, and the README is candid that contributing simulation infrastructure was still described as forthcoming.

The second is that federated averaging assumes clients are roughly comparable. Where the data is heavily non-IID, where clients have wildly different amounts of data, or where personalization matters more than a single global model, the plain algorithm is a weak baseline. This is what the `personalization` and `stateful_clients` examples are about, and why the Federated Core layer exists for anyone who needs a different rule.

The third is operational simplicity. A framework with gRPC, privacy accounting, Bazel builds and a typed functional core is not something you drop into an existing training script in an afternoon. The path of least resistance for most teams is to write the aggregation loop yourself over a standard training loop, because that is often forty lines and requires no new dependency.

The distinction between federated and distributed is worth holding onto while evaluating this. A distributed computation splits work across machines that share the data store or coordinate tightly. A federated computation keeps data where it lives and moves the model instead, which is a different constraint and a different set of failure modes. TFF's simulation runtime exists to let you experiment with the second problem without hardware for the first.

PySyft and the wider landscape of privacy-preserving ML

The related searches for this project surface PySyft, which is the closest alternative in spirit and a useful comparison. PySyft targets privacy-preserving machine learning more broadly, covering the case where computation itself should happen away from the data owner rather than only the training data. TFF's emphasis is narrower and more research-oriented: express federated computations declaratively, simulate them locally, and try new aggregation algorithms.

That difference shows up in what each is good for. If you need differential privacy with a formal accounting story around a specific training loop, the presence of `dp-accounting` in TFF's dependencies signals that the framework supports it. If you need to run arbitrary computation against data someone else holds, in an environment designed for that, PySyft's model is closer to what you want. If you want to experiment with a new federated aggregation algorithm in TensorFlow, TFF's Federated Core is the more direct instrument.

Two more comparisons are worth making. Against a plain TensorFlow or PyTorch distributed setup, TFF's advantage is that the federated structure is the type, so a computation written against it does not silently assume a shared data store. Its disadvantage is the same one: you are committing to a framework, its build system and its release cadence rather than composing two existing libraries.

Against writing the loop by hand, TFF wins on the things that are easy to get subtly wrong: consistent aggregation, typed grouping of client values, privacy accounting, and a runtime you can switch between. It loses on everything you already have working. The deciding question is whether your research question is about the algorithm or about the model.

Editorial conclusion

TensorFlow Federated is a research framework for federated learning and other computations over decentralized data, and it suits you if you intend to try an algorithm rather than just train a model. The `tff.learning` layer covers the common case of applying federated averaging to a model you already have, and Federated Core is where the real flexibility lives, and also the real complexity. Two practical cautions first: the repository you would install from here is the `google-parfait` mirror rather than the upstream project, and the newest tagged release is 0.88.0 from 2024-09-26 even though commits continue into 2026. Start with the simulation runtime and the `simple_fedavg` example, and read `docs/get_started.md` before planning a deployment.

Frequently asked questions

What is the purpose of federated learning?

Federated learning trains a shared global model across many participating clients that keep their training data locally. The README uses prediction models for mobile keyboards as the example, trained without uploading sensitive typing data to servers.

What is the difference between federated and distributed computation?

Distributed computation splits work across machines that can coordinate tightly and often share storage. Federated computation keeps data where it lives and moves the model to the data instead, which is the constraint TensorFlow Federated is built to express through its typed computation environment.

How do I install TensorFlow Federated?

The README points to `docs/install.md` for instructions, covering both installing as a package and building from source. The package metadata requires Python 3.12 or newer and pins TensorFlow to the 2.21 series, and the repository builds with Bazel.

Which API should I start with, tff.learning or Federated Core?

Start with `tff.learning`, the high-level Federated Learning API, if you want to apply the included federated training and evaluation implementations to a model you already have. Move to Federated Core when you need to express an algorithm the built-in implementations do not cover.

Official sources

  1. google-parfait/tensorflow-federated on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-parfait-tensorflow-federated.svg)](https://hysenlabs.com/projects/google-parfait-tensorflow-federated)