# PySyft v2: Running Remote Data Science Jobs Through Google Drive

> PySyft 0.10+ split into a sync engine and a Remote Data Science client. It lets a data owner approve a peer, share mock data, and execute submitted Python jobs against the private copy without new infrastructure.

**OpenMined/PySyft** — Perform data science on data that remains in someone else's server

- Repository: https://github.com/OpenMined/PySyft
- Website: https://www.openmined.org/
- Stars: 10,041 · Forks: 2,007
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/openmined-pysyft

## The problem PySyft solves is approval, not cryptography

Most privacy tooling starts from the assumption that you need math to protect data. PySyft starts from an organizational fact: the person who wants an analysis and the person who holds the data are different people, often at different organizations, and the data holder will not hand over a copy. The README states the goal plainly: private data never leaves the data owner's machine, and only approved results are shared.

The intended user is a data scientist who has been given mock data and wants to run real analysis, paired with a data owner who is willing to run code but not to export a table. That is a narrower audience than the phrase federated learning suggests. There is no aggregation of gradients from thousands of phones here. The unit of work is a submitted Python file, reviewed and approved by one human, executed once.

## How the sync engine and the RDS client divide the work

The README draws a line that older PySyft users will not expect. Since 0.10, `syft` is the sync engine and helper layer, imported as `import syft as sy`. Datasets and jobs live in a separate package, `syft-rds`, which exposes `login_do` and `login_ds`. The README tells anyone depending on the legacy API to pin `syft<0.10`, which is a clear signal that this is a break, not an addition.

The transport is deliberately dull. Connections are described in docs/connections.md as a Google Drive transport layer, and the README calls the design transport-agnostic, working over Google Drive today and extensible to any file-based transport. Sync is offline-first: peers can be disconnected and changes reconcile when connectivity returns. Job execution is isolated, described as sandboxed Python virtual environments with controlled access to private data.

The data flow has a distinctive property. Inside a job, `resolve_dataset_file_path` resolves to the private data, while outside it the data scientist only ever sees the mock. The same code path therefore behaves differently depending on where it runs, which is what makes the mock/private split work without the analyst rewriting their script.

## Installing PySyft and running a first job

The README's quick start assumes two parties, a data owner (DO) and a data scientist (DS). Each installs on their own machine. The data scientist needs two packages; the data owner adds the background services package, which the README describes as providing email notifications, auto-approval and a TUI dashboard.

```bash
# Data scientist
uv pip install "syft>=0.10.0" "syft-rds>=0.6.0"
# Data owner (adds the background services)
uv pip install "syft>=0.10.0" "syft-rds>=0.6.0" "syft-bg>=0.3.12"
```

Both parties then log in. The README notes that the default path is colab auth, and that non-colab users should pass `token_path` instead. The peer request must be approved before anything flows, which the README lists as a feature: data owners approve each collaborator explicitly.

```python
import syft as sy                          # sync engine + helpers
from syft_rds import login_do, login_ds    # datasets + jobs

do = login_do(email="do@org.com")
ds = login_ds(email="ds@org.com")

ds.add_peer("do@org.com")
do.approve_peer_request("ds@org.com")
```

The data owner creates a dataset with two paths, a mock file and a private file, and names the user who may see it. After `do.sync()` and `ds.sync()`, the data scientist can list datasets with `ds.datasets.get_all()`. The README's example uses `mock.txt` and `private.txt` that you create yourself and fill with text.

The submitted script is ordinary Python. It imports `syft`, calls `resolve_dataset_file_path` with the dataset name, and writes a result into `outputs/`. In the README's example the result is just the length of the file content.

```python
# analysis.py
import json
import syft as sy

data_path = sy.resolve_dataset_file_path("census")
with open(data_path, "r") as f:
    data = f.read()

with open("outputs/result.json", "w") as f:
    json.dump({"length": len(data)}, f)
```

Submission is one call from the data scientist, and the data owner approves and processes. The README shows `do.jobs[0].approve()` followed by `do.process_approved_jobs(share_outputs_with_submitter=True)`, then a final sync on both sides before the data scientist reads `ds.jobs[-1].output_paths[0]`. Note the ordering: nothing runs until a human on the data owner side approves the specific job.

## Where PySyft v2 is the wrong tool

The honest limitation is stated in the project's own metadata. `pyproject.toml` classifies the package as Development Status 3 - Alpha, and the README's own banner describes 0.10+ as the successor to `syft-client`, with datasets and jobs moved out into `syft-rds`. Documentation written for PySyft 0.9 and earlier will not run against this. The README's instruction to pin `syft<0.10` exists because the two APIs are not compatible.

The transport is the second constraint. Google Drive is a file store, not a message bus. The offline-first design is a direct consequence: work is expressed as files that reconcile later. If your use case needs low-latency rounds of model updates across many clients, file synchronization is the wrong primitive, and the README does not claim otherwise.

Third, the approval model is manual by default. The background services package can auto-approve, but the quick start has a human calling `approve()` on each job. For a single collaboration that is a feature. For a pipeline that submits hundreds of jobs a day it is a queue with a person in it.

Finally, the README does not document rollback or how to revoke a dataset after it has synced to a peer. The permissions documentation lives in packages/syft-permissions, but the top-level README does not summarize what revocation does to already-synced copies.

## PySyft compared with Flower and other federated frameworks

The comparison people search for is PySyft versus Flower, and the difference is architectural rather than a matter of feature lists. Flower's model is a long-running server coordinating clients that each train locally and return model updates; the round is the unit of work. PySyft v2's unit of work is a submitted script plus a result file, moved over storage the organizations already use.

That has consequences in both directions. PySyft does not require the data owner to run a persistent coordination service, which is often the blocker in a two-organization collaboration where nobody wants to open a port. Flower, by contrast, is built for the case where many clients participate in repeated training rounds and the coordination overhead is worth paying.

If your actual need is differential privacy guarantees on released statistics, neither framing matches: PySyft's protection here comes from the approval step and from data not leaving the owner's machine, not from a noise mechanism. The repository does carry `syft-crypto-python` as a dependency, but the top-level README does not describe a differential privacy mechanism, so do not assume one from the topic tags.

## Licence, maintenance and the cost of tracking this API

PySyft is Apache-2.0, declared in both the README badge and `pyproject.toml`. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your own code. It also means the project places no restriction on how you deploy the client inside a commercial setting. This is not legal advice; read the LICENSE file in the repository root for the actual terms.

The last push to the repository was on 2026-09-10, so the codebase is being touched. The most recent tagged release, however, is v0.9.6b6 from 2025-04-13, and that tag predates the 0.10 line the README now documents. The practical reading: the dev branch moves, the release tags lag. If you install from PyPI with the README's version specifiers you are on a different artifact than someone building from `dev`.

The upgrade cost is concentrated in the 0.9 to 0.10 break. The README's own advice is to pin `syft<0.10` if you depend on the legacy API, which means any codebase written against the old surface is frozen there until someone ports it. The workspace layout in `pyproject.toml` lists a large set of member packages (`syft-job`, `syft-dataset`, `syft-bg`, `syft-permissions`, `syft-perms`, `syft-enclave`, `syft-restrict`, `syft-migration`, and others), so the install surface is a set of separately versioned packages rather than one monolith. Expect to track versions across several of them, and note that `syft` pins `syft-permissions` and `syft-perms` at exactly 0.1.15.

## Conclusion

Adopt PySyft if your data owner already works out of Google Drive and you want job submission with mock/private separation and an explicit approval step, and you can accept an API that the README itself calls alpha and that was restructured between 0.9 and 0.10. Do not adopt it if you need a mature federated training stack for many clients, or a transport other than a file-based one today. Before committing, verify three things: that your data owner can complete the Google Cloud OAuth setup described in docs/auth.md, that every dependency you need is reachable at the versions the README pins (syft>=0.10.0, syft-rds>=0.6.0, syft-bg>=0.3.12), and that the permissions model in packages/syft-permissions covers roles beyond the single DO/DS pair shown in the quick start.

## FAQ

### How do I install PySyft?

The README installs it with uv: `uv pip install "syft>=0.10.0" "syft-rds>=0.6.0"` for a data scientist, adding `"syft-bg>=0.3.12"` on the data owner side. Python 3.10 or newer is required.

### What is the difference between syft and syft-rds?

Since 0.10, `syft` is the sync engine and helper layer imported as `import syft as sy`, while datasets and jobs live in `syft-rds`, which provides `login_do` and `login_ds`. Code depending on the older PySyft API should pin `syft<0.10`.

### Does PySyft work with federated learning?

The repository carries federated-learning as a topic, but the README describes job submission against a data owner's private data with mock/private separation, not multi-client training rounds. The README does not document a federated averaging mechanism.

### How does a data scientist see the private data in a PySyft job?

They do not. The data scientist works against a mock file, and inside a submitted job `resolve_dataset_file_path` resolves to the private data instead. Only approved results are shared back, according to the README.

## Sources

- [License: Apache-2.0](https://github.com/OpenMined/PySyft/blob/dev/LICENSE)
- [OpenMined/PySyft on GitHub](https://github.com/OpenMined/PySyft)
- [Project website](https://www.openmined.org/)
- [README](https://github.com/OpenMined/PySyft/blob/dev/README.md)
- [Releases](https://github.com/OpenMined/PySyft/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/openmined-pysyft
