Framework
secretflow/secretflow avatar
secretflow/secretflow

SecretFlow's four layers, its 2024 dependency pins, and a beta-only tag history

A unified framework for privacy-preserving data analysis and machine learning

2,714 stars468 forksPythonApache-2.0

At a glance

What is it?
SecretFlow splits privacy-preserving machine learning into device, flow, algorithm and workflow layers, and leaves the cryptography to SPU, HEU and YACL. What the repository states about itself is thin: no install command, no algorithm list, no benchmark numbers, and a dependency set whose oldest exact pins are two-year-old development builds.
Who is it for?
SecretFlow fits teams already committed to SPU and YACL, and it does not fit teams that need a stable tag or an arm64 install this year. Before adopting it, read the two x86_64 markers in pyproject.toml and the dev pins dated 2024, check what installation.md and the API Reference say for your specific case, and decide which of the three build toolchains at the repository root produces the artifact you intend to pin.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 164 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Plain devices and secret devices sit under one abstract layer

Everything above the algorithm layer in SecretFlow is written against a single abstraction: the device. Two kinds exist. Plain devices behave like ordinary local computation. Secret devices encapsulate the cryptographic protocols, so the same analysis code can run either over raw data on one machine or over data that never leaves its owner.

That split is the project's central design claim, and it explains the shape of the repository. The Python package under `secretflow/` does not implement the cryptography itself. It models higher algorithms as a flow of device objects and a DAG, then carries data processing, model training and hyperparameter tuning through that flow.

The consequence for anyone reading the source is that the protocol work is not in this repository at all. The device layer is an interface, and what sits behind it is named in a list of related projects rather than in the code tree. SPU is described as a provable, measurable secure computation device, HEU as a high-performance homomorphic encryption algorithm library, and YACL as the C++ library holding cryptography, network and io modules that other SecretFlow code depends on. Kuscia handles task orchestration on K3s, and SCQL covers joint analysis between parties that do not trust each other.

YACL is the sharpest example of the split between interface and implementation, since it is C++ while the tree described here is Python with its own `tests/` directory. A read of the device layer tells you what protocol the code expects to find; it cannot tell you what that protocol costs, because the measuring happens in a repository you have to install separately.

The algorithm layer covers two partitioning schemes in one line

The algorithm layer is where data analysis and machine learning happen on partitioned data, and the partitioning is named in a single phrase: horizontal or vertical. That phrase is the whole specification at this level. No algorithm is listed in the repository README, no table maps routines to partitioning modes, and nothing states what happens when a horizontal routine receives vertically partitioned tables.

The detail lives in documentation the README links out to rather than in the repository: a Getting Started page, a User Guide, an API Reference and a Tutorial, all on the project's documentation site. The API Reference is where to confirm which routines exist, because the package tree in this repository is where such an answer is generated from, not where it is summarized for a reader.

Horizontal and vertical partitioning are different problems with different communication costs, and offering both under one import path still leaves the caller to know which one each routine assumes. Anyone building on SecretFlow should read the API Reference for the specific estimator or query they need rather than trusting the one-line description at the top of the README, and should confirm from the package source whether a routine needs both parties holding aligned row sets or complementary column sets.

The dependency list hints at what this layer is expected to talk to. `sqlglot==25.5.1` parses SQL, `duckdb==1.0.0` executes it locally, `pyarrow==14.0.2` moves columnar data between partitions, and `s3fs==2024.2.0` with `aiobotocore==2.17.0` reach object storage. Each of those sits in the mandatory dependency block rather than an extra, so SQL parsing, a columnar engine and object storage are part of every install whether the routine you call needs them or not.

The workflow layer promises tuning, and the repository ships no configuration for it

The top of the architecture is a workflow layer that integrates data processing, model training and hyperparameter tuning. In this repository that promise is one sentence long. There is no YAML, no TOML, no Python entry script and no command line shown for wiring those three steps together. There is no installation command either. Both the Install and Deployment sections are pointers, one to `docs/getting_started/installation.md` and the other to `docs/getting_started/deployment.md`.

Two example directories sit under `examples/`. `examples/benchmark/` backs the Benchmarks section, which is itself a pointer to `docs/developer/benchmark/overall_benchmark.md`; the README prints no timing, no hardware description and no comparison figure of its own. `examples/my_plugin/` is the more telling one, since a plugin example implies an extension point for writing your own device or algorithm. Nothing in the README says what interface a plugin must implement, where it gets registered, or how it is loaded at runtime.

So the two things a new user needs first, the install path and the extension path, both sit outside the repository README, while a `docker/` directory is present at the root with no prose describing what any image in it is built for.

Learning material follows the same route out. The README's PETs section points at `docs/awesome-pets/awesome-pets.md`, described as a curated list of papers and tutorials on privacy-enhancing technologies. So the repository ships a reading list, a benchmark pointer and a documentation site, and the wiring instructions for its own four layers live in a file the README never shows.

Two packages install only on x86_64

Two entries in the dependency list carry an architecture marker, and they are the SPU API and SDK packages:

toml
    "sdc-apis==0.1.0.dev240320; platform_machine == 'x86_64'",
    "sdc-sdk==0.1.0.dev240320; platform_machine == 'x86_64'",

pip evaluates `platform_machine == 'x86_64'` while resolving the wheel, so on any other architecture those two distributions are simply not installed by this metadata. That covers Apple Silicon workstations, which are a normal machine for data work, and ARM servers. What takes their place is stated nowhere in the repository: there is no alternative marker, no conditional extra, and no note about a fallback device implementation. The rest of the cryptography still arrives, because SPU itself is pinned separately.

The numeric stack is equally narrow, and narrow in a way that fixes the era of the surrounding tooling. `numpy` is capped below 2.0, `pandas` is pinned to exactly 1.5.3, `pyarrow` to 14.0.2 and `duckdb` to 1.0.0. JAX and jaxlib stop at 0.4.26, with an inline comment giving the reason as a fix to a jnp.select numerical problem. A resolved environment here is a 2024 shaped environment whenever the install runs.

One detail makes the marker easier to miss. The condition sits on the dependency line rather than on an extra named per architecture, so an install on a non-x86_64 machine resolves and finishes with fewer packages instead of stopping with an error. Nothing announces the omission at install time either, since the install log on that machine simply has no `sdc-apis` entry to point at.

The oldest exact pins are dev builds dated March 2024

The dependency list pins by date in a way worth reading closely. `sdc-apis` and `sdc-sdk` are both `0.1.0.dev240320`, a development build from 2024-03-20, held at that exact string with no upper bound and no stable alternative. Alongside them sit more dated dev builds: `spu==0.9.4.dev20250618`, `sf-sml==0.1.0.dev20250623`, `sf-heu==0.6.0.dev20250514` and `secretflow_serving_lib==0.10.0.dev20250414`.

Then there is the pre-release internal stack. `kuscia==0.0.3b0`, `secretflow-dataproxy==0.5.0b0` and `secretflow-spec==1.1.0b0` are all beta pins, so the framework's own runtime components, meaning the orchestration layer, the data proxy and the spec package, are declared against releases that were never marked stable.

The distance between those pins and the repository's own history is the practical point. The last recorded push to the default branch is 2026-04-24, more than two years past the 2024-03-20 build that two dependencies are frozen on. Nothing in `pyproject.toml` says whether those two pins are still needed for compatibility with SPU, or whether they have simply not been revisited since the day they were written.

The dated pins also run on more than one clock. `sf-heu` and `sf-sml` carry May and June 2025 dates, `spu` a June 2025 date, and the serving library an April 2025 date, while the two `sdc-` packages sit two years earlier at the same version. A dependency set assembled from four different dates is harder to reason about than a single freeze, and no comment in the file groups them.

The dev extra mixes 2025 pins with a much older xgboost

The optional `dev` extra shows the same exact-pin habit, with one loose name per line. `pytest==8.4.1`, `build==1.2.2`, `statsmodels==0.13.2`, `polars==0.19.3`, `filelock==3.18.0` and `liac-arff==2.5.0` are all fixed; `matplotlib` and `psutil` are bare names, and so is `tweedie`. Then `xgboost==1.7.5` stands out. It is the oldest exact pin in the file, several releases behind the 2024 and 2025 pins around it, and a resolved dev environment holds that version whatever else moves.

The `doc` extra follows the same pattern with looser bounds: `autodocsumm~=0.2.14`, `linkify-it-py~=2.0`, `matplotlib~=3.10`, `mdutils==1.6`, `myst-nb~=1.2` and `secretflow-doctools~=0.8.5`. That list ends inside the extra, right after `secretflow-doctools~=0.8.5`, without a closing bracket, so any extra declared after `doc` is not visible in `pyproject.toml` as the file stands. The partial list is the whole of the documentation toolchain a reader can confirm from this repository.

The two loose names, `matplotlib` and `psutil`, are the entries free to move between installs, and `tweedie` is a third. Everything else in the extra is fixed to a patch version, which is the sensible choice for a group whose job is to produce comparable numbers.

Three build toolchains share the root directory

The repository root holds artifacts from three build toolchains. Packaging runs through `pdm-backend`, named in the build-system table, with a `pdm_build.py` hook beside the metadata because `[project]` sets `dynamic = ["version"]`. Locking is handled by uv, since `uv.lock` sits at the root. A Bazel-style `WORKSPACE` file also sits at the root, next to a `pytest.ini` for test discovery and a `.pylintrc` for linting.

Around them sits the housekeeping set: `.licenserc.yaml` for header checks, `.pre-commit-config.yaml`, `.coveragerc` for coverage settings, `.gitmodules`, and a `docker/` directory. Documentation is built as a site through `.readthedocs.yaml` and lives under `docs/`, with named files such as `REPO_LAYOUT.md`, `VERSIONING.md`, `SECURITY.md`, `LEGAL.md` and `CHANGELOG.md`, plus a Chinese `README.zh-CN.md` linked from the English one.

Nothing in the README says which of those toolchains produces the artifact you install. The version string is not written in `pyproject.toml` at all, and with `pdm_build.py` at the root and a `WORKSPACE` file beside it, a version seen inside an installed environment has more than one possible origin. Teams pinning SecretFlow internally should settle that question first and record the answer.

A `.gitmodules` file adds a fourth wrinkle to reproducing any of it, since a submodule arrives as a checkout rather than as a wheel and can point at a commit that no longer exists. Whatever lockfile you settle on, the submodule commit is the part a lockfile cannot describe for you.

Every recent tag ends in 0b0, and the disclaimer bans non-releases

The three most recent releases are `v1.12.0b0` on 2025-04-09, `v1.13.0b0` on 2025-07-04 and `v1.14.0b0` on 2025-09-26. Each is the zeroth beta of its minor line, and no stable tag appears among them. The last recorded push to `main` is 2026-04-24, close to seven months after the newest tag, so the tag history is not a full picture of what sits on the default branch.

The disclaimer at the end of the README reads oddly against that list. Non-release versions of SecretFlow are prohibited from use in any production environment, for reasons including bugs, glitches, missing functionality and security issues. The tags a new user is pointed at are all beta builds, and the version itself is dynamic, so working out which artifacts count as a release is left to the reader rather than settled by the repository.

Two further root facts bear on that decision. The package metadata declares Apache-2.0 and the repository carries both a LICENSE file and a LEGAL.md, so licensing is settled. Maintenance is not something this README claims: no line in it says the project is actively developed, and the dated push record is the only evidence on offer.

The repository's own counters point the same way. It lists 2715 stars, 469 forks and 92 open issues, and none of those three numbers is broken down by release, so none of them says which of the three tags is the one people are running.

Editorial conclusion

SecretFlow fits teams already committed to SPU and YACL, and it does not fit teams that need a stable tag or an arm64 install this year. Before adopting it, read the two x86_64 markers in pyproject.toml and the dev pins dated 2024, check what installation.md and the API Reference say for your specific case, and decide which of the three build toolchains at the repository root produces the artifact you intend to pin.

Frequently asked questions

What is SecretFlow actually built for?

It is described as a unified framework for privacy-preserving data intelligence and machine learning, whose algorithm layer works on horizontally or vertically partitioned data.

Which Python versions does SecretFlow support?

pyproject.toml sets requires-python to >=3.10,<3.12, and the classifiers list only Programming Language :: Python :: 3.10 and 3.11.

Does SecretFlow install on Apple Silicon or other non-x86_64 machines?

The sdc-apis and sdc-sdk entries carry the marker platform_machine == 'x86_64', so pip skips both everywhere else, and no fallback device is named in the repository.

Where does SecretFlow keep its cryptography?

The device layer delegates to secret devices, and the primitives live in sibling repositories: SPU for secure computation, HEU for homomorphic encryption, and YACL for cryptography, network and io.

Does SecretFlow have stable releases to pin?

The three most recent tags, v1.12.0b0, v1.13.0b0 and v1.14.0b0, are all beta builds, and the README disclaimer bars non-release versions from production use.

Where does SecretFlow document installation and deployment?

The README holds no install command of its own and points to docs/getting_started/installation.md and docs/getting_started/deployment.md.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. secretflow/secretflow on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/secretflow-secretflow.svg)](https://hysenlabs.com/projects/secretflow-secretflow)