# fklearn is a functional machine learning library that its own owner has deprecated

> Nubank's chapter-owned Python library, now marked deprecated with limited support. The readme is 231 words; the real detail sits in pyproject.toml, where eight extras, a Python range capped below 3.15, and one exact MarkupSafe pin do the talking.

**nubank/fklearn** — fklearn: Functional Machine Learning

- Repository: https://github.com/nubank/fklearn
- Stars: 1,552 · Forks: 176
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nubank-fklearn

## The maintenance notice sits above the four principles

fklearn is a Python library for machine learning written in a functional style, and the first thing its readme tells you is that it is deprecated. Support is limited, the library is chapter-owned, and no specific team maintains it. That is an unusual thing to find at the top of a project with 1552 stars and 175 forks, and it changes how the rest of the document should be read: the four principles that follow are a statement of intent by a team that has handed the work over rather than a promise of upkeep.

The timing is precise. The default branch is `master`, the last commit is dated 2026-09-24, and the newest release, 4.3.0, was published on 2026-09-24 as well, so the final commit of record is the release that carries the deprecation. Two earlier tags sit behind it: 4.2.3 from 2026-06-10 and 4.2.2 from 2026-04-27, the only one of the three with a release name attached, Workflow Feature.

The license is Apache-2.0 with a LICENSE file in the root. The documentation is hosted on readthedocs with separate pages for getting started, the API reference and contributing, and the community link points at a Gitter room rather than a forum.

## Four principles written as acceptance criteria

The readme states its purpose in one line, that functional programming principles should make real machine learning problems easier to solve, and names scikit-learn as the reference point for the project's name. The four principles that follow are not aspirations, they are checks a change is supposed to pass:

Validation should reflect real life situations. Production models should match validated models. Models should be production ready with few extra steps. Reproducibility and in depth analysis of model results should be easy to achieve.

Each one is a property of the pipeline shape rather than of a model, and each maps onto something a code reviewer could object to: a validation split that does not look like production traffic, a fitted artifact that differs from the object in the notebook, a handoff that needs a week of glue code. The dependency list backs the style without explaining it. `toolz` is a runtime requirement, which is the dictionary of small composable functions, and the rest of the core is a conventional set: joblib, numpy, pandas, scikit-learn, statsmodels and tqdm. What the API actually looks like is not in the repository; that lives in the API reference on the documentation site.

## The readme shows five extras, the manifest defines eight

Installation is a single line, `pip install fklearn`, and everything else is optional. The readme offers five extras, each one naming the backend it pulls in:

```
pip install fklearn[lgbm]       # LightGBM support
pip install fklearn[xgboost]    # XGBoost support
pip install fklearn[catboost]   # CatBoost support
pip install fklearn[all_models] # All model backends
pip install fklearn[all]        # All models + tools
```

The project configuration defines more than the readme shows. Besides the three single backend extras there are tools, demos and docs, and the two aggregate extras are written as references to the package's own extras: all_models expands to lgbm, xgboost and catboost, and all expands to all_models plus tools. So a self referential expression in the manifest is what the two catch all install lines actually resolve.

The extras a user would reach for are thin. LightGBM is bounded to 4 through 5, XGBoost to 2 through 3, CatBoost to 1.2.2 through 2. In the tools extra, shap is bounded to 0.43 through 0.49 and swifter to 0.24 through 2, which is a pair of version ceilings rather than pins. Under the hood the model backend is not part of the core install at all: scikit-learn is the only estimator library in the default dependency set.

## The docs extra is held years behind the rest of the manifest

The `docs` extra is the outlier in the dependency list. Everything else in the manifest is written as a modern range, with numpy and pandas both capped below 3, scikit-learn between 1.5 and 1.8, and the whole package restricted to Python 3.10 and newer with an upper bound below 3.15. The docs extra goes the other way: Sphinx below 6, nbsphinx from 0.4.2, sphinx-rtd-theme from 0.4.3, jinja2 below 3, and markupsafe pinned to the single version 2.0.1.

An exact pin on a transitive library is the kind of thing a security scanner flags and a maintainer defends, and here it sits next to a jinja2 ceiling that is also several major versions behind. Nothing in the repository explains which of those bounds is a compatibility floor rather than a preference.

The rest of the tooling is current. ruff is the formatter and the linter, the dependency group carries pytest, pytest-cov, pytest-xdist, mypy and hypothesis, and the group is installed by default, which the configuration records as the default dependency group. Coverage is configured from its own file, and the documentation build is wired through a readthedocs configuration file, so the version ceilings in the docs extra are what that site actually builds against.

## uv runs the development loop while setuptools builds the package

Two build systems share the repository. The project is built by setuptools, with the build backend declared as setuptools.build_meta, and its version 4.3.0 is written directly into the project table. Development, however, is driven entirely by uv, which is what the readme tells contributors to use:

```bash
uv sync                 # core deps + dev group
uv sync --all-extras    # also installs lgbm / xgboost / catboost / tools / demos / docs
```

A plain `uv sync` is enough for most work, because the configuration sets the default dependency group to `dev`, so pytest, ruff, mypy and hypothesis arrive without asking. The environment installs the library itself in editable mode, which is uv's default, so edits under `src/` take effect without a reinstall. Tests run with coverage over that source directory, and the two lint commands are a ruff check and a ruff format check over `src/` and `tests/`.

A lockfile is committed at the root, and adding to it is a separate command from adding a dependency: `uv add` for runtime, `uv add --dev` for the group. The Python range is 3.10 up to but not including 3.15, and the classifier list stops at 3.14, so the newest interpreter is excluded by design rather than by an oversight.

## One lockfile rule exists to keep internal credentials out of git

The most specific instruction in the readme is addressed to a single audience, contributors inside Nubank, and it is about the lockfile rather than about code. A lockfile written on a company machine can record the private index the machine was configured to use, including credentials, so the rule is to regenerate it with the configuration turned off and the index pointed at public PyPI:

```bash
uv lock --no-config --default-index https://pypi.org/simple/
```

`--no-config` is what excludes local index settings such as the internal registry, and the check that follows is manual: before committing, look at the package source URLs in the lockfile and confirm they point at public PyPI and carry no credentials. That is a concrete leak guard sitting in a readme, not a general warning.

The rest of the repository is conventional and fairly complete: tests and their configuration, a scripts directory, a docs directory, a notebooks directory, MANIFEST.in for what goes into the distribution, a Brewfile for macOS tooling, a coverage configuration, a changelog, a code of conduct and a contributing guide.

## The repository's main language is a notebook

The language statistics for this repository report Jupyter Notebook rather than Python, which means the notebooks directory outweighs the source directory by line count even though every line that runs comes from `src/`. For a library that exists to be imported, that is a telling ratio: the explanation is weighted more heavily than the implementation.

The same imbalance shows up in what is versioned. The history is a changelog file plus a changelog directory, the documentation site is built from a docs directory with its own readthedocs configuration, and the four principles that define the library are stated once, in prose, at the top of the readme rather than enforced anywhere in the tooling. The tooling is thorough about format, types and tests and silent about the principles.

A deprecated project is also a stable one to pin. Version 4.3.0 matches the last commit date exactly, the version lives in one place in the project table, and the release history is short enough to read in full: three tags across five months, one of them named.

## Conclusion

fklearn is fine to read, cite and learn from, and the deprecated label is the single fact that should decide whether you install it. Version 4.3.0 is the last one carrying that notice, so a project starting today would begin on a library its owner will not maintain, with 42 open issues and no named team behind the answers. If you do adopt it for something that matters, pin the version, expect no security or dependency updates, and plan the migration now rather than after the first breaking change. If you only want the ideas, the four principles in the readme and the composition style they describe travel well to any scikit-learn code base.

## FAQ

### Is fklearn still maintained?

No. The readme carries a maintenance notice that fklearn is deprecated and support is limited, and that the library is chapter owned with no specific team maintaining it. The last commit and the 4.3.0 release share the date 2026-09-24.

### What does installing fklearn pull in by default?

The core install is joblib, numpy, pandas, scikit-learn, statsmodels, toolz and tqdm, plus the build backend. Model backends are separate extras for LightGBM, XGBoost and CatBoost, and the docs, tools and demos groups are extras too.

### Which Python versions does fklearn support?

The project requires Python 3.10 and higher, with an upper bound below 3.15, and the classifiers list runs from 3.10 to 3.14. The default dependency group installs ruff, pytest, mypy and hypothesis, and it is installed by a plain uv sync.

### Why do fklearn contributors regenerate the lockfile with extra flags?

The command is uv lock with --no-config and the public PyPI index as the default, because --no-config excludes local index settings such as Nubank's internal registry. The readme then asks that package source URLs in the lockfile be checked for public PyPI addresses and for credentials before committing.

## Sources

- [Issues](https://github.com/nubank/fklearn/issues)
- [License: Apache-2.0](https://github.com/nubank/fklearn/blob/master/LICENSE)
- [nubank/fklearn on GitHub](https://github.com/nubank/fklearn)
- [README](https://github.com/nubank/fklearn/blob/master/README.md)
- [Releases](https://github.com/nubank/fklearn/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nubank-fklearn
