# frugally-deep: running Keras models in C++ without linking TensorFlow

> A header-only C++14 library that re-implements the prediction subset of TensorFlow, so a trained Keras model can run as a plain dependency in your binary. The trade-off is a conversion step, one CPU core per prediction, and a layer list that stops short of Lambda and stateful RNNs.

**Dobiasd/frugally-deep** — A lightweight header-only library for using Keras (TensorFlow) models in C++.

- Repository: https://github.com/Dobiasd/frugally-deep
- Stars: 1,127 · Forks: 238
- Language: C++
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dobiasd-frugally-deep

## The problem frugally-deep solves for C++ teams

Training a model in Keras and then shipping it inside a C++ application usually means linking the application against TensorFlow. The README describes frugally-deep as the alternative: it "re-implements a (small) subset of TensorFlow, i.e., the operations needed to support prediction", and the stated consequence is "a much smaller binary size than linking against TensorFlow". The intended user is a C++ developer who already has a model, already trusts Keras as the training tool, and only needs the forward pass at runtime. Python stays in the pipeline, not in the deployed artifact.

The library is header-only and written in modern C++. Its own dependencies are FunctionalPlus, Eigen and nlohmann/json, all header-only as well. That matters for build integration: there is no shared object to ship, no runtime library to match against a specific TensorFlow version, and no ABI question about which compiler built the framework. The cost of that choice appears later, in the conversion step and in the layer coverage.

## How the conversion and forward pass actually work

The workflow has three stages. In Python you build, compile, train and evaluate the model as usual, then save it to a single file with model.save('....keras'). The README requires the model's image_data_format to be channels_last, which is the TensorFlow backend default; models created with a different image_data_format or another backend are not supported. That is a hard gate, not a warning.

A script in the keras_export directory converts the saved model into the frugally-deep file format. During conversion, a test case is generated automatically: an input and its corresponding output values, saved alongside the model. On the C++ side, fdeep::load_model runs that test to confirm the forward pass in frugally-deep matches the result Keras produced. This is the most useful design decision in the project. A silent numerical mismatch between two implementations of the same operations is the failure mode that would otherwise surface in production, and here it surfaces at load time instead.

The C++ side then calls model.predict with fdeep tensors and prints or consumes the result. The README is explicit that this covers inference only, and that convolutions avoid materializing the im2col input matrix, which is where a large temporary allocation would otherwise appear. Supported architectures go beyond sequential stacks: the functional API, multiple inputs and outputs, nested models, residual connections, shared layers, variable input shapes and custom layers passed as factory functions to load_model are all listed as supported.

## Installing frugally-deep and running a first prediction

The README points to INSTALL.md for the different ways to install the library, so the repository itself is the place to follow. The tested toolchain is a C++14-compatible compiler (GCC 4.9, Clang 3.7 with libc++ 3.7, or Visual C++ 2015 and later), Python 3.9 or higher, TensorFlow 2.21.0 and Keras 3.14.0. The README notes these are the tested versions and that somewhat older ones might work, which is a caveat rather than a guarantee.

The full workflow starts with a small Keras model saved to a .keras file:

```python
import numpy as np
from keras.layers import Input, Dense
from keras.models import Model

inputs = Input(shape=(4,))
x = Dense(5, activation='relu')(inputs)
predictions = Dense(3, activation='softmax')(x)
model = Model(inputs=inputs, outputs=predictions)
model.compile(loss='categorical_crossentropy', optimizer='nadam')
model.fit(np.asarray([[1, 2, 3, 4], [2, 3, 4, 5]]),
          np.asarray([[1, 0, 0], [0, 0, 1]]), epochs=10)
model.save('keras_model.keras')
```

Then the conversion script turns that file into the frugally-deep format. Run it from the repository root, since the path keras_export/convert_model.py is relative to the checkout:

```bash
python3 keras_export/convert_model.py keras_model.keras fdeep_model.json
```

The output is fdeep_model.json, and it carries the generated test case with it. Loading that file in C++ runs the test before you ever call predict:

```cpp
#include <fdeep/fdeep.hpp>
int main()
{
    const auto model = fdeep::load_model("fdeep_model.json");
    const auto result = model.predict(
        {fdeep::tensor(fdeep::tensor_shape(static_cast<std::size_t>(4)),
        std::vector<float>{1, 2, 3, 4})});
    std::cout << fdeep::show_tensors(result) << std::endl;
}
```

If the test fails, load_model does not return a usable model, so a mismatch shows up on the first run rather than as a wrong number later. The README does not document what to do when the test fails, and it does not document rollback or versioning of the converted file format, so treat a failed load as a signal to check the model's layer list against the supported set.

## Where frugally-deep stops: unsupported layers and single-core inference

The README lists what is currently not supported, and the list is specific enough to plan around: Lambda (with a link to a FAQ entry explaining why), Hashing, HashedCrossing, MelSpectrogram, STFTSpectrogram, StringLookup, TextVectorization, stateful recurrent layers, and temporal models. A model that uses any of these cannot be converted, and there is no partial-conversion mode described. If your pipeline preprocesses text with TextVectorization inside the model graph, that preprocessing has to move outside the graph or the model has to be redesigned.

The second limitation is stated with unusual bluntness in the README: the library "utterly ignores even the most powerful GPU in your system and uses only one CPU core per prediction". A single prediction will not get faster because the machine has a GPU. The README's own answer is throughput rather than latency: you can run multiple predictions in parallel and use as many CPUs as you like to improve overall prediction throughput. That is a different engineering target. If your application needs one low-latency prediction per request, this design will not deliver it; if it needs to process a stream of inputs, parallelism across predictions is the intended path, and you are responsible for the threading.

There is also a numerical surface to consider. The library re-implements the operations itself, and the auto-generated test case checks one input and output pair. That is a sanity check, not a full validation of your model's behaviour across its input domain. The README does not claim otherwise.

## frugally-deep compared with ONNX Runtime and TensorFlow Lite

The obvious alternatives are ONNX Runtime and TensorFlow Lite, and the difference is where the model format comes from. ONNX Runtime consumes ONNX, so the workflow is Keras to ONNX to a runtime, with an intermediate format that is not Keras-specific. TensorFlow Lite consumes its own converted flatbuffer and brings a runtime library and a larger surface of tooling. frugally-deep consumes a JSON file produced by its own convert_model.py, and its selling point is that the C++ side is headers only, with no framework to link.

That narrows the comparison to build and deployment. If you want a vendor-supported runtime with a broad operator set and GPU delegates, frugally-deep is the wrong tool; the README makes no claim about GPU execution at all. If you want to drop a model into an existing C++14 codebase and keep the dependency list to three header-only libraries, the conversion script and the load-time test are the whole integration. The price is the supported-layer list and the absence of a training path in C++.

## Licence, maintenance and upgrade cost

frugally-deep is MIT licensed, and the dependencies it names (FunctionalPlus, Eigen, nlohmann/json) carry their own licences, which you should check separately for your distribution model. Nothing here is legal advice; the practical point is that a header-only library means the licence text travels with the headers you copy into your build.

Maintenance looks current: the last push to the default branch was on 2026-05-06, and releases v0.19.4, v0.19.5 and v0.20.0 landed in late April and early May 2026. The upgrade cost is tied to the conversion step rather than to the C++ headers. The README pins tested versions of TensorFlow (2.21.0) and Keras (3.14.0), so a Keras upgrade can change the saved model format or the layer set in ways that affect convert_model.py. Because the converted file embeds a test case and load_model checks it, a mismatch after an upgrade should appear at load time rather than silently. Budget for re-running the conversion and the load test whenever the training-side stack moves, not just when the C++ library does.

## Conclusion

Adopt frugally-deep if you train in Keras and need prediction inside a C++ binary where linking TensorFlow is too heavy, and your model uses layers from the supported list. Do not adopt it if you rely on Lambda layers, stateful recurrent layers, text vectorization layers, or GPU execution, and do not expect a training path in C++. Before committing, verify three things: that convert_model.py accepts your saved .keras file, that the auto-generated test case passes on load, and that the prediction throughput you get from running several predictions in parallel on separate cores meets your target.

## FAQ

### Can frugally-deep run a Keras model in C++ without TensorFlow?

Yes. The README states that the library re-implements a small subset of TensorFlow, specifically the operations needed to support prediction, and that its only dependencies are the header-only libraries FunctionalPlus, Eigen and nlohmann/json. The model is converted to a JSON file first and then loaded with fdeep::load_model.

### Which Keras layers does frugally-deep not support?

The README lists Lambda, Hashing, HashedCrossing, MelSpectrogram, STFTSpectrogram, StringLookup, TextVectorization, stateful recurrent layers and temporal models as currently not supported. A model using any of them cannot be converted.

### Does frugally-deep use the GPU or multiple CPU cores for a prediction?

The README says it uses only one CPU core per prediction and ignores the GPU. It notes that you can run multiple predictions in parallel to use more CPUs and improve overall throughput, but a single forward pass stays on one core.

### How do you convert a Keras model to the frugally-deep format?

Save the model from Keras with model.save('....keras'), then run keras_export/convert_model.py with the saved file and the output path as arguments. The README's example is python3 keras_export/convert_model.py keras_model.keras fdeep_model.json.

### How does frugally-deep check that its predictions match Keras?

The README states that convert_model.py generates a test case with input and corresponding output values and saves it along with the model, and that fdeep::load_model runs this test to confirm the forward pass matches Keras.

### What Python, TensorFlow and Keras versions does frugally-deep require?

The README lists Python 3.9 or higher, TensorFlow 2.21.0 and Keras 3.14.0 as the tested versions, and notes that somewhat older versions might work too. The C++ side needs a C++14-compatible compiler such as GCC 4.9, Clang 3.7 or Visual C++ 2015 and later.

## Sources

- [Dobiasd/frugally-deep on GitHub](https://github.com/Dobiasd/frugally-deep)
- [Issues](https://github.com/Dobiasd/frugally-deep/issues)
- [License: MIT](https://github.com/Dobiasd/frugally-deep/blob/master/LICENSE)
- [README](https://github.com/Dobiasd/frugally-deep/blob/master/README.md)
- [Releases](https://github.com/Dobiasd/frugally-deep/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dobiasd-frugally-deep
