# google-research/morph-net: shrinking convolutional networks with resource regularizers

> MorphNet learns which filters to remove during training by adding a resource regularizer to the loss. It is for TensorFlow users with a working seed network and a fixed FLOPs or latency budget.

**google-research/morph-net** — Fast & Simple Resource-Constrained Learning of Deep Network Structure

- Repository: https://github.com/google-research/morph-net
- Stars: 1,039 · Forks: 151
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-research-morph-net

## The problem morph-net addresses: a seed network that is bigger than its budget

Most teams do not start from an empty architecture. They start from a convolutional network that trains and reaches acceptable accuracy, then discover it is too large for the device or the latency target they have to meet. The usual response is manual channel tuning: pick a layer, halve its filters, retrain, measure, repeat. MorphNet attacks that loop differently. The README describes it as a method for learning deep network structure during training, where a regularizer pushes the influence of filters down and, once they are small enough, the corresponding output channels are marked for removal. The output is a proposal, not a new topology. The README states plainly that MorphNet does not change the topology of the network; the proposed model has the same number of layers and connectivity pattern as the seed network. That constraint is the whole design. If your problem is really about connectivity (skip connections, kernel sizes, block counts), this library is the wrong instrument. It only adjusts the number of output channels in each convolution layer.

## How the regularizer crawls the graph and marks channels

The mechanism is a constrained optimization. You add a regularization term to the training loss and minimize the sum with stochastic gradient descent or a similar optimizer, so the structure becomes a constraint rather than a post-hoc decision. The regularizer is initialized with a threshold and the output boundary ops of your model, optionally the input boundary ops too. According to the README, the regularizer crawls the graph starting from the output boundary and applies regularization to some of the ops it encounters; when it reaches any input boundary op it stops, so ops inside the input boundary are left alone. The threshold decides which output channels can be eliminated. Two calls matter and the README warns against confusing them: get_regularization_term() is the loss you add to training, while get_cost() is the estimated cost of the network if the proposed structure is applied. The README also notes that the regularization loss is not currently added to the tf.GraphKeys.REGULARIZATION_LOSSES collection, a choice it says is subject to revision, so you must wire it into your loss yourself. Regularizer choice is not a hyperparameter in the usual sense. It is determined by your target cost and by whether you can add layers: LogisticSigmoid (FiGS) is the recommended path and requires gating layers, Gamma works when the seed network has BatchNorm with scale enabled, and GroupLasso is for models without BatchNorm and is marked deprecated.

## Installing morph-net and running a first structure-learning pass

The repository ships a setup.py with version 0.2.1, package name morph_net, and no published release artifacts in the retrieved metadata, so the practical route is installing from the checkout. The package metadata declares Python 2 and 3 classifiers, but the surrounding code is TensorFlow, and the examples live under examples/keras/ and examples/slim/, so pick the example that matches your stack. Clone the repository and install it in editable mode from the top level, where setup.py sits next to morph_net/ and examples/.

## The two-round workflow is the real cost of admission

MorphNet splits training into structure learning and retraining, and that split is where most of the effort goes. The first round trains with the regularizer; the second round retrains the resulting model from scratch with normal hyperparameters. The README recommends a fixed learning rate, no decay, for the structure-learning round, which is a deviation from what most training recipes do and easy to get wrong if you reuse an existing configuration. There is no published rule for when to stop the first round. The README says there are no specific guidelines on the stopping time, only that you would likely wait for the regularization loss to stabilize. That leaves the most consequential decision, when the proposed structure is good enough, to your own monitoring. The exported JSON is also a moving target: the README notes that as training progresses the proposed model structure changes, so a file pulled at the wrong moment encodes a different network than the one you intended. Budget for the second training run as a full run, not a fine-tune, because that is what the workflow specifies.

## Where morph-net is the wrong tool

Three cases stand out. First, topology changes. If your cost problem comes from a layer that should be removed entirely, a kernel that should shrink, or a block that should be replaced, morph-net cannot express that; it preserves layer count and connectivity by design. Second, non-TensorFlow stacks. The regularizers operate on a TensorFlow graph and the examples are Keras and Slim, so a PyTorch model has no path here. Third, models that cannot host gating layers. The recommended LogisticSigmoid regularizer requires adding a gating op after each layer you want to prune, and if you cannot modify the graph, you fall back to Gamma (which requires BatchNorm with scale enabled) or GroupLasso (deprecated). The README also warns that if you use BatchNorm you must set scale=True on tf.keras.layers.BatchNormalization; a model built without that flag will not give the Gamma regularizer what it needs, and the failure is silent at the API level. There is also a hard boundary on what gets touched: ops inside the input boundary are not regularized, so layers you place behind that boundary will never be candidates for removal.

## How morph-net differs from post-training pruning toolkits

The closest alternative in practice is magnitude-based pruning, the kind of workflow where you train to convergence, rank weights or channels by a norm, delete a percentage, and fine-tune. The difference is where the decision is made. Post-training pruning decides after the fact using a fixed heuristic and a global sparsity target you set by hand. MorphNet decides during training, with the target expressed as a resource cost (FLOPs or model size) and the removal driven by a regularizer rather than a threshold on weights. That means the accuracy and the cost are negotiated in the same optimization, and the README frames it exactly that way, as a constrained optimization of structure under the constraint represented by the regularizer. The trade-off is control. A pruning toolkit lets you say "remove 30 percent of channels in this layer" and see the result immediately. MorphNet gives you a regularization strength and an alive threshold, and you search. If you need per-layer guarantees, the pruning approach is more direct. If you need the cost target to drive the search, and you can afford two training rounds, morph-net is the more principled fit.

## Maintenance, licensing and what to check before you commit

The repository is not archived and its last push was on 2026-07-02, so it is not abandoned, but the retrieved metadata shows no releases, and setup.py still lists Python 2 among its classifiers, which is a signal about how much of the packaging has been revisited. The licence is Apache-2.0, declared both in the LICENSE file at the top level and in the setup.py classifier. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve notices and state changes. That is the general shape of the licence, not advice about your situation, and if you plan to redistribute a modified morph-net inside a product, read the LICENSE text and the CONTRIBUTING.md in the repository rather than relying on a summary. On upgrade cost: with no releases to pin against, you are tracking the master branch, and the README's own note that the regularization-loss collection behaviour is subject to revision is the kind of change that can move under you. Before adopting, confirm the TensorFlow version your seed network needs actually works with the code in morph_net/, since no compatibility matrix is published.

## Conclusion

Adopt morph-net if you already have a trained TensorFlow seed network and a cost target to hit, and you are willing to run two training rounds plus a topology-preserving rewrite of your model. Skip it if you need to change connectivity, if you cannot add gating layers or enable BatchNorm scale, or if you are not on TensorFlow. Before committing, verify that your seed network's BatchNorm layers were built with scale=True, that you can produce the StructureExporter JSON, and that you have budget for the retraining pass, because morph-net does not ship a rollback path for a structure you dislike.

## FAQ

### Does morph-net change the architecture of my model, or only the number of channels?

It only changes the number of output channels in each convolution layer. The README states that MorphNet does not change the topology of the network, so the proposed model keeps the same number of layers and the same connectivity pattern as the seed network.

### Which regularizer should I use in morph-net, and do I need to modify my model?

The README recommends the LogisticSigmoid regularizer (FiGS), but it requires adding a probabilistic gating operation after any layer you want to prune. If you cannot add layers, use Gamma when the seed network has BatchNorm with scale enabled, or GroupLasso otherwise, which the README marks as deprecated.

### How many training rounds does the morph-net workflow require?

Two. The README calls the first round structure learning, where the regularizer is added to the loss, and the second round retraining, where the exported structure is rebuilt and trained from scratch without the regularizer using standard hyperparameter values.

### What values should I try for the morph-net regularization strength?

The README recommends searching along a logarithmic scale spanning a few orders of magnitude around 1/(initial cost). Its example is a seed network starting at 1e9 FLOPs, where you would explore regularization strength around 1e-9.

## Sources

- [google-research/morph-net on GitHub](https://github.com/google-research/morph-net)
- [Issues](https://github.com/google-research/morph-net/issues)
- [License: Apache-2.0](https://github.com/google-research/morph-net/blob/master/LICENSE)
- [README](https://github.com/google-research/morph-net/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-research-morph-net
