Open-source project
Kaixhin/nninit avatar
Kaixhin/nninit

nninit: weight initialisation for Torch7 modules, picked by an accessor and an initialiser

Weight initialisation schemes for Torch7 neural network modules

100 stars13 forksLuaMIT

At a glance

What is it?
nninit is a single-file Lua library that adds one method to Torch7's nn.Module: module:init(accessor, initialiser, ...). The accessor says which tensor to touch, the initialiser says how, and the chain of schemes runs from constant fills to xavier, kaiming, orthogonal, sparse, and convolution-aware init. It was last pushed on 2017-06-21.
Who is it for?
nninit fits a Torch7 project that initialises weights through nn and nngraph and wants the choice of scheme to be a data value rather than hand-written tensor code, particularly for the schemes that pick gains by nonlinearity. Skip it if you are on a framework other than Torch7, since the whole surface is a Torch module method, and skip it if you need schemes it does not implement.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Probably not. The repository last received commits 113 months ago, on June 21, 2017.
What is it written in?
Mainly Lua, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One method is the whole API: module:init(accessor, initialiser, ...)

The library's entire interface is a single method added to nn.Module:

lua
module:init(accessor, initialiser, ...)

The three parts divide the work cleanly. The accessor extracts the tensor to be initialised from the module. The initialiser is a function that takes the module, the tensor, and further options, adjusts the tensor, and returns the module, which is why init calls can be chained. The trailing ... is simply additional arguments for that initialiser. Installation is a rocks command:

sh
luarocks install nninit

Because initialisers return the module rather than a boolean or nothing, a whole layer's weights, its biases, and any subtensor can be initialised in one expression. That chaining is the design decision worth internalising before choosing a scheme.

Three accessor shapes, and the middle one indexes into a tensor

The accessor can be a string, a table, or a function, and the README gives an example of each. A string accesses the tensor as a property of the module:

lua
module:init('weight', nninit.constant, 1)

A table is the one that does real work. The tensor is first accessed as a property from the first element, and a subtensor is then extracted using Torch's indexing operator applied to the second element:

lua
module:init({'weight', {{1, 5}, {}}}, nninit.uniform, -1, 1)

A function must return the tensor as the result of being applied to the module:

lua
module:init(function(m) return m.weight:narrow(1, 1, 10) end, nninit.normal, 0, 0.01)

So arbitrary indexing is where the library earns its name: you can initialise a whole parameter or five rows of it.

Most initialisers come in pairs, one that fills and one that adds

Read the initialiser list as a family rather than a catalogue. nninit.copy copies the init tensor to the tensor being initialised. nninit.constant fills the tensor with a constant val, and nninit.addConstant adds the constant val to the current tensor, while nninit.mulConstant multiplies the current tensor by it. The same pairing repeats for the distributions: nninit.normal fills the tensor from N(mean, stdv) and nninit.addNormal adds to the current tensor from the same distribution, and nninit.uniform fills from U(a, b) with nninit.addUniform adding to it. The distinction is the operation, not the distribution, and it matters when you are initialising an existing tensor to a target state rather than starting from empty.

nninit.eye changes what it fills based on the module type

One initialiser is explicitly limited, and the limitation is documented in bold: nninit.eye only supports the module weights as the tensor, and it relies on the module type to determine the appropriate identity. What that identity is depends on what you are initialising. For linear layers and lookup tables, weights are filled with the identity matrix. For convolutional layers, filters are filled with the Dirac delta function, and the result is normalised by the number of input layers. So a single name covers two different mathematical objects, and the library decides which one you meant by inspecting the module. That is convenient for convolutional architectures, where starting from an identity-like filter bank is a deliberate choice, and wrong to rely on if you want uniform behaviour across module types.

xavier and kaiming differ in one term and in their default distribution

The two most used schemes are distinguished by a single term in the standard deviation and by which distribution they reach for. nninit.xavier fills the tensor with stdv = gain * sqrt(2 / (fanIn + fanOut)), uses the uniform distribution by default, and is also known as Glorot initialisation. nninit.kaiming fills it with stdv = gain * sqrt(1 / fanIn), uses the normal distribution by default, and is also known as He initialisation. Both accept optional named parameters dist and gain passed in via a table, and both come with their references attached in the documentation. One practical detail applies to kaiming specifically: the scheme typically includes the gain for ReLU units, which has to be specified manually in nninit.kaiming with the option {gain = 'relu'}. The default is not inferred from your module.

Gains are a lookup keyed by the nonlinearity that follows

The gain table is where the library encodes knowledge about activations, and it is worth reading as a whole. Gains can be calculated depending on the succeeding nonlinearity. If gain is a number it is used directly; if it is a string the mapping is applied. The entries are 'linear' at 1, 'sigmoid' at 1, 'tanh' at 5 / 3, 'relu' at sqrt(2), and 'lrelu' at sqrt(2 / (1 + leakiness^2)). Where applicable, gains default to 1. The leaky case is the only one that takes an extra parameter, and the documentation shows how to pass it, with the string first in the table and the named parameter alongside:

lua
module:init('weight', nninit.kaiming, {gain = {'lrelu', leakiness = 0.3}})

Two distributions, and three schemes that refuse certain tensors

The distribution choice is small: the two types supported are 'normal' and 'uniform', which is what the optional dist parameter on xavier and kaiming selects. The restrictions are more interesting. nninit.orthogonal only supports tensors with at least 2 dimensions, and fills the tensor with a normally distributed random orthogonal matrix, with an optional gain. nninit.sparse takes a sparsity value between 0 and 1 and sets (1 - sparsity) percent of the tensor to 0, so a sparsity of 0.2 drops out 80% of the tensor. nninit.convolutionAware only supports 2D convolutions with a symmetric filter size, fills the convolution tensor with matrices that are orthogonal in the frequency space, and takes an optional std that specifies the noise to break symmetry in the inverse Fourier transform.

One Lua file and a rocks directory, and the module list it covers

The repository is small in the way a focused library should be. The top level holds nninit.lua, the rocks directory, the README, a LICENSE file, and .gitignore, so the implementation is a single file. What it applies to is stated in the opening paragraph: it works with nn, and therefore nngraph, and allows arbitrary indexing of weights, biases, and parameters. The supported modules are nn.Linear and nn.LinearNoBias, nn.LookupTable, nn.TemporalConvolution, nn.SpatialConvolution and cudnn.SpatialConvolution, and nn.VolumetricConvolution and cudnn.VolumetricConvolution. The published example requires nn, cunn, cudnn, rnn, and nninit together, which shows the intended pairing with recurrent and GPU modules. The licence is MIT, and the last push to master is dated 2017-06-21.

Editorial conclusion

nninit fits a Torch7 project that initialises weights through nn and nngraph and wants the choice of scheme to be a data value rather than hand-written tensor code, particularly for the schemes that pick gains by nonlinearity. Skip it if you are on a framework other than Torch7, since the whole surface is a Torch module method, and skip it if you need schemes it does not implement. Before you adopt it, note the three narrowing constraints in its own documentation: eye and convolutionAware are restricted to particular module types and filter shapes, orthogonal needs at least two dimensions, and for ReLU-family gains you must pass the gain yourself with {gain = 'relu'}. The last push to master is dated 2017-06-21 and there are no tagged releases.

Frequently asked questions

What is nninit used for?

nninit provides parameter initialisation schemes for Torch7 neural network modules. It works with nn and therefore nngraph, and it allows arbitrary indexing of weights, biases, and parameters.

How do I install nninit?

Install it from LuaRocks with luarocks install nninit. The repository ships nninit.lua plus a rocks directory, and the published example requires nn, cunn, cudnn, rnn, and nninit.

Which Torch7 modules does nninit support?

nn.Linear and nn.LinearNoBias, nn.LookupTable, nn.TemporalConvolution, nn.SpatialConvolution and cudnn.SpatialConvolution, and nn.VolumetricConvolution and cudnn.VolumetricConvolution.

How do I set the gain for leaky ReLU in nninit?

Pass a table whose first element is the gain string and whose named element carries the extra parameter, for example {gain = {'lrelu', leakiness = 0.3}}. The documented mapping is 'linear' 1, 'sigmoid' 1, 'tanh' 5 / 3, 'relu' sqrt(2), and 'lrelu' sqrt(2 / (1 + leakiness^2)).

What does nninit.sparse do to a tensor?

It sets (1 - sparsity) percent of the tensor to 0, with sparsity between 0 and 1, so a sparsity of 0.2 drops out 80% of the tensor.

Official sources

  1. Official README
  2. Project repository