The package is named attention and it installs TensorFlow
Keras Attention Layer (Luong and Bahdanau scores).
At a glance
- What is it?
- keras-attention is a single Keras layer offering the Luong and Bahdanau score functions, three input dimensions in and one vector out. Its PyPI name is the single word attention, its install requirement is tensorflow>=2.1 with no ceiling, and its examples import from tensorflow.keras rather than from Keras.
- Who is it for?
- keras-attention is worth reading if you want one readable attention layer instead of a framework's built-in one, and the two score functions are implemented as described. Before you depend on it, check three things.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One word on PyPI, and TensorFlow in the requirements
The install line is one command.
pip install attentionThe distribution name is the single generic word attention, not keras-attention, and nothing in the command disambiguates it. What it does declare is heavier than the layer it delivers. The packaging file lists exactly two requirements, numpy at 1.18.1 or newer and tensorflow at 2.1 or newer, so requesting this one small Keras layer installs a machine learning framework as a side effect. There is no extra marker for the framework, no optional dependency group, and no upper bound on either requirement. The package itself is one directory: the root of the repository holds the attention package, the examples directory, the setup file, a tox configuration and a licence, with nothing else in the source tree to read.
The tested line stops at 2.14 and the floor is 2.1
Two version statements sit on the page and they do not line up. The tested sentence covers TensorFlow 2.8, 2.9, 2.10, 2.11, 2.12, 2.13 and 2.14, and it carries its own date, Sep 26, 2023. The packaging requirement is a floor of 2.1, seven minor releases below the oldest version anyone is said to have tried, with nothing above 2.14 declared or excluded. The release history explains the gap between the two dates at least partly: the current release is 5.0.0, tagged in March 2023, and the one before it that is visible is named 3.0, tagged in September 2020, with a note that it added the attention layer to Sequential. So the tag naming is inconsistent, 3.0 against 5.0.0, and the most recent recorded push to the repository is 12 March 2026, some six months after the newest release and nearly three after the newest tested version. Nothing on the page says what changed in between.
Three empty image blocks where the demonstrations should be
Two of the three worked demonstrations end with a sentence promising a picture and then an empty element. The adding example says an overview of the training is shown below, where the top represents the attention map and the bottom the ground truth, and that as training progresses the attention map converges to the ground truth. Below that text is an empty centred paragraph tag and nothing else. The finding-the-maximum example does the same thing: it says the attention layer converges perfectly to what was expected after a few epochs, and is followed by another empty centred paragraph. The images themselves are not missing from the repository. The examples directory holds an animated GIF of the attention output and a PNG of equations, plus a readme subdirectory, so the assets are there and the page does not point at them.
The headline example stops in the middle of a keyword argument
The example section imports numpy, pulls Input from tensorflow.keras, Dense and LSTM from tensorflow.keras.layers, and the model helpers from tensorflow.keras.models, then imports the layer itself from the attention package. The model definition begins with an input of shape time_steps by input_dim and a comment explaining that the data is dummy and there is nothing to learn in this example. Then it stops, mid-expression, at x = LSTM(64, return_sequences=. There is no error handling, no output shape, no call to Attention and no line that fits anything. What the reader is left with is a function named main and a set of local names. The working version of this is a separate file in the examples directory, and the examples have their own requirements file which has to be installed separately before any of them run, so the snippet on the page is an index into the repository rather than something you can paste.
The IMDB result takes a maximum over ten epochs on the test set
This is the only measurement on the page and the method matters as much as the numbers. Two LSTM networks are compared, one with the attention layer and one with a fully connected layer, both held at 250K parameters for fairness. The results are over 10 runs, and for every run the recorded figure is the maximum accuracy on the test set across 10 epochs.
| Measure | No Attention | Attention | | ------------- | ------------- | ------------- | | MAX Accuracy | 88.22 | 88.76 | | AVG Accuracy | 87.02 | 87.62 | | STDDEV Accuracy | 0.18 | 0.14 |
So the difference is 0.54 on the maximum and 0.60 on the average, and the run to run standard deviation is reported next to it, 0.18 against 0.14, which the page reads as a boost in accuracy plus reduced variability. Two things are missing that you would want before treating that as settled. A maximum over ten epochs is a selected statistic, and taking it on the test set makes the test set the selection set rather than a held-out one. And no per-run values, seeds or significance test are given, so the 0.6 gap has to be read against the 0.18 spread by eye.
The badge row has two identical links and no build status
The row of badges at the top of the page has four entries and only two distinct destinations. Two of them are the same package download counter at pepy.tech for this project, linked twice. The third is the licence file in the repository, and the fourth is a link to the TensorFlow website, which is not a status indicator for this project at all. There is no build badge, no test badge and no coverage badge, which is consistent with the rest of the page: the tox configuration file in the repository root is never mentioned, and the only instruction about testing anywhere is the tested-versions sentence. The page is also thin on provenance. The author field in the packaging file names one person, the licence is declared there as Apache 2.0 while the licence field in the repository's own settings gives the SPDX form Apache-2.0, and the long description handed to the package index is the README read straight off disk.
One layer, two score functions, and the time axis disappears
The entire public API is a single class with two named arguments.
Attention(
units=128,
score='luong',
**kwargs
)units is the number of output units in the attention vector, and score selects the scoring function applied to the decoder state and the encoder states, with two possible values. Luong's is described as the multiplicative style and Bahdanau's as the additive style, each with a link to its paper. The input is a 3D tensor of batch_size, timesteps and input_dim. The output is a 2D tensor of batch_size by num_units, so the layer collapses the time axis and returns one vector per sample rather than a per-step weighting. Nothing in the signature exposes the attention weights themselves: to see them the page sends you to the adding example, and the two other examples in that directory are a sequence-maximum task and the IMDB comparison.
Editorial conclusion
keras-attention is worth reading if you want one readable attention layer instead of a framework's built-in one, and the two score functions are implemented as described. Before you depend on it, check three things. The install name is generic enough that a wrong package could satisfy your requirement, so pin it deliberately. The declared dependency is a floor with no upper bound while the tested line stops at TensorFlow 2.14 as of September 2023, so you are choosing how much of the gap your own testing covers. And if you rely on the IMDB result, note that it takes the maximum test accuracy over ten epochs, which uses the test set for selection.
Frequently asked questions
How do I install the Keras attention layer?
With pip install attention. The distribution is published under the single word name attention rather than keras-attention, and it declares numpy at 1.18.1 or newer and tensorflow at 2.1 or newer, so the install brings TensorFlow with it. The examples have a separate requirements file at examples/examples-requirements.txt that you install with pip install -r.
Which Keras does keras-attention work with?
The examples import from tensorflow.keras, from tensorflow.keras.layers and from tensorflow.keras.models, rather than from a standalone keras package, so the code as written targets the Keras bundled with TensorFlow. The page's tested line covers TensorFlow 2.8 through 2.14 as of Sep 26, 2023, while the declared requirement only sets a floor at tensorflow 2.1 with no upper bound.
What is the difference between the luong and bahdanau score functions?
The page describes Luong's as the multiplicative style and Bahdanau's as the additive style, and links a paper for each. You choose with the score argument, whose stated possible values are exactly those two. Both take the same input shape of batch_size, timesteps and input_dim and return a 2D tensor of batch_size by num_units, collapsing the time axis.
Does the attention layer improve accuracy on the IMDB dataset?
The page reports that it does, on its own experiment. Two 250K parameter networks are compared over 10 runs, recording the maximum test accuracy across 10 epochs: max accuracy goes from 88.22 to 88.76, average from 87.02 to 87.62, and standard deviation from 0.18 to 0.14. No per-run values, seeds or significance test are given.
How can I see the attention weights?
The layer does not return them. Its documented output is a single vector of batch_size by num_units, and the page points you at examples/add_two_numbers.py for a demonstration of the weights on a small sequence task. The examples directory also holds a sequence-maximum task, the IMDB comparison, and image assets including an animated GIF and an equations image.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/philipperemy-keras-attention)