# kyegomez/zeta: modular PyTorch blocks for attention, MoE and quantization

> Zeta is a PyTorch component library published on PyPI as zetascale. It gives you attention variants, feedforward modules, a BitLinear quantized layer and full encoder-decoder structs, but its release history and documentation are uneven.

**kyegomez/zeta** — Build high-performance AI models with modular building blocks

- Repository: https://github.com/kyegomez/zeta
- Website: https://zeta.apac.ai
- Stars: 603 · Forks: 58
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/kyegomez-zeta

## What zeta is for, and who should reach for it

Zeta is a collection of PyTorch modules rather than a training framework or a model zoo. The README calls it "a modular PyTorch framework designed to simplify the development of AI models by providing reusable, high-performance building blocks." The intended user is someone who is already writing PyTorch and is tired of reimplementing multi-query attention, SwiGLU, relative position bias or a quantized linear layer every time a paper lands.

The repository topics confirm the focus: attention-mechanism, transformer-architecture, ffns, pytorch-implementation. There is no serving layer, no distributed training launcher and no dataset pipeline in the top-level layout. The directories are zeta/, tests/, examples/, docs/ and experimental/, plus a stray multi_query_attention.py at the root.

If your problem is "I need a working MultiQueryAttention class with a known tensor signature in the next ten minutes," zeta is aimed at you. If your problem is "I need to fine-tune a 7B model on eight GPUs with checkpoint resumption," zeta does not address it and the README does not claim to.

## The module surface: attention, MoE, feedforward, quantization, structs

The README lists six families. Attention mechanisms, including multi-query, sigmoid and flash attention. Mixture of Experts with routing and gating. Neural network modules such as feedforward networks, activations and normalization. Quantization, specifically BitLinear and dynamic quantization. Architectures, meaning transformers, encoders, decoders, vision transformers and "complete model implementations." Training utilities, meaning optimization algorithms, logging and performance monitoring.

The practical entry points are the two import namespaces shown in the examples: `from zeta import MultiQueryAttention` for the top-level attention module, and `from zeta.nn import SwiGLUStacked, RelativePositionBias, FeedForward` for the smaller layers. Quantization lives under `zeta.quant`, and the composition primitives (Encoder, Decoder, Transformer, ViTransformerWrapper, AutoRegressiveWrapper) live under `zeta.structs`.

That split matters when you read the code. The structs are assembled from the nn and attention pieces, so a bug in a base layer propagates into every architecture built on top of it. The README does not describe a versioning policy for those layers.

## Installing zetascale and running a first attention block

The README gives a single install line. The distribution name on PyPI is zetascale, not zeta, so the import name and the install name differ.

```bash
pip3 install -U zetascale
```

After that, the README's first example builds a multi-query attention layer with dim=512 and heads=8, runs a random tensor of shape (2, 4, 512) through it, and prints the output shape. Note that the forward pass returns a tuple, not a tensor, so you unpack three values.

```python
import torch
from zeta import MultiQueryAttention

model = MultiQueryAttention(dim=512, heads=8)
text = torch.randn(2, 4, 512)
output, _, _ = model(text)
print(output.shape)  # torch.Size([2, 4, 512])
```

A smaller first check is the SwiGLU stack, which takes a (5, 10) input and returns (5, 20) when constructed as SwiGLUStacked(10, 20).

```python
import torch
from zeta.nn import SwiGLUStacked

x = torch.randn(5, 10)
swiglu = SwiGLUStacked(10, 20)
print(swiglu(x).shape)  # torch.Size([5, 20])
```

If either script fails at import time rather than at the forward pass, the cause is usually the dependency set, not the module. The pyproject.toml pins joblib to >=1.3.0,<1.4.0 and scikit-learn to >=1.5.0,<1.6.0 with the comment "Pin compatible versions to prevent import errors," which tells you the authors have hit that failure mode themselves.

## BitLinear and the quantization path

BitLinear is the one component in the README tied to a specific paper: the code comment attributes it to BitNet: Scaling 1-bit Transformers for Large Language Models. It replaces nn.Linear and performs quantization plus dequantization around the linear transform.

```python
import torch
from torch import nn
import zeta.quant as qt

class MyModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.linear = qt.BitLinear(10, 20)

    def forward(self, x):
        return self.linear(x)

model = MyModel()
print(model(torch.randn(128, 10)).size())  # torch.Size([128, 20])
```

The README claims it reduces memory usage "while maintaining performance" but gives no measurement, no comparison against a full-precision baseline and no note on which layers are safe to swap. Treat the substitution as an experiment you have to validate on your own evaluation set. The dependency list also includes bitsandbytes, so the quantization story is not self-contained in pure PyTorch.

## Where zeta stops being the right tool

The README does not document rollback, deprecation or a migration path between versions. The release list shows 2.3.7 in April 2024 and 0.0.111 in July 2023, while pyproject.toml declares version 2.8.8. That gap between the published releases and the in-tree version means the code on master and the code you get from pip may not be the same thing, and the README does not explain how to reconcile them.

Licensing is also inconsistent. The README badge says MIT, the pyproject.toml license field says MIT, but the repository metadata for this project records Apache-2.0. If you are redistributing zeta inside a product, resolve that before you ship, and the README does not help you do it.

The third limit is scope. Zeta gives you building blocks and a handful of architecture examples, such as the PalmE class in the README that wires a ViTransformerWrapper encoder to a Transformer decoder. It does not give you a trainer, a checkpoint format or a config system. For anything past a research prototype you will be writing that layer yourself, and the "Training Utilities" bullet in the README is the only description of what exists there.

## How it compares to lucidrains libraries and plain torch.nn

The repository topics include lucidrains, and the overlap is real: multi-query attention, relative position bias, feedforward with GLU, and vector quantization all appear in that family of packages too. The difference in approach is packaging. Lucidrains projects tend to ship one paper per repository with a narrow, documented API and its own README, so you install exactly the component you want. Zeta bundles many of them into one distribution, zetascale, and re-exports them under zeta, zeta.nn, zeta.quant and zeta.structs. You get one dependency to manage and a consistent import path; you also get the whole dependency graph, including transformers, torchvision, accelerate and bitsandbytes, whether or not your module needs them.

The other comparison is just writing the layer yourself. A multi-query attention block is not enormous, and the README's own example is nine lines. The case for zeta is breadth and the claim of fused kernels "where applicable," not novelty. If you only need one layer, the cost of a dependency tree this size probably exceeds the cost of writing it.

## Maintenance, versioning and what the metadata actually says

The last push to the repository was on 2026-09-07, so the code is being touched. That is not the same as a release cadence. The most recent release listed is 2.3.7 from 2024-04-06, roughly two and a half years before that push, and pyproject.toml carries 2.8.8. Nothing in the README or the repository files explains the relationship between the tag, the PyPI artifact and master, and there is no changelog in the top-level entries.

The upgrade cost follows from that. You cannot read a release note to learn what changed between 2.3.7 and the current tree, so pinning zetascale and reading the diff yourself is the only reliable check. The lint group pins ruff to >=0.5.1,<0.5.2, which suggests the project holds tooling versions tightly while leaving the version number of the package itself loose.

On licence, the README badge and pyproject.toml both say MIT, and the repository metadata says Apache-2.0. Both are permissive, but they carry different notice and patent terms. This is not legal advice; it is a reason to confirm the actual LICENSE file before you redistribute.

## Conclusion

Adopt zeta if you are prototyping transformer variants and want ready-made attention, feedforward and quantization modules you can import instead of writing from scratch; read the source of the specific module you import before you commit to it, because the README does not document rollback, deprecation or version compatibility between releases. Skip it if you need a stable, versioned API with a changelog, or if your production stack cannot absorb the dependency list in requirements.txt. Before adopting, check whether the module you need is covered by tests/ and confirm which version pyproject.toml declares, since the package metadata and the README badge disagree on the licence.

## FAQ

### What is the PyPI package name for kyegomez/zeta?

The distribution is published as zetascale, and the README's install command is pip3 install -U zetascale. The import name is different: you write from zeta import MultiQueryAttention.

### Does zeta include a training loop or trainer?

The README lists training utilities among the components, described only as optimization algorithms, logging and performance monitoring. There is no trainer class shown in the examples, and the repository layout has no top-level training module outside examples/training/.

### Which Python version does zeta require?

The pyproject.toml declares python = "^3.10" and classifies the package for Python 3.10. The README does not state a supported version range beyond that.

## Sources

- [kyegomez/zeta on GitHub](https://github.com/kyegomez/zeta)
- [License: Apache-2.0](https://github.com/kyegomez/zeta/blob/master/LICENSE)
- [Project website](https://zeta.apac.ai)
- [README](https://github.com/kyegomez/zeta/blob/master/README.md)
- [Releases](https://github.com/kyegomez/zeta/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kyegomez-zeta
