kyegomez/Gemini: Open-Source PyTorch Implementation of the Gemini Architecture
The open source implementation of Gemini, the model that will "eclipse ChatGPT" by Google
At a glance
- What is it?
- kyegomez/Gemini is a community PyTorch implementation of the multimodal transformer architecture described in Google DeepMind's Gemini paper. It is a research starting point, not a trained model, and several features from the paper remain unimplemented.
- Who is it for?
- kyegomez/Gemini is a useful reference for researchers who want to study or extend the multimodal transformer architecture described in the Gemini paper, without waiting for official code releases. It is not a replacement for the actual Gemini API and should not be treated as a production model.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What This Repository Is and What It Is Not
kyegomez/Gemini is a community implementation of the Gemini multimodal transformer architecture published by Google DeepMind. The repository is not affiliated with Google and does not contain trained weights. The README describes its goal as implementing the architecture from the paper so researchers can study, run, and extend it without access to Google's internal codebase.
The project is aimed at ML researchers and engineers who want to experiment with a multimodal transformer that processes text, image, and audio inputs natively. The README explains the core idea: instead of using a separate vision encoder like a ViT before the transformer, Gemini feeds image embeddings directly into the transformer. Audio is ingested at 16 kHz using features from Google's Universal Speech Model, as described in the paper. This direct multi-modal processing is what distinguishes the architecture from approaches that use adapter layers or cross-attention between a language model and a frozen vision encoder.
Architecture: Transformer with Flash Attention and Multi-Modal Inputs
The core class is Gemini in gemini_torch/model.py. The README documents several configurable attention mechanisms: Multi Grouped Query Attention, Flash Attention, RoPE positional encoding, ALiBi, xPos, and QK normalization. The no-pos-embeds option and a KV cache are also listed.
Inputs are encoded as tokens with special modality markers. The README mentions [IMG], [AUDIO], and similar markers for indicating the start and end of non-text modalities. Images are passed as embeddings fed directly to the transformer. Audio is represented as features, matching the USM-based approach described in the paper.
A second class, LongGemini, adds Ring Attention for long-context processing. The README notes that LongGemini does not include multi-modal processing yet. It takes a ring_seq_size parameter that controls the sequence chunking for distributed attention.
The tokenizer class, MultimodalSentencePieceTokenizer, wraps the LLaMA tokenizer from Hugging Face and adds special tokens for each modality. The README notes it does not yet fully process images, audio, or video and that help is needed in that area.
Installing gemini-torch and Running the Base Transformer
The package installs from PyPI:
pip3 install gemini-torchDependencies include torch, einops, sentencepiece, and ring-attention-pytorch, plus zetascale. Python 3.9 or newer is required, as listed in pyproject.toml.
The README shows a minimal usage example for the text-only transformer with reduced dimensions for local testing:
import torch
from gemini_torch.model import Gemini
model = Gemini(
num_tokens=50432,
max_seq_len=4096,
dim=1280,
depth=16,
dim_head=64,
heads=12,
use_abs_pos_emb=False,
attn_flash=True,
attn_kv_heads=2,
qk_norm=True,
attn_qk_norm=True,
attn_qk_norm_dim_scale=True,
)
text = torch.randint(0, 50432, (1, 4096))The full multi-modal version adds post_fusion_norm and post_modal_transform_norm flags. The README presents the multi-modal example with further reduced dimensions to keep memory use manageable on smaller machines.
LongGemini: Ring Attention for Extended Context
The repository includes a separate LongGemini class that wraps the transformer with Ring Attention. Ring Attention is a distributed attention mechanism that splits long sequences into chunks processed in a ring topology, enabling longer context windows than a single device can fit in memory. The README gives a usage example:
import torch
from gemini_torch import LongGemini
x = torch.randint(0, 10000, (1, 1024))
model = LongGemini(
dim=512,
depth=32,
dim_head=128,
long_gemini_depth=9,
heads=24,
qk_norm=True,
ring_seq_size=512,
)The long_gemini_depth parameter controls how many of the transformer layers use the Ring Attention mechanism. The ring_seq_size parameter sets the chunk size. The README notes that LongGemini does not yet include multi-modal processing; it handles only text token sequences at this stage.
What Is Implemented and What Remains in the TODO List
The README's TODO section is explicit about what is missing. Video processing is listed as unimplemented. The specific technique described in the paper, encoding video as a sequence of frames in a long context window interleaved with text, is in the TODO. A chain-of-thought prompting approach that the paper found improved accuracy on difficult benchmarks is also listed as unimplemented. Training a 1.8B and a 3.25B parameter model (the Nano-1 and Nano-2 sizes referenced in the paper) is a further TODO item.
What is marked as complete: image feature embedding with alignment to text, and audio processing using USM features. These are checked off in the TODO list, meaning the code structure for those paths exists.
The repository has no GitHub releases. There are no published model weights or evaluation results. The architecture can be instantiated and run as a forward pass, but the model is randomly initialised. Any output it produces reflects random weights, not trained knowledge.
The Multimodal Tokenizer and Sentencepiece Integration
Handling multiple modalities requires a tokenizer that can represent non-text inputs as token sequences. The repository includes MultimodalSentencePieceTokenizer in gemini_torch/tokenizer/, which wraps the LLaMA tokenizer from Hugging Face's transformers library and extends it with special tokens for each modality.
The README usage example shows encoding an audio description string with a modality argument:
from gemini_torch.tokenizer import MultimodalSentencePieceTokenizer
tokenizer_name = "hf-internal-testing/llama-tokenizer"
tokenizer = MultimodalSentencePieceTokenizer(tokenizer_name=tokenizer_name)
encoded_audio = tokenizer.encode("Audio description", modality="audio")
decoded_audio = tokenizer.decode(encoded_audio)The README explicitly notes that the tokenizer does not yet fully process images, audio, or video inputs and that the project is looking for contributors to help extend it. This is one of the more significant gaps between the current implementation and what a training run would require.
Architectural Relationship to Fuyu and Design Rationale
The README explicitly describes the design rationale for the multimodal approach. The README states that the architecture bears resemblance to Fuyu's architecture but is expanded to cover multiple modalities. The key difference from approaches that use a separate ViT encoder is that Gemini feeds image embeddings directly into the transformer without an intermediate encoder step. This reduces the pipeline complexity and allows the transformer to handle multiple modalities within a single forward pass.
The README also explains the conditional decoding approach: Codi, described as a component of Gemini, uses conditional generation and tokenised outputs to produce image outputs as discrete image tokens. This approach follows the method described in the paper rather than using a separate diffusion-based image decoder.
For audio, the README states that audio is ingested at 16 kHz using features from Google's Universal Speech Model (USM). The motivation given is that mapping audio directly to text typically loses nuance captured by USM features. This is the technique described in the Gemini paper that the implementation aims to reproduce.
The README's References section lists further improvements that the project intends to incorporate: combining reinforcement learning with the pretrained transformer, self-improving mechanisms inspired by RoboCat, speculative decoding, the Algorithm of Thoughts prompting technique, and RLHF. These are listed without implementation status, suggesting they are future directions rather than completed features.
Limitations and the Alternative of the Official Gemini API
The most significant limitation is the absence of trained weights. Instantiating the model produces a randomly initialised transformer. To get meaningful outputs, a user would need to train from scratch on a large multi-modal dataset, which requires substantial compute and data infrastructure that the README does not discuss.
The zetascale dependency is a third-party package not part of the standard PyTorch ecosystem. Its API and stability are outside the scope of this repository's maintenance. If zetascale changes in a breaking way, this implementation breaks too. The requirements.txt pins zetascale without a version constraint, which means pip may install an incompatible version.
The direct alternative for anyone who needs a working Gemini-class model is the official Gemini API from Google DeepMind. The official API provides access to trained production models, handles infrastructure, and is updated by Google. It does not provide model weights or allow architectural modifications. kyegomez/Gemini fills the opposite role: it is open, modifiable, and extensible, but requires training before it can do useful work.
The repository is MIT licensed. The last push was on 2026-09-14.
Editorial conclusion
kyegomez/Gemini is a useful reference for researchers who want to study or extend the multimodal transformer architecture described in the Gemini paper, without waiting for official code releases. It is not a replacement for the actual Gemini API and should not be treated as a production model. The TODO list and the absence of trained weights mean it produces architecturally correct but untrained outputs. Before building on this, verify that the zetascale and ring-attention-pytorch dependencies install cleanly in your environment, since these are third-party packages not from the PyTorch ecosystem and may have their own compatibility constraints.
Frequently asked questions
Is kyegomez/Gemini the same as Google's official Gemini AI?
No. kyegomez/Gemini is a community PyTorch implementation of the architecture described in the Gemini paper. It is not affiliated with Google and contains no trained weights. The official Gemini models are proprietary and accessible only through Google DeepMind's API.
Does kyegomez/Gemini include pre-trained model weights?
No. The repository contains architecture code only. Any instance of the model is randomly initialised. Training from scratch requires external datasets and compute infrastructure not described in the README.
Which attention mechanisms does kyegomez/Gemini support?
The README lists Multi Grouped Query Attention, Flash Attention, RoPE, ALiBi, xPos, and QK normalization, all configurable through constructor arguments to the Gemini class.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kyegomez-gemini)