Model or dataset
salesforce/CodeGen avatar
salesforce/CodeGen

Salesforce CodeGen: Open-Source LLMs for Program Synthesis

CodeGen is a family of open-source model for program synthesis. Trained on TPU-v4. Competitive with OpenAI Codex.

5,180 stars420 forksPythonApache-2.0

At a glance

What is it?
CodeGen is a family of open-source language models from Salesforce AI Research, trained to generate code from natural language prompts. Three generations have been released: CodeGen1, CodeGen2, and CodeGen2.5, all available as weights on Hugging Face and loadable with the transformers library.
Who is it for?
Research teams and engineers who need open-weight code generation models to experiment with or fine-tune should look at CodeGen as one reference point in the code LLM landscape. The models are research releases, not production inference services, and the ethics disclaimer in the repository is explicit that they were not evaluated for all downstream applications.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 120 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Program Synthesis as a Hugging Face Checkpoint

CodeGen addresses the problem of generating code from natural language descriptions using open-weight models. Unlike GitHub Copilot or similar hosted services, CodeGen releases the model weights directly, so teams can run inference locally, fine-tune the models on private codebases, or study the model behavior without using an external API.

Salesforce AI Research released three generations. CodeGen1, announced in March 2022, matched OpenAI Codex performance at the time according to the README. CodeGen2, released in May 2023, added infill sampling capability, which allows the model to complete code given both a prefix and a suffix. CodeGen2.5, released in July 2023, achieves performance that the README describes as outperforming 16B parameter models with only 7B parameters.

The audience is research teams and engineers who need open-weight code generation checkpoints. The README includes an ethics disclaimer stating the models are 'for research purposes only in support of an academic paper' and were not evaluated for all downstream applications.

Three Generations: CodeGen1, CodeGen2, and CodeGen2.5

Each generation introduces a different capability focus. CodeGen1 targeted multi-turn program synthesis, where the model builds a function across multiple conversation turns. The architecture and training details are described in the ICLR 2023 paper listed in the README. Model sizes for CodeGen1 are 350M, 1B, 3B, 7B, and 16B parameters.

CodeGen2 added infill sampling, which allows the model to generate code that fits between a known prefix and suffix. This is the capability that powers features like filling in the middle of a function signature. Model sizes for CodeGen2 are the same range as CodeGen1.

CodeGen2.5 targets the 7B parameter range with improved performance. The README states it 'outperforms 16B parameter models with only 7B.' The models are named with a suffix indicating the training data: mono for monolingual Python training and multi for multilingual training.

All weights are hosted on the Hugging Face Hub under the Salesforce organization. The model naming follows the pattern Salesforce/codegen-2B-mono for CodeGen1 and Salesforce/codegen2-7B for CodeGen2.

Loading a CodeGen Model with the transformers Library

All three generations load through the Hugging Face transformers library using AutoTokenizer and AutoModelForCausalLM.

For CodeGen1:

python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("Salesforce/codegen-2B-mono")
model = AutoModelForCausalLM.from_pretrained("Salesforce/codegen-2B-mono")
inputs = tokenizer("# this function prints hello world", return_tensors="pt")
sample = model.generate(**inputs, max_length=128)
print(tokenizer.decode(sample[0], truncate_before_pattern=[r"\n\n^#", "^'''", "\n\n\n"]))

For CodeGen2, the README adds trust_remote_code=True and revision="main" to the model load:

python
model = AutoModelForCausalLM.from_pretrained("Salesforce/codegen2-7B", trust_remote_code=True, revision="main")

For CodeGen2.5, the tokenizer also requires trust_remote_code=True:

python
tokenizer = AutoTokenizer.from_pretrained("Salesforce/codegen25-7b-mono", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Salesforce/codegen25-7b-mono")

Loading a 7B or 16B model requires significant GPU memory. The README does not document minimum hardware requirements or quantization options. Teams running these models on consumer hardware should check Hugging Face model pages for community notes on memory requirements.

Multi-Turn Program Synthesis and the Jaxformer Training Library

CodeGen1's defining capability is multi-turn program synthesis, which the README paper title names explicitly: 'An Open Large Language Model for Code with Multi-Turn Program Synthesis.' The model is trained to accept a conversation history where the user iteratively refines a code generation request across multiple turns, with the model building the final function incrementally.

The Jaxformer library handles data preprocessing, training, and fine-tuning for all CodeGen generations. The README links to it as a separate repository at github.com/salesforce/jaxformer. Jaxformer is the required tool for anyone reproducing the training setup or fine-tuning on a custom codebase. The CodeGen repository itself contains only inference code and model references; it does not include training scripts.

CodeGen2's infill capability changes the input format. The model accepts a prefix and suffix, and generates the code segment that would fit between them. This is distinct from standard left-to-right completion and requires a specific prompting format that the model was trained to recognize.

Limitations: Research Release, Ethics Constraints, and Hardware Requirements

The README's ethics disclaimer is direct: the models 'are not specifically designed or evaluated for all downstream purposes.' Salesforce AI Research recommends that users evaluate accuracy, safety, and fairness before deploying the models in applications. The disclaimer links to Salesforce's Acceptable Use Policy and AI Acceptable Use Policy.

There is no fine-tuning recipe in the CodeGen repository itself. Fine-tuning requires Jaxformer, which is a separate project with its own dependencies and setup process.

The repository has no GitHub releases. The last push was on 2026-06-02. The three model generations were all published between March 2022 and July 2023. There are no indications in the README of a CodeGen3 or a subsequent generation.

Loading the 16B parameter model requires roughly 32 GB of GPU memory in fp16, though the README does not specify this directly. Teams without access to that hardware cannot run the largest models locally.

CodeGen vs GitHub Copilot: Open Weights vs Hosted Inference

GitHub Copilot is a hosted code generation service built on top of OpenAI Codex and subsequent models. It is available through a subscription and runs entirely on remote infrastructure. There are no downloadable weights, no fine-tuning API, and no way to inspect the model internals.

CodeGen offers the opposite trade-off: weights are public, fine-tuning is possible through Jaxformer, and the full training methodology is described in peer-reviewed papers. The cost is infrastructure: running a 7B or 16B model locally requires hardware that most individual developers do not own.

For teams that want to understand code generation model behavior, contribute to research, or fine-tune on a private codebase without sending code to an external API, CodeGen provides the access that Copilot does not. For teams that want a fully managed, low-setup coding assistant without infrastructure concerns, Copilot's hosted model is the more practical choice.

License, Citation Requirements, and Maintenance Status

CodeGen is licensed under Apache-2.0, which permits commercial use, modification, and distribution. The ethics disclaimer is not a license restriction but a recommendation to evaluate the model before deployment.

The README includes two BibTeX entries for citing the CodeGen and CodeGen2 papers in academic work. The Jaxformer library is referenced by URL rather than as a direct dependency of the repository.

The last push was on 2026-06-02. No new model generations have been announced in the repository. The assets/ directory in the repository root contains images used in the README, and the codegen1/, codegen2/, and codegen25/ subdirectories contain generation-specific README files and any supplementary code.

Editorial conclusion

Research teams and engineers who need open-weight code generation models to experiment with or fine-tune should look at CodeGen as one reference point in the code LLM landscape. The models are research releases, not production inference services, and the ethics disclaimer in the repository is explicit that they were not evaluated for all downstream applications. Teams building a coding assistant for production use should evaluate whether fine-tuning a CodeGen checkpoint or using a managed inference endpoint is the right trade-off. The Jaxformer library, referenced in the README for data preprocessing and training, is a separate repository that is required if reproducing the training setup rather than just running inference.

Frequently asked questions

What is Salesforce CodeGen and what is it used for?

Salesforce CodeGen is a family of open-source language models trained to generate code from natural language prompts. It includes models from 350M to 16B parameters across three generations: CodeGen1 for multi-turn program synthesis, CodeGen2 with infill capability, and CodeGen2.5 with improved efficiency at the 7B scale.

How do I load a CodeGen model for inference?

All CodeGen variants load through the Hugging Face transformers library using AutoTokenizer.from_pretrained and AutoModelForCausalLM.from_pretrained with the appropriate Salesforce model name. CodeGen2 and CodeGen2.5 require trust_remote_code=True in the from_pretrained call.

Can CodeGen models be fine-tuned on a private codebase?

Yes. The Jaxformer library, linked in the CodeGen README, handles data preprocessing, training, and fine-tuning. Fine-tuning requires Jaxformer as a separate setup; the CodeGen repository itself contains only inference code and model references.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. salesforce/CodeGen on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/salesforce-codegen.svg)](https://hysenlabs.com/projects/salesforce-codegen)