Model or dataset
salesforce/xgen avatar
salesforce/xgen

XGen-7B: Salesforce's Research LLM with an 8,000-Token Context Window

Salesforce open-source LLMs with 8k sequence length.

726 stars40 forksPythonApache-2.0

At a glance

What is it?
XGen is a family of 7-billion-parameter language models from Salesforce AI Research, trained specifically to process sequences of up to 8,000 tokens. The release is research-only and ships no fine-tuning scripts or evaluation tooling beyond what HuggingFace model cards provide.
Who is it for?
XGen suits researchers who need a documented 7B baseline for long-context NLP benchmarks and who are comfortable loading large model weights from HuggingFace. It is not for production deployments: the ethics disclaimer in the repository explicitly limits it to research in support of an academic paper, and the requirements file pins transformers to version 4.29.2, which is far behind current releases.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 120 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem XGen addresses and who it is for

Most open-weight language models at the time of XGen's release were trained on sequences of 2,048 tokens or fewer. Extending context length requires changes to both the training data and the positional encoding scheme, and those changes have real costs in compute time and memory. XGen was built specifically to demonstrate that a 7B model can be trained reliably on 8,000-token inputs using the OpenAI Tiktoken tokenizer, which was an unusual choice for an open-weight model at the time.

The target audience is researchers running reproducible experiments that require a documented baseline with long context. The repository's ethics disclaimer makes the scope explicit: this is a release in support of an academic paper, not a foundation for production systems. Salesforce published a technical report on arXiv (arXiv:2309.03450) as the primary reference for the model's design, training procedure, and reported capabilities. Teams building applications are directed toward that paper rather than toward any operational guide in the repository.

Three model variants and what distinguishes them

The repository provides three model cards published to HuggingFace:

- XGen-7B-4K-Base: a base language model trained with a 4,000-token sequence length. - XGen-7B-8K-Base: a base language model trained with an 8,000-token sequence length. - XGen-7B-8k-Inst: the 8K model with instruction fine-tuning, released for research purposes only.

The distinction between the 4K and 8K base models is useful for ablation studies: a researcher can compare long-context and standard-context performance on the same architecture by switching between them. The instruction-tuned variant supports prompted dialogue but carries an explicit research-only restriction, making it unsuitable for user-facing deployments under the terms the repository describes.

All three variants use the OpenAI Tiktoken package for tokenization. The repository's README points out that Tiktoken must be installed separately before loading any of the models, as it is not automatically pulled in by the transformers dependency.

Loading the model and running inference

The only dependency listed in requirements.txt is a pinned version of the transformers library:

bash
pip install tiktoken

After installing Tiktoken, the model loads through the standard HuggingFace AutoModel interface. The README provides a complete example that generates text in bfloat16 precision:

python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("Salesforce/xgen-7b-8k-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Salesforce/xgen-7b-8k-base", torch_dtype=torch.bfloat16)
inputs = tokenizer("The world is", return_tensors="pt")
sample = model.generate(**inputs, max_length=128)
print(tokenizer.decode(sample[0]))

The trust_remote_code=True flag is required because the tokenizer includes custom code for Tiktoken integration. That flag means the model will execute code downloaded from HuggingFace, which is a security consideration in restricted environments. The bfloat16 dtype reduces memory use compared to float32 but still requires hardware that supports it, which rules out older GPUs without bfloat16 support such as consumer Ampere cards with limited bfloat16 throughput.

Dependency pinning and what it costs

The requirements.txt file contains a single line: transformers==4.29.2. That version was current in mid-2023 when the paper was submitted. The transformers library has had dozens of releases since then, including changes to model loading, attention implementations, and the tokenizer interface. Pinning to 4.29.2 means XGen will not benefit from those changes without manual intervention, and it may conflict with other libraries in an environment that expects a newer transformers version.

This is a deliberate choice for reproducibility. Research code is typically pinned to the exact environment used to produce the paper's results. Anyone who wants to run XGen in a more modern environment will need to test whether the Tiktoken integration still functions and whether the model loading behavior has changed in a breaking way. The repository does not provide migration guidance for newer transformers versions.

The sample.py file at the repository root reproduces the basic loading example from the README. It exists as a runnable script for verifying that the environment is set up correctly, not as a broader demonstration of capabilities.

Limitations: research scope, no training code, and ethics boundaries

XGen ships no training scripts, no fine-tuning code, and no evaluation harness. It is a model release in the sense of weights and a loading example, not a platform for building on top of. Researchers who want to fine-tune XGen must write their own training loop and manage the transformers version compatibility themselves.

The architecture paper describes the model but the repository does not reproduce the training setup. The README refers readers to the arXiv paper for implementation details. The paper is titled Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length and is available at arxiv.org/abs/2309.03450. The README includes a BibTeX citation entry for academic attribution.

The ethics statement directs users to Salesforce's standard Acceptable Use Policy and its AI-specific AUP before deploying the model. The ethics disclaimer notes explicitly that the models were not evaluated for all downstream purposes and that accuracy, safety, and fairness concerns require user-side evaluation before any deployment.

How XGen compares to other open-weight long-context models

At the time of its 2023 release, XGen's 8K context was longer than what most comparable open-weight models offered. Since then, models such as those in the Mistral family have been released with 32K or longer context windows, with active maintenance, fine-tuning tooling, and GGUF quantized versions for consumer hardware. Those models are also governed by permissive licenses.

The difference in approach between XGen and a model like Mistral is not just context length. Mistral's releases include quantized variants, instruction-following models, and community tooling. XGen is a single paper's artifact: useful as a historical reference point and as a reproducible baseline, but not positioned as a platform for building applications. For practitioners who need long-context inference today, XGen's pinned dependencies and research-only restriction make it a poor fit compared to models with active ecosystem support.

The last push to the repository was on 2026-06-02, and the repository has no GitHub releases. The Apache 2.0 license permits commercial use in principle, but the ethics statement asks users to consult Salesforce's Acceptable Use Policy before deploying the model in any context.

Editorial conclusion

XGen suits researchers who need a documented 7B baseline for long-context NLP benchmarks and who are comfortable loading large model weights from HuggingFace. It is not for production deployments: the ethics disclaimer in the repository explicitly limits it to research in support of an academic paper, and the requirements file pins transformers to version 4.29.2, which is far behind current releases. Engineers looking for a production-ready, actively maintained open-weight model with long context will find more recent alternatives in the LLaMA or Mistral families. Before using XGen, confirm that your hardware can hold a bfloat16 7B model in VRAM and that the Salesforce Acceptable Use Policy permits your intended use.

Frequently asked questions

What is XGen-7B used for?

XGen-7B is a research-only language model from Salesforce designed for long-context NLP tasks. The repository states it is released in support of an academic paper and directs users to evaluate safety and accuracy before any deployment.

Where can I download the XGen model weights?

The weights are published on HuggingFace under the names Salesforce/xgen-7b-4k-base, Salesforce/xgen-7b-8k-base, and Salesforce/xgen-7b-8k-inst. They load through the standard AutoModelForCausalLM interface with trust_remote_code=True.

What version of the transformers library does XGen require?

The requirements.txt file in the repository pins transformers to version 4.29.2. Using a newer version requires manual testing for compatibility with the Tiktoken-based tokenizer.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. salesforce/xgen on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/salesforce-xgen.svg)](https://hysenlabs.com/projects/salesforce-xgen)