Model or dataset
SalesforceAIResearch/xLAM avatar
SalesforceAIResearch/xLAM

xLAM: Salesforce's Large Action Models for Function Calling and Agent Training

xLAM: A Family of Large Action Models to Empower AI Agent Systems

636 stars60 forksPythonApache-2.0

At a glance

What is it?
xLAM is a family of large language models from Salesforce AI Research, fine-tuned specifically for function calling, multi-turn conversation, and agent task execution. It covers models from 1B to 70B parameters and includes ActionStudio, a framework for agentic training data preparation. The repository is provided for research purposes only.
Who is it for?
xLAM is the right starting point for researchers who want a pre-built, Salesforce-trained function-calling model without starting from scratch. The repository note says research purposes only, and part of the data remains unreleased due to internal regulations, so teams building production agents should treat xLAM as a research baseline rather than a deployable foundation.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 120 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What xLAM Addresses in Agent Research

Function calling and multi-step tool use are the main gaps between a general-purpose language model and a model that can reliably execute agent workflows. xLAM is Salesforce AI Research's attempt to fill that gap with a family of models trained on aggregated agent trajectory data. The repository introduces a pipeline that collects trajectories from multiple environments, converts them into a unified format, and trains models to act as agents rather than to converse.

The project also includes ActionStudio, a sub-framework for creating agentic training data. ActionStudio was accepted at the EMNLP 2025 main conference. The repository is marked for research purposes only, and the README states that some data was partially released due to internal regulations.

Model Family: Size, Context, and Release Dates

The xLAM family covers a range of sizes. The README's model table lists the following as of the last push on 2026-06-02:

- Llama-xLAM-2-70b-fc-r: 70B parameters, 128k context, released March 26, 2025 - Llama-xLAM-2-8b-fc-r: 8B parameters, 128k context, released March 26, 2025. GGUF files are available for this model. - xLAM-2-32b-fc-r: 32B parameters, 32k context (max 128k), released March 26, 2025

Older models in the family include xLAM-1b-fc-r and xLAM-7b-fc-r, which were released in July 2024. All models are available via the Salesforce collection on Hugging Face at `Salesforce/xlam-models-65f00e2a0a63bbcd1c2dade4`.

The naming convention uses `fc` for function-calling and `r` for released. The `xLAM-2` prefix marks the second-generation models built on Llama base weights.

Installing and Training with ActionStudio

The repository installs as the `actionstudio` package, with version `v2.0` in setup.py. The requirements.txt pins specific versions of the main dependencies:

bash
transformers>=4.55.0
deepspeed>=0.14.4
ray==2.30.0
peft==0.14.0
accelerate==1.4.0
bitsandbytes==0.45.5

The stack uses DeepSpeed for distributed training, Ray for data handling, and PEFT with bitsandbytes for parameter-efficient fine-tuning. The requirements also include TRL (`trl==0.8.6`) for reinforcement learning from human feedback, WandB for experiment tracking, and OpenAI's Python SDK (`openai>=1.66.2`), which ActionStudio uses for data generation pipelines.

The `requirements.sh` file at the repository root suggests some dependencies require shell-level setup steps beyond pip. The README directs users to the ActionStudio README for the full training configuration details.

APIGen-MT and the Training Data Pipeline

One of the key contributions alongside the models is APIGen-MT, a pipeline for generating multi-turn function-calling training data through simulated agent-human interaction. The project open-sourced the APIGen-MT-5k dataset on Hugging Face at `Salesforce/APIGen-MT-5k`, described as a compact dataset for exploring multi-turn function calling. The full APIGen-MT paper and project website are linked in the README.

The earlier dataset `xlam-function-calling-60k` at `Salesforce/xlam-function-calling-60k` contains 60,000 examples and was used to train the first-generation models. The xLAM training pipeline aggregates trajectories from multiple heterogeneous sources, normalizes them into a consistent format, and maintains balance across data sources during partitioning.

This data pipeline is what distinguishes xLAM from simply fine-tuning a base model on a single function-calling dataset: the emphasis is on handling multi-environment, multi-format trajectory data at scale.

Research-Only Scope and Data Restrictions

The repository README includes a prominent note: the repository is for research purposes only, and data is partially released due to internal regulations. This is not a minor caveat. It means the full training pipeline cannot be reproduced outside Salesforce, and the released datasets represent a subset of what was used to achieve the published results.

Teams considering xLAM for production agent deployment face a related constraint: no production support, no SLA, and no commitment to backward compatibility. The model weights themselves are on Hugging Face and can be used under Apache-2.0, but the training data and pipeline are not fully reproducible from the public repository alone.

The GGUF format is available only for the 8B model. Anyone running the 32B or 70B models needs a GPU environment capable of loading those weights without quantization, unless they quantize the checkpoint independently.

Comparison with Gorilla

Gorilla is a research project from Berkeley focused on LLM tool use and function calling, developed independently of Salesforce. Like xLAM, it targets the gap between general-purpose language models and models that call external APIs reliably. Gorilla is also a research tool, not a production product.

The approaches differ in training emphasis. The Berkeley Function-Calling Leaderboard (BFCL) tracks both; xLAM-2-fc-r is cited in the README as achieving top-1 performance on BFCL as of April 2025. The xLAM family places stronger emphasis on multi-turn agent trajectories and the data pipeline that generates them, while Gorilla's work has concentrated more on API retrieval from large tool sets. Researchers working on multi-turn agent benchmarks will find the two complementary as evaluation baselines.

License and Maintenance Status

xLAM is released under Apache-2.0 in the root LICENSE.txt. Individual components may have separate terms, documented in `license_info.md`. The repository has a CODE_OF_CONDUCT.md, CONTRIBUTING.md, and SECURITY.md, which indicates some level of governance structure.

The last push was on 2026-06-02. The repository has no GitHub releases; version tracking is informal. ActionStudio carries `v2.0` in setup.py. An AI_ETHICS.md file is present at the repository root, which documents the project's ethical use guidelines and is worth reviewing given the research-only designation.

Editorial conclusion

xLAM is the right starting point for researchers who want a pre-built, Salesforce-trained function-calling model without starting from scratch. The repository note says research purposes only, and part of the data remains unreleased due to internal regulations, so teams building production agents should treat xLAM as a research baseline rather than a deployable foundation. Before running training, confirm that the pinned dependency versions in requirements.txt are compatible with the installed CUDA and PyTorch versions, since the stack includes DeepSpeed, Ray, and bitsandbytes, which all have specific hardware requirements.

Frequently asked questions

What is xLAM and which models are available?

xLAM (Large Action Models) is a family of language models from Salesforce AI Research, trained for function calling and multi-turn agent tasks. Available models on Hugging Face include Llama-xLAM-2-70b-fc-r, Llama-xLAM-2-8b-fc-r, and xLAM-2-32b-fc-r, plus older 1B and 7B models.

Can xLAM be used in production applications?

The README states the repository is for research purposes only. The models themselves are Apache-2.0 licensed on Hugging Face, but there is no production support, no SLA, and the training data is only partially released.

What is the Salesforce APIGen-MT-5k dataset?

APIGen-MT-5k is a compact multi-turn function-calling dataset released by Salesforce AI Research on Hugging Face at Salesforce/APIGen-MT-5k. It was generated using the APIGen-MT pipeline, which simulates agent-human interaction to produce training data for multi-turn function calling.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. SalesforceAIResearch/xLAM on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/salesforceairesearch-xlam.svg)](https://hysenlabs.com/projects/salesforceairesearch-xlam)