Model or dataset
ShishirPatil/gorilla avatar
ShishirPatil/gorilla

Gorilla: an LLM tool-calling research stack, not a drop-in SDK

Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)

13,044 stars1,409 forksPythonApache-2.0

At a glance

What is it?
The ShishirPatil/gorilla repository bundles the Berkeley Function Calling Leaderboard, an inference path for Gorilla finetuned models, GoEx and Agent Arena. It is a research and evaluation project, and its README tells you almost nothing about running it yourself.
Who is it for?
Adopt Gorilla if you are evaluating function-calling behaviour across models, or you want the BFCL harness and GoEx in one place, and you are willing to read the subdirectory READMEs because the top-level one does not walk you through setup. Do not adopt it expecting a pip-installable tool-calling library with a stable API surface; the repository is a collection of research components, each with its own layout, and the top-level README documents none of them end to end.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 170 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Gorilla actually is, and who it is built for

The README states the goal plainly: "Given a natural language query, Gorilla comes up with the semantically- and syntactically- correct API to invoke." That is function calling, or tool calling, expressed as a research problem. The repository is the artifact around that problem rather than a single library. Its top-level entries are agent-arena/, berkeley-function-call-leaderboard/, data/, goex/, gorilla/, openfunctions/ and raft/, plus the usual .github/, .devcontainer/ and LICENSE. Each of those is a separate project with its own scope.

The audience follows from that layout. If you are choosing a model for an agent that must emit valid tool calls, the Berkeley Function Calling Leaderboard is the part you care about, and it has its own changelog and its own release cadence. If you are reproducing the original paper, you want gorilla/eval and the APIBench data under data/. If you are building runtime safety around generated actions, goex/ is the relevant directory. A team that just wants an OpenAI-compatible function-calling client is not the audience; nothing in the README describes that use.

The README also notes that the project has "served ~500k requests" since its initial release. Treat that as a statement about the hosted service the authors operate, not about the repository you would clone. Nothing in the README describes a self-hosted serving component.

How the pieces fit: BFCL, GoEx and the finetuned models

The mechanism differs per subdirectory, and the README describes them at different levels of detail.

The leaderboard is the most documented. BFCL V2 introduced enterprise-contributed data and real-world scenarios; V3 added a state-based evaluation system for multi-turn and multi-step function calling, testing sequential functions and service states; V4 Agentic, announced 2025-07-17, covers tool calling in agentic settings with web search plus multi-hop reasoning and error recovery, agent memory management, and format sensitivity. Format sensitivity is the interesting one: it means the same model is scored against prompt variations, so a model that only works with one tool-schema formatting is measured as weaker. Cost and latency metrics were added to the leaderboard on 2024-04-01, so the ranking is not accuracy alone.

GoEx is a different mechanism. The README calls it "a runtime for LLM-generated actions like code, API calls, and more", with "post-facto validation" for assessing actions after execution, plus "undo" and "damage confinement" abstractions. That is a containment model: the action runs, then validation and reversal logic decide what to do about it. It is not a pre-execution permission system, and the README does not claim it is.

The finetuned models are the third path. The README points at inference/README.md for the CLI interface to chat with Gorilla, and at Hugging Face under gorilla-llm for published weights. The original release was gorilla-7b-hf-delta-v0, and the README notes commercially usable Apache 2.0 licensed Gorilla models were released on 2023-06-06. The repository's own LICENSE is Apache-2.0.

Installing Gorilla: where the instructions actually live

The top-level README does not contain install steps. It links to subdirectories, and the setup you need depends on which one you want. For the inference path, the README points to inference/README.md for the CLI interface, so that file is the place to start rather than the root of the repository.

Clone the repository first, since every path below is relative to it:

bash
git clone https://github.com/ShishirPatil/gorilla.git
cd gorilla

From there, the README's own links tell you where to go. The finetuned models are published under the gorilla-llm organization on Hugging Face, and the README references the original gorilla-7b-hf-delta-v0 checkpoint. The evaluation code lives under gorilla/eval, and the APIBench data under data/.

For the leaderboard, the README links directly to the subdirectory and its changelog:

bash
cd berkeley-function-call-leaderboard
# follow this directory's own README and CHANGELOG

Do not expect a single pip install at the repository root to set up all of these. The README does not document a unified installer, a Python version requirement, or a dependency file at the top level, and it does not describe how to run a first evaluation end to end. Those details, if they exist, are in the subdirectory READMEs. That is the honest state of the documentation: the root file is an index, not a quickstart.

What the leaderboard does not tell you

A leaderboard is a measurement instrument, and its limits are the limits of the dataset behind it. BFCL V2 added enterprise-contributed data specifically because the earlier set was not representative enough; V3 added multi-turn and state-based scoring because single-call accuracy did not capture workflows; V4 added memory, web search and format sensitivity because agentic use exposed gaps. Each version exists because the previous one missed something. That is a healthy research trajectory and also a warning: a model ranked highly on one BFCL version has been measured against that version's assumptions.

The format sensitivity work is the sharpest example. If a model's score moves when the prompt format changes, then the number you quote depends on how your own application serializes tool schemas. BFCL can tell you a model is format-sensitive; it cannot tell you whether your format is the one that model handles well.

There is a second gap. The README describes GoEx as providing undo and damage confinement, but undo is only meaningful for actions that can be reversed. A sent email, a charged payment, or a deleted record in a system without soft deletes does not undo. The README does not enumerate which action classes are reversible, so the containment guarantee is bounded by whatever the underlying tool supports. Read goex/ before assuming it covers your write paths.

Compared with a general agent framework

The obvious comparison is a general-purpose agent framework, where tool calling is one component among planning, memory and orchestration. Gorilla inverts that. Its center of gravity is measurement: the leaderboard exists to score how well models emit correct calls, and the finetuned models and GoEx exist as artifacts around that question. A general framework assumes the model can call tools and spends its effort on control flow. Gorilla asks whether the model can call tools at all, and under what prompt formatting.

That difference shows up in what you get. A general framework hands you a runtime you embed in an application. Gorilla hands you an evaluation harness, a set of finetuned checkpoints, and a research runtime for validating and reversing generated actions. If your question is "which model should back our tool-calling layer, and does the answer change when we reformat the schema", Gorilla is aimed at you. If your question is "how do I wire three tools into a chat loop by Friday", the leaderboard will not answer it and the README will not show you how.

The repository is also not a single dependency. Adopting "Gorilla" in practice means adopting one subdirectory, and the maintenance you inherit is that subdirectory's maintenance. The root repository's last push was on 2026-04-13, but the release history shown is dominated by leaderboard versions: v1.1 in 2024-08, v1.2 in 2025-01, v1.3 in 2025-07. The cadence of the leaderboard is not the cadence of every other directory in the tree.

Licence and the cost of tracking a moving benchmark

The repository is Apache-2.0, and the README states that commercially usable Apache 2.0 licensed Gorilla models were released on 2023-06-06. That covers the code and those model releases; it does not automatically cover every asset in the tree, and the README does not break the licence down per subdirectory. If you vendor one directory, check the LICENSE at the root and confirm the files you copy carry the same terms. This is a description of what the repository states, not legal advice.

The upgrade cost is the more practical concern. BFCL has moved through V2, V3 and V4, and each version changed what is being measured, not just the scores. A ranking you cite from one version is not comparable to a ranking from another. If you pin your model choice to a BFCL result, record which version produced it, because the next release can reorder the table by changing the task rather than by changing the models. The changelog under berkeley-function-call-leaderboard/ is the file that tracks this, and it is the one to watch. Everything else in the repository can sit still for a year without affecting your decision.

Editorial conclusion

Adopt Gorilla if you are evaluating function-calling behaviour across models, or you want the BFCL harness and GoEx in one place, and you are willing to read the subdirectory READMEs because the top-level one does not walk you through setup. Do not adopt it expecting a pip-installable tool-calling library with a stable API surface; the repository is a collection of research components, each with its own layout, and the top-level README documents none of them end to end. Before committing, open berkeley-function-call-leaderboard/ and read its own README and CHANGELOG, check the gorilla/inference README for the model path you intend to run, and confirm the Apache-2.0 LICENSE covers the subdirectory you plan to vendor, since the repository mixes several projects under one tree.

Frequently asked questions

What does Gorilla do for LLM function calling?

The README states that Gorilla lets LLMs use tools by invoking APIs, producing a semantically and syntactically correct API call from a natural language query. The repository also ships the Berkeley Function Calling Leaderboard, which scores how well models perform that task.

How do I install Gorilla?

The top-level README gives no install steps. It links to subdirectories instead: inference/README.md for the CLI interface to chat with Gorilla, gorilla/eval for evaluation code, and berkeley-function-call-leaderboard/ for the leaderboard, each with its own setup.

Which models does the Gorilla repository publish?

The README points to the gorilla-llm organization on Hugging Face, and references the original gorilla-7b-hf-delta-v0 checkpoint. It also states that commercially usable Apache 2.0 licensed Gorilla models were released on 2023-06-06.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. ShishirPatil/gorilla on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/shishirpatil-gorilla.svg)](https://hysenlabs.com/projects/shishirpatil-gorilla)