# Functionary: A Deprecated Tool-Calling Language Model with vLLM and SGLang Serving

> Functionary was MeetKai's open-source language model specialized for tool and function calling, serving a JSON Schema-compatible interface through vLLM or SGLang. The repository is explicitly deprecated and no longer maintained.

**MeetKai/functionary** — Chat language model that can use tools and interpret the results

- Repository: https://github.com/MeetKai/functionary
- Stars: 1,595 · Forks: 118
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/meetkai-functionary

## What Functionary Was Designed to Do

Functionary is a language model that interprets and executes functions and plugins. The key design constraint is that the model decides when to call a function based on the conversation, whether to invoke functions in parallel or serially, and how to use the results. It does not call a function on every turn; it only triggers functions when the conversation requires them.

Function definitions are given as JSON Schema objects, following the same format as OpenAI GPT function calls. This made Functionary a drop-in candidate for teams that had already built OpenAI-compatible function schemas but wanted an open model they could self-host.

The README's deprecation notice is explicit: this repository reflects a very old snapshot of Functionary and does not represent the current state of the project. The code, models, and documentation are significantly out of date. Issues and pull requests may not be reviewed. The README does not link to a replacement repository from MeetKai, so teams that used Functionary should evaluate alternatives independently.

The project is MIT-licensed. The last push to the repository was on 2026-06-30.

## How the Server and JSON Schema Function Protocol Work

Functionary ships as a server process, not as a library you call from within your application. The server exposes an API compatible with OpenAI's function calling interface, so clients that already use the openai Python client can point the base URL at the Functionary server and switch model names.

Function definitions are passed as part of the request using JSON Schema syntax. The model receives the function signatures in its context and generates either a natural language response or a function call with arguments. If it generates a function call, the client is responsible for executing the function and sending the result back in the next message. The model then uses the result to continue the conversation or call another function.

The README describes support for parallel tool calls and serial tool calls. It also documents a code interpreter capability added in version 2.4: passing `{type: "code_interpreter"}` in the tools list enables a code execution mode. The changelog in the README shows the model has gone through versions 2.0 through 4.x, with a reasoning-capable variant (`functionary-v4r-small-preview`) added in December 2024.

## Deploying Functionary with vLLM or SGLang

Two serving backends are supported: vLLM and SGLang. Install the server dependencies for whichever backend you prefer:

```bash
pip install -e .[vllm]
```

or:

```bash
pip install -e .[sglang] --find-links https://flashinfer.ai/whl/cu124/torch2.5/flashinfer-python
```

To start the small model server with vLLM:

```bash
python3 server_vllm.py --model "meetkai/functionary-v4r-small-preview" --host 0.0.0.0 --port 8000 --max-model-len 8192
```

The medium models are larger: the README states that `functionary-medium-v3.1` requires four A6000s or two A100 80GB GPUs and uses tensor parallelism:

```bash
python server_vllm.py --model "meetkai/functionary-medium-v3.1" --host 0.0.0.0 --port 8000 --max-model-len 8192 --tensor-parallel-size 2
```

The README also documents a Text-Generation-Inference (TGI) server, a Modal.com serverless deployment option (`modal_server_vllm.py`), and LoRA adapter support. Dynamic LoRA loading is available through a `/v1/load_lora_adapter` endpoint in the vLLM server.

## Limitations: Deprecation, No Releases, and GPU Requirements

The most significant limitation is the explicit deprecation. The README states this is a reference-only snapshot that will not receive updates, bug fixes, or support. There are no GitHub releases in the repository. Model weights live on HuggingFace under the meetkai organization.

The medium models require hardware that many developers do not have available: four A6000s or two A100 80GB GPUs. The small preview model can fit on lesser hardware, but the README does not document specific GPU memory requirements for each model size.

The pyproject.toml shows the installable package (version 0.0.1) depends only on `jsonref`, `json_source_map`, and `PyYAML`; the actual inference dependencies (vllm or sglang) are installed through optional extras. This means the base package has minimal footprint, but the optional extras pull in large dependencies with specific CUDA version requirements.

The vLLM extra pins vllm to 0.8.2 for non-macOS platforms. The sglang extra pins sglang to 0.4.4.post1 and transformers to 4.48.3. These pins will conflict with environments that need newer versions.

## Functionary vs. Hosted Tool-Use APIs

The openai Python client with GPT-4o or GPT-4o-mini provides tool-calling through OpenAI's hosted service. The model, the serving infrastructure, and the updates are all managed by OpenAI. Functionary was the self-hosted counterpart: you run the model on your own GPUs, you own the data path, and you can modify the server code.

The trade-offs are clear. Functionary gives you data residency and no per-call API costs. The hosted service gives you current model improvements, reliability guarantees, and no GPU infrastructure to maintain. For the specific use case of function calling with Llama 3.1-based models, the Llama 3.1 family now includes native function calling support per Meta's documentation, making a separately trained function-calling model less necessary than it was when Functionary first launched.

The README's own changelog notes that `functionary-small-v3.1` uses Meta's original prompt template from the Llama 3.1 user-defined tool calling documentation, while `functionary-small-v3.2` uses MeetKai's own prompt template and is described as better. This versioning history shows the project tracked Llama 3.1's native tool support while adding its own modifications.

## Licence, Maintenance Status, and What Comes Next

Functionary is licensed under MIT, which permits unrestricted use, modification, and redistribution for both commercial and non-commercial purposes. The deprecation notice in the README does not add any licensing restriction; it only communicates that the maintainers will no longer act on issues or pull requests.

The last push was on 2026-06-30. The changelog documents changes up through December 2024, and the README's deprecation notice explains the current state: the repository reflects a very old snapshot.

Teams that are running Functionary servers today and want to migrate should consider that the vLLM and SGLang APIs are stable enough to swap models with minimal code changes. A migration to a different model served through the same vLLM infrastructure would require only a model path update in the startup command, assuming the new model exposes a compatible function calling interface.

## Conclusion

Functionary should not be chosen for new projects. The README opens with a deprecation notice stating this is a very old snapshot, significantly out of date, available for reference only, with no updates, bug fixes, or support going forward. For teams that ran earlier Functionary deployments, the vLLM and SGLang server scripts remain usable as-is, but the model weights and the codebase are frozen. The last push was on 2026-06-30. Any production tool-calling requirement is better served by a maintained alternative that receives ongoing security patches and model improvements.

## FAQ

### What is Functionary and what is it used for?

Functionary is a language model from MeetKai specialized for interpreting and executing functions and plugins. It receives function definitions as JSON Schema objects and decides when to call them, enabling LLM-powered workflows that invoke external tools without calling a function on every turn.

### Is Functionary still maintained?

No. The README explicitly states that Functionary is deprecated and no longer actively maintained. The code, models, and documentation are significantly out of date, and issues and pull requests may not be reviewed. The last push was on 2026-06-30.

### What GPU hardware does Functionary require?

The README states that the medium models (such as functionary-medium-v3.1) require four A6000s or two A100 80GB GPUs with tensor parallelism. The README does not document minimum GPU requirements for the small preview model.

## Sources

- [Issues](https://github.com/MeetKai/functionary/issues)
- [License: MIT](https://github.com/MeetKai/functionary/blob/main/LICENSE)
- [MeetKai/functionary on GitHub](https://github.com/MeetKai/functionary)
- [README](https://github.com/MeetKai/functionary/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meetkai-functionary
