# Sayna builds with no features on by default and ships an image with all of them

> A Rust voice server that puts one interface in front of Deepgram, ElevenLabs, Google Cloud, and Azure for speech to text and text to speech over WebSocket and REST. The interesting parts are the manifest, the feature flags, and which endpoint is left unauthenticated.

**SaynaAI/sayna** — Sayna is a unified Voice Layer for AI Agents with a seemless integration to an existing agentic frameworks

- Repository: https://github.com/SaynaAI/sayna
- Website: https://sayna.ai
- Stars: 313 · Forks: 45
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/saynaai-sayna

## Cargo.toml says 0.1.15 while the newest release tag is v0.1.16

The manifest and the release feed disagree by one version. The newest published releases are v0.1.14, v0.1.15, and v0.1.16, with v0.1.16 dated 2026-05-25, while the version field in Cargo.toml still reads `0.1.15`. The default branch was last pushed on 2026-05-31, so this is not an unbuilt tag either.

How this reaches you is the other half of the story. The manifest sets `publish = false` with a comment saying binary-only distribution via GitHub releases, so the crate is not on crates.io and you install a binary, not a library. That makes the manifest version a build-time string rather than the thing a user resolves, which is why a one-version drift is survivable.

The release machinery is conventional-commit based. `cliff.toml` configures git-cliff for changelog generation and `release-plz.toml` handles the release, with the generated CHANGELOG.md sitting at the root. The default branch is `master`.

So if you pin anything, pin the release tag you downloaded, not the version in the source tree.

## No features are on by default and the image turns all of them on

`default = []` in the features table means a plain `cargo build` gives you none of the optional machinery. There are three feature groups and each one pulls a real dependency tree.

`stt-vad` adds ONNX Runtime bindings through `ort`, plus `ndarray` and `rustfft`, and it is what provides Silero voice activity detection with machine-learned end-of-turn detection. `noise-filter` is the heavier one, bringing in `deep_filter` and four `tract` crates for the DeepFilterNet path along with `num_cpus`. `openapi` adds `utoipa`.

Now look at the Dockerfile. It declares `ARG CARGO_BUILD_FEATURES="--all-features"` and uses that argument when it builds, so the published image is compiled with all three groups enabled.

That is the discrepancy worth internalizing before you debug anything: the binary you build from source with `cargo build` and the binary in `saynaai/sayna` are not the same program. If you are not using Docker and you skipped the features, you have no voice activity detection at all, which changes turn-taking behaviour rather than just adding a feature.

## The image pins Rust 1.95.0 while the build instructions ask for 1.88.0

The development prerequisites say Rust 1.88.0 or later, with ONNX Runtime described as optional and needed only for the `stt-vad` feature. The Dockerfile declares `ARG RUST_VERSION=1.95.0` and builds on `rust:1.95.0-slim-bookworm`. The crate uses edition 2024.

The lower documentation floor is the more useful number for you, since it is what the project claims to support, but the higher number is what the image is actually compiled with.

The rest of the image explains why the build is heavier than a typical Rust service. It installs clang, cmake, pkg-config, libssl-dev, and libzstd-dev, then a block of graphics libraries including libva, libdrm, libglib2.0, libgbm, libx11, libxext, libxrandr, libxcomposite, libxdamage, and libxfixes, because webrtc-sys needs them. The apt sources are rewritten with a mirror fallback for debian-security, with a comment tying it to ARM64 compatibility, and the ONNX Runtime download is switched on `TARGETARCH`.

The pattern repeats for caching: the compose file mounts a named volume at `/data/cache` and sets `CACHE_PATH` to match, so a cached artefact survives container replacement rather than living inside the image layer.

## One POST endpoint is deliberately left unauthenticated

Authentication is optional and off unless you set `AUTH_REQUIRED=true`. When it is on, token validation is delegated to an external service at `AUTH_SERVICE_URL`, and the signature key path plus a five second timeout are configured alongside it. The signing keys are RSA, generated with a 2048 bit private key and a public half extracted to hand to the auth service:

```bash
openssl genrsa -out auth_private_key.pem 2048
openssl rsa -in auth_private_key.pem -pubout -out auth_public_key.pem
```

A token is accepted either as `Authorization: Bearer <token>` or as a `?api_key=<token>` query parameter. The second form is the convenient one and the one that ends up in access logs and shell history.

The exception is `POST /livekit/webhook`. It is documented as unauthenticated because LiveKit calls it, and it validates requests with LiveKit's own JWT signature mechanism instead. It is also the endpoint that logs SIP-related attributes, which is there for phone call troubleshooting and is the reason that endpoint carries more metadata than the others.

## Per request provider credentials are all or nothing

The WebSocket config message carries `stt_config` and `tts_config` blocks, and each one may include its own `auth` object. That object is optional. If it is omitted, or sent as an empty object, Sayna falls back to the credentials configured on the server. If it is provided, it has to be complete for that provider.

There are three shapes. Plain API key providers take `{ "api_key": "..." }`. Google Cloud takes either `{ "credentials": "/path/to/creds.json" }` or the service account JSON inline. Azure Speech takes `{ "api_key": "...", "region": "eastus" }`, so the region travels with the key per request rather than coming from the server.

A partially filled auth object does not merge with the server configuration, which is the sharp edge here. Passing a key for a provider that also needs a region, or naming a credentials file that only exists on the server, silently falls into the incomplete case rather than raising a configuration error at startup.

The rest of the config message shows the sampling asymmetry. STT is configured for 16000 with `linear16` encoding, while TTS is configured for 44100 with an mp3 output format, so the two halves of a conversation do not share a sample rate.

## TLS is rustls everywhere so the binary can cross-compile to musl

The dependency table is organised by role, and the pattern running through it is avoiding a system OpenSSL. The HTTP client is declared with `default-features = false` and the `rustls-tls`, `stream`, and `http2` features, with a comment saying this is important for cross-compilation to musl targets. `Cross.toml` at the repository root is the tool that would do that work.

The same choice is repeated three more times. The TLS crate itself is `rustls` with default features off and the `ring` provider. The JSON Web Token library uses the `rust_crypto` feature rather than a system crypto backend. The WebSocket client uses `rustls-tls-webpki-roots`, which bundles the root store instead of reading the system one.

The cost of that consistency is a longer build: the ring provider and the ONNX runtime are both compiled from source in the Docker image. What you get in return is a static binary that does not need OpenSSL present on the target, which is what makes the Cross.toml workflow plausible rather than aspirational.

The server side is axum with the `ws` feature, tokio with `full`, and tower-http with trace enabled.

## Audio-disabled mode exists so the control plane can be tested without keys

You can run the server with no provider credentials at all. The trick is a WebSocket configuration message that turns audio off:

```json
{
  "type": "config",
  "audio": false
}
```

The stated uses are local development and testing, user interface work without audio processing, exercising WebSocket message flows, and debugging non-audio features. In other words, the config plane and the routing are testable without a single provider key, which is what makes the server pleasant to develop against.

The environment file shows the shape of the credential set. `DEEPGRAM_API_KEY`, `PORT`, and `ELEVENLABS_API_KEY` are set, and the Azure pair is commented out with a warning that the key must match the region or you get 401 errors. Google Cloud has no entry at all, which matches the fact that its credentials are a file path or a service account document rather than a key.

One mismatch to notice between the docs and the compose file: the compose example passes Deepgram and ElevenLabs keys plus a cache path on a named volume at `/data/cache`, while the environment example documents Azure as well. Compose is a starting point, not the full key list.

## Conclusion

Sayna fits if you want one voice interface across several providers, you need the WebSocket control plane, and you are willing to pick your own feature set rather than take a default. Read the manifest before you start, because `default = []` means a local `cargo build` has no voice activity detection, no noise filtering, and no OpenAPI surface, while the published container has all three, and a bug you reproduce locally may simply not exist in the image you deployed. Check that your Azure region matches the key, since a mismatch returns a 401 the documentation calls out explicitly. Decide early whether the query-parameter token form is acceptable in your environment. And if you rely on the LiveKit webhook, remember it is the one unauthenticated POST and leans entirely on signature verification.

## FAQ

### What is Sayna?

It is a real-time voice processing server written in Rust that exposes unified speech-to-text and text-to-speech through WebSocket and REST. Providers are pluggable through a trait-based abstraction and Deepgram, ElevenLabs, Google Cloud, and Microsoft Azure are supported.

### Which speech providers does Sayna support?

Deepgram and ElevenLabs for both STT and TTS, Google Cloud with WaveNet, Neural2, and Studio voices, and Microsoft Azure with 400+ neural voices across 140+ languages. Adding one means implementing the provider trait rather than editing a config file.

### How do I run Sayna without any API keys?

Use audio-disabled mode. Send a WebSocket configuration message with `audio: false`, which exercises the control plane without initializing STT or TTS. It is meant for local development, user interface work, testing message flows, and debugging non-audio features.

### Does Sayna need authentication?

Not by default. Set `AUTH_REQUIRED=true` to require a token on the protected endpoints, validated by an external service at `AUTH_SERVICE_URL` using an RSA signing key. The token goes in an `Authorization: Bearer` header or in an `api_key` query parameter.

### What do the Sayna cargo features enable?

`stt-vad` enables Silero voice activity detection with end-of-turn detection and pulls in ONNX Runtime bindings. `noise-filter` enables DeepFilterNet. `openapi` adds the documentation generator. The default feature set is empty, while the published Docker image builds with all features.

## Sources

- [License: Apache-2.0](https://github.com/SaynaAI/sayna/blob/master/LICENSE)
- [Project website](https://sayna.ai)
- [README](https://github.com/SaynaAI/sayna/blob/master/README.md)
- [Releases](https://github.com/SaynaAI/sayna/releases)
- [SaynaAI/sayna on GitHub](https://github.com/SaynaAI/sayna)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/saynaai-sayna
