CLI tool
kxxt/aspeak avatar
kxxt/aspeak

aspeak ships one crate as both a binary and a cdylib, with five ways to pick TLS

A simple text-to-speech client for Azure TTS API.

497 stars60 forksRustMIT

At a glance

What is it?
aspeak is a small Rust client for the Azure text-to-speech API, and almost all of its design decisions live in the Cargo feature graph rather than in the code. One crate compiles as both a command line binary and a cdylib loadable from Python, the two network modes are separate features behind a unified trait, and TLS is a five-way choice. The documentation is generated from a template, which is why the profile section reads the way it does.
Who is it for?
aspeak suits someone who wants Azure speech from a shell script, a Rust program, or a Python process without pulling in a vendor SDK, and the profile system is what makes it usable repeatedly rather than once. Three things to check before you commit.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 163 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One crate compiles two ways, and TLS is a five way choice

The Cargo manifest describes a package that is deliberately two artefacts.

The library target is built as both a `cdylib` and an `rlib`. The dynamic library is what makes a Python extension module possible at all, since CPython imports a shared object rather than a Rust library, and the Rust library form is what other Rust code links. A separate binary target carries a required feature, so the command line tool cannot be built without pulling in the audio and synthesizer stack.

The features are where the design decisions live. There is an `audio` feature that pulls in a playback library, so speaking without an output device is possible. There is a `python` feature that requires audio, adds the Python bindings, logging, an error reporting helper, and the synthesizers as a group. And there are two separate synthesizer features, one for the REST path and one for the WebSocket path, with a third unified feature providing a trait both implement, and a group feature that turns on all three.

Then there are five mutually exclusive TLS options: a default that selects native TLS, native TLS itself, a vendored native TLS build, and two rustls variants differing in whether roots come from the native store or from a bundled webpki set.

The release profile is set for a shipped binary with full link time optimisation, stripped symbols, and a single codegen unit. The edition is 2024 and the minimum Rust version is 1.85.

Version 6 changed the default transport

There is one behavioural change in this tool that will silently alter what your script produces, and it is stated at the top of the documentation as a note.

Starting from version 6.0.0, aspeak uses the REST interface to the Azure speech service by default. Previously the WebSocket interface was the primary path. If you want the older behaviour you can pass a mode flag when invoking it, or set a mode key in the authentication section of your profile to the websocket value.

Why it matters is the shape of the two paths rather than the transport itself. REST is a request and a response, which is simpler and easier to retry. WebSocket is a persistent connection that can stream, and streaming is what you want for long text. The examples directory shows the project treating them as peers: one example uses the REST synthesizer and another uses the WebSocket one, each in its simplest form.

So the default is now the boring option, and the streaming option is something you opt into per invocation or per profile.

The version boundary is worth internalising. If you have a profile written before version 6 with no mode set, upgrading changes the transport without changing your configuration, and any difference in latency or in how partial output appears is the transport change rather than anything you did.

The Python wheel is two builds and a merge script

Building the Python distribution is a three command sequence, and the shape of it explains the packaging problem.

You need maturin installed, then you build twice. The first build produces the Python extension, selecting the Python binding target and the Python interpreter, writing to one output directory. The second produces the binary flavour, selecting the binary binding target, writing to a different output directory.

bash
maturin build --release --strip -F python --bindings pyo3 --interpreter python --manifest-path Cargo.toml --out dist-pyo3
maturin build --release --strip --bindings bin -F binary --interpreter python --manifest-path Cargo.toml --out dist-bin
bash merge-wheel.bash

Then a shell script merges the two into one wheel file in the output directory.

So one Rust crate has to appear inside a Python wheel in two forms, and the merge step exists because a wheel carries a compiled extension under a specific platform tag while the merge combines them into something installable.

The distribution story around that has two gaps stated plainly. Prebuilt wheels are only available for the x86_64 architecture, and the source distribution has not been uploaded to the package index because of technical issues. Linux wheels are also absent because of manylinux compatibility issues, though they can still be built locally.

The practical consequence is that installing from the index is a good path on x86_64 and a dead end on an Apple Silicon machine or on Linux.

The README is generated, and the seam is visible

The documentation is not hand written, and the repository makes that explicit in two places.

There is a source file for the README, a script that generates the real one from it, and a make target that runs that script. The make file also has a clean target, and the clean target removes the generated README along with sample audio files.

So two things follow. The README is an output, and editing it directly will be overwritten the next time anyone runs the generation step. And running the clean target on a fresh clone deletes the README, which is an unusual thing for a clean step to do and worth knowing before you run it.

The generation is also visible in the rendered text. The section describing the profile file contains a sentence announcing that the default profile looks like a certain way, immediately followed by a sentence telling you to check the comments in the config file for the available options, and only then the configuration block itself. The prose and the sample have been spliced together by the template rather than written in sequence.

And the configuration sample is cut off partway through, in the middle of the authentication section's comments.

Neither is fatal, because the comments in the file are described as the authoritative reference. But it means the README is a rough index to the actual configuration rather than a complete description of it.

Authentication options go before the subcommand

The command line has an ordering rule that is easy to get wrong, and the documentation calls it out explicitly.

Authentication options must be placed before any subcommand. To use a subscription key against an official regional endpoint, the key and the region both go first and the text subcommand comes last.

If you are not using an official endpoint, there is an endpoint option in place of the region option.

Two later additions give you ways to avoid passing the secret at all. From version 5.2.0 the secrets can come from environment variables, one for a subscription key and one for an authorization token, so a secret can stay out of your shell history.

From version 4.3.0 there is proxy support, and it is narrower than the name suggests. Only plain HTTP and SOCKS5 proxies are supported, and the documentation says explicitly that HTTPS proxies are not yet supported. The proxy is passed as an option with a scheme and a host and port. The tool also respects the conventional HTTP proxy environment variable in either case, so an existing corporate proxy setting works without any flag.

For a client that runs in a locked down network, that gap matters: if your egress is an HTTPS proxy, the flag will not help you and the environment variable will not either.

Profiles are TOML, and verbosity above one needs a debug build

Profiles arrived in version 4 and are the reason this tool is usable in a script rather than only once.

A profile is a configuration file holding default values for the command line options, in TOML. Three subcommands manage it: one creates a default profile, one opens it in an editor, and one prints the path so you can find and edit it yourself if the editor subcommand misbehaves.

The default profile documents its own options in comments, and the first setting is output verbosity on a numeric scale. Zero is default and one is verbose. Two is debug and three or more is trace, and both of those levels are stated to work only on a debug build.

That is a small but real constraint. If you are chasing a failing call and reach for trace level in a release binary, the level you set is not available, so the diagnostic has to be reproduced from a debug build.

Two flags control profile selection at invocation time. One points at an alternate profile by path, for when you have a second account or a second region. The other disables profile loading entirely, which is what you want when you want one command to be hermetic rather than silently inheriting defaults from a file.

The authentication section carries the endpoint or region, the transport mode, and the key, each commented out by default.

Rate takes four different syntaxes

The prosody controls are the most elaborately specified part of the documentation, and the reason is that the service accepts four different notations and the client forwards all of them.

A bare float is interpreted as a percentage. A value of `0.5` becomes fifty percent.

A named keyword is passed through: extra slow, slow, medium, fast, extra fast, and default.

A percentage can be given directly, as `+10%`.

And a relative float with an `f` suffix is treated as a multiplier of the default rather than a percentage, following the service's own documentation. Under that reading `1f` changes nothing, `0.5f` halves the rate, and `3f` triples it.

So the same value can be written four ways, and the difference between `0.5` and `0.5f` is a factor of two rather than a rounding difference. That is a genuine trap: half speed and double speed look almost identical on the page.

The pitch control is documented in the same shape, with a float treated as a percentage, so `-0.5` becomes minus fifty percent, and its own set of named keywords beginning with extra low.

The documentation for pitch is cut off in the middle of that keyword list, so the full set is not readable there.

The newest tag is a year older than the newest commit

The repository state is worth reading before you pin anything.

The default branch is not master or main. It is named after the current major version. That is a deliberate choice for a tool where the transport default and the Python packaging both changed at a major boundary, and it means a clone gives you version 6 code by default.

The release history is short and oddly spaced. There is a patch release from October 2023, then nothing until March 2025, when a release candidate and the corresponding final release were published about a minute apart.

The last push to the branch is dated 2026-04-23, which is more than a year after the newest tag. So anyone installing from a release gets code from March 2025 while anyone building from the default branch gets something considerably newer, and nothing in the release list tells you which you have.

That gap shows up in the documentation too. The install command for the package index pins a specific version from the 6.0 line while the crate manifest is at 6.1.0. So following the documented install command gets you an older release than building from source, by design or by oversight.

The service itself has a free tier of half a million characters per month, which is enough to try the tool and not enough to build on.

Editorial conclusion

aspeak suits someone who wants Azure speech from a shell script, a Rust program, or a Python process without pulling in a vendor SDK, and the profile system is what makes it usable repeatedly rather than once. Three things to check before you commit. Decide the network mode first, because version 6 defaults to REST and the WebSocket path is opt in through a flag or a profile key. Pick your TLS feature deliberately, since there are five mutually exclusive options and the default pulls in native TLS, which matters on a host without OpenSSL development headers. And read the install section carefully rather than following it, because the PyPI command still pins an older version than the crate, and there is no source distribution on PyPI at all, so anything that is not a prebuilt x86_64 wheel has to be built from source. The free tier is half a million characters a month, which is generous for a script and useless for a product.

Frequently asked questions

Does aspeak use the WebSocket or REST interface by default?

REST. Starting from version 6.0.0 aspeak uses the Azure REST interface by default. To get the WebSocket interface you pass a mode flag when invoking it or set mode to websocket in the auth section of your profile.

How do I install aspeak?

The recommended route is to download a release archive from the repository and put the extracted binary anywhere on your PATH. Arch users can install aspeak-bin from the AUR, and for the command line only you can also use cargo install aspeak with the binary feature.

Can I install aspeak from PyPI on an ARM Mac or on Linux?

Not from prebuilt wheels. Only x86_64 wheels are available, no source distribution has been uploaded, and Linux wheels are absent because of manylinux compatibility issues. On other platforms you have to build the wheel yourself with maturin and then run the merge script.

Where do aspeak authentication options go on the command line?

Before any subcommand. You pass your subscription key and either a region for an official endpoint or an endpoint for a custom one, then the subcommand. From version 5.2.0 the key can also come from environment variables instead.

Does aspeak support HTTPS proxies?

No. Proxy support arrived in version 4.3.0 and covers only plain HTTP and SOCKS5 proxies; the documentation states that HTTPS proxy support is not there yet. The tool does respect the conventional HTTP proxy environment variable.

What is the difference between 0.5 and 0.5f for aspeak rate?

A bare float is treated as a percentage, so 0.5 becomes fifty percent. A float with an f suffix is a multiplier of the default, so 0.5f halves the rate, 1f changes nothing, and 3f triples it. Named keywords and explicit percentages like +10% are also accepted.

Official sources

  1. Issues
  2. kxxt/aspeak on GitHub
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kxxt-aspeak.svg)](https://hysenlabs.com/projects/kxxt-aspeak)