# autopreso: a voice-driven Excalidraw whiteboard that draws while you talk

> autopreso turns live speech into an Excalidraw scene through a local Express and WebSocket server, with a choice of OpenAI, Codex or Ollama as the drawing agent. It is a macOS-first alpha, and the README is candid about rough edges.

**kunchenguid/autopreso** — Realtime speech to presentation. Let the whiteboard whiteboard itself.

- Repository: https://github.com/kunchenguid/autopreso
- Stars: 446 · Forks: 64
- Language: JavaScript
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/kunchenguid-autopreso

## What autopreso solves, and who it is aimed at

The README opens with a grievance rather than a feature list: "You wanted to give the talk, not build the deck." That is the whole pitch. autopreso is for people who already think out loud in front of a whiteboard and who find the deck-building step a tax on the part they actually enjoy. You stage a few seed elements (the README suggests a title and an agenda), press Start Preso, and then talk. An agent edits the Excalidraw scene as your words arrive.

The target user is narrow in a useful way. You need a microphone, a browser, and either an OpenAI API key, a signed-in Codex CLI, or a local Ollama model. You also need macOS, or at least the willingness to run without the bundled transcription sidecars, because the two optional dependencies that ship Moonshine binaries are scoped to darwin-arm64 and darwin-x64. The platform badge in the README says macOS plainly. Nothing in the repository layout suggests a Windows or Linux build path for those sidecars.

It is not a slide tool. There is no export to PDF, no template gallery, no speaker notes. The output is an Excalidraw scene that the agent keeps rearranging, and the value is in the live performance rather than in an artifact you hand to someone afterward.

## The audio-to-canvas pipeline and its two modes

The README's diagram lays out the data flow in one line: microphone in the browser, 24kHz audio into a speech-to-text stage (Moonshine or the OpenAI Realtime WebSocket), text chunks into a whiteboard agent (OpenAI, Codex or Ollama), and tool calls out to the Excalidraw scene. The server is Express plus ws, bound to 127.0.0.1, so the browser and the agent talk to each other over a loopback WebSocket rather than over your LAN.

The mode switch matters more than it first appears. In staging mode the canvas is yours; you sketch the seed content client-side with no agent interference. In live mode the canvas is handed to the agent, and the README states that OpenAI Realtime transcription is biased toward your staging text and labels. That biasing is a real design decision: it means the seed content does double duty as both visual scaffolding and a vocabulary hint for the transcriber, which should reduce the chance that your own terminology gets mangled into something the agent then draws.

There is also a warmup loop. After you hit start, the agent primes itself against the staging content and your Agent instructions before it begins consuming transcripts, so the first sentence you speak does not hit a cold model. That is a small thing that says the author has actually presented with this.

## Installing autopreso and running a first session

The fastest path is npx, which needs no global install. The README's quick start shows the server booting and printing its address.

```bash
npx autopreso
# autopreso listening at http://127.0.0.1:3210
```

If you would rather keep it on the machine, install globally and run the binary. The package declares an engines field of node >=24, so check that first.

```bash
npm install -g autopreso
autopreso
```

Once the browser opens, the README's three-step sequence is the whole onboarding: drop reference materials onto the staging canvas, pick your microphone, transcription model, agent model and optional Agent instructions, then click Start Preso and talk. The status panel is where you change providers after the first run, because auto-detection only fires when no settings file exists.

Building from source is the path if you want to modify the agent. Note that npm start runs node ./src/cli.js directly, and there is a separate build:moonshine-sidecars script for the transcription binaries.

```bash
git clone https://github.com/kunchenguid/autopreso.git
cd autopreso
npm install
npm start
```

The CLI surface is deliberately tiny. The reference table lists exactly two commands, autopreso and autopreso -h, and two flags, --no-open and -h/--help. If you are running headless or on a remote box with port forwarding, --no-open is the one you want.

## Provider auto-detection is convenient once and confusing twice

On first run, autopreso inspects your environment and picks providers. The precedence table is explicit: Codex CLI auth wins over OLLAMA_MODEL, which wins over OPENAI_API_KEY. Transcription flips to OpenAI Realtime whenever an OpenAI key is present, and otherwise falls back to Moonshine medium on macOS. With nothing configured at all, you get OpenAI gpt-5.5 as the agent (which needs a key) and Moonshine medium for transcription.

The sharp edge is in the same paragraph: after first run, this auto-detection no longer applies. Provider environment variables only seed settings.json on the first launch; once the file exists they are ignored, and you must edit ~/.config/autopreso/settings.json or use the in-app panel. That is a defensible choice for reproducibility, but it produces a specific failure mode. You export a new OPENAI_API_KEY, restart, and nothing changes, because the old value is already persisted. The README does say this, but it is the kind of sentence people skim past on the way to the quick start.

Settings that do persist include models, API keys, STT engine choices and Agent instructions. Agent instructions can run to 100,000 characters and take effect on the next Start Preso, not mid-session, which is worth knowing if you are iterating on prompts while presenting.

## Cost tracking, and where its numbers stop being trustworthy

The live Session cost card estimates agent token costs and OpenAI Realtime audio costs for the current presentation, and it resets on Start Preso or a session reset. The README is unusually honest about the boundaries of that estimate. OpenAI prices come from a built-in May 2026 rate table, which means the numbers drift as OpenAI changes pricing and the table is not updated. Local providers always show $0.0000, which is correct for money but hides the electricity and latency cost of running a model locally. Codex shows token volume instead of dollars because it routes through your subscription. Unknown models show n/a.

So the cost card is a rough instrument, not an invoice. If you are presenting for an hour on OpenAI Realtime, treat the figure as an order-of-magnitude signal. The fact that the rate table is baked into the release rather than fetched is a trade-off: no network dependency and no surprise, but also no automatic correction when prices move.

## The alpha warning is not decoration

The README carries a warning block that says autopreso is in alpha and under active development, and asks you to expect rough edges, breaking changes, and "the occasional weird drawing." The last push was on 2026-08-23, and the most recent release, autopreso-v0.1.8, landed on 2026-07-31. The version number is still 0.1.x. This is early software, and the release history shows the pace of a project that is still finding its shape rather than one that has settled.

The practical limitations follow from the architecture. The server binds to 127.0.0.1 only, so there is no built-in way to let a remote collaborator watch the canvas or to run the agent on one machine and the browser on another without your own tunneling. The bundled Moonshine transcription binaries are optional dependencies scoped to two macOS architectures, so on any other platform you are relying on the OpenAI Realtime path, which means audio leaves your machine. The "can run locally" claim in the feature list is real, but it requires Moonshine for transcription and Ollama for the agent together; running one locally and the other remotely gives you a partial local setup, not a private one. And because the agent edits a live scene, a bad transcription does not produce a wrong bullet point, it produces a wrong drawing that you then have to talk your way out of or undo by hand.

If you need a stable, cross-platform deck builder with a predictable output file, this is the wrong tool today.

## How autopreso differs from a generic Excalidraw plus LLM script

The obvious alternative is wiring an Excalidraw canvas to a model yourself: capture audio, transcribe it, prompt a model, apply the resulting element mutations. That is essentially what autopreso does, and the repository is small enough that the comparison is fair. The difference is in the parts you would have to rediscover. autopreso already handles the staging-to-live handoff, biases the transcriber toward your seed vocabulary, warms the agent before the first transcript, persists provider settings across restarts, and tracks session cost with a rate table. It also ships the Moonshine sidecar packaging, which is the least interesting part to build and the most annoying to get right.

A second alternative is the manual route: draw on the whiteboard yourself while you talk. That is free, works on every platform, and never mishears you. autopreso's bet is that the cognitive load of drawing while speaking is high enough that offloading it to an agent is worth the accuracy risk. Whether that bet pays off depends on your material. A talk about system architecture, where boxes and arrows follow predictably from what you say, suits it. A talk where the visual is the argument, not a restatement of it, does not.

## Licence, maintenance and what upgrading costs you

autopreso is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive licence with few obligations, but it says nothing about the models you point it at: your OpenAI usage is governed by OpenAI's terms, and Ollama models carry their own licences. The MIT grant covers the autopreso code, not the weights or the API.

Upgrade cost is currently low in absolute terms and high in uncertainty. The package is at 0.1.8 and the README warns about breaking changes, so pinning a version is reasonable if you depend on the settings file format or the CLI flags. The settings file at ~/.config/autopreso/settings.json is the main piece of state to back up before upgrading, since it holds your API keys and provider choices. The CLI has only two commands and two flags, so there is little surface to break there; the risk sits in the settings schema and in the provider auto-detection rules, which could shift between releases. Release automation runs through release-please, and the manifest file at the repository root tracks versions, which suggests releases are cut deliberately rather than continuously.

## Conclusion

Adopt autopreso if you present from a whiteboard rather than slides and you are willing to run an alpha on macOS with a Node 24 toolchain. Skip it if you need a stable deck builder, a Windows or Linux target, or a tool that never sends audio off your machine without you configuring that yourself. Before you commit, verify that a Moonshine sidecar exists for your CPU architecture, that your Codex or OpenAI credentials are the ones you want the auto-detection to pick, and that the settings file at ~/.config/autopreso/settings.json holds the providers you actually intend to use, because auto-detection only runs on first launch.

## FAQ

### Does autopreso work on Windows or Linux?

The README badges the platform as macOS, and the two bundled Moonshine transcription packages are optional dependencies for darwin-arm64 and darwin-x64 only. On other platforms you would need the OpenAI Realtime transcription path, which sends audio to OpenAI.

### Can I run autopreso entirely locally without sending audio anywhere?

Yes, according to the README, if you use Moonshine for transcription and Ollama for the agent together. That combination requires macOS for the Moonshine sidecar, and local providers show $0.0000 on the session cost card.

### Why did my new API key not take effect after restarting autopreso?

Provider environment variables only seed ~/.config/autopreso/settings.json on the first run. Once that file exists, the variables are ignored, so you have to edit the file or change providers in the in-app status panel.

## Sources

- [Issues](https://github.com/kunchenguid/autopreso/issues)
- [kunchenguid/autopreso on GitHub](https://github.com/kunchenguid/autopreso)
- [License: MIT](https://github.com/kunchenguid/autopreso/blob/main/LICENSE)
- [README](https://github.com/kunchenguid/autopreso/blob/main/README.md)
- [Releases](https://github.com/kunchenguid/autopreso/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kunchenguid-autopreso
