ElatoAI: Realtime Voice AI on ESP32 with 100-Plus Model Providers
Realtime Voice AI with 100+ Models on Arduino ESP32 with Secure Websockets and Edge Functions for AI Companions, and Devices
At a glance
- What is it?
- ElatoAI is an open-source platform that connects an Arduino ESP32 microcontroller to cloud-hosted voice AI models via WebSocket edge functions, enabling conversational AI in physical devices for over 20 minutes without interruption. It covers the full stack from ESP32 firmware to a Next.js control app, supporting providers from OpenAI Realtime to Hume AI EVI-4.
- Who is it for?
- Hardware builders and developers who want to put a real-time conversational AI into a physical device will find ElatoAI's full-stack approach, from ESP32 firmware through edge functions to a web management app, covers most of the infrastructure work. The platform requires cloud API credentials for every voice pipeline, so there is no fully offline operation unless the local-ai-toys companion project is used.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ElatoAI Is and the Problem It Solves
Building a voice AI device on an ESP32 requires solving multiple separate problems: getting audio into and out of the microcontroller with acceptable quality, managing the WebSocket connection to a cloud AI API, handling voice activity detection to know when a human has finished speaking, maintaining conversation context across multiple turns, and providing a way to manage the device remotely.
ElatoAI packages solutions to all of these into a single open-source repository. The project grew out of a Kickstarter campaign for a voice AI toy product and is documented at elatoai.com, with a reference in the README to a featured example in the OpenAI Cookbook. The README targets hardware builders, makers, and developers who want to add voice AI capabilities to a physical device without building the full infrastructure from scratch.
The README states the platform supports over 20-minute uninterrupted conversations globally, which is relevant because maintaining a long-lived WebSocket connection across a consumer network to a geographically distributed edge function is not trivial. The Deno Edge Function and Cloudflare Workers server options are specifically designed to handle this with low latency.
Three-Component Architecture: ESP32, Edge, and Frontend
ElatoAI is organized into three distinct components that each live in a separate directory in the repository.
The ESP32 firmware, in firmware-arduino/, is built with the Arduino framework and managed via PlatformIO. It handles microphone input with Opus audio compression, WebSocket communication to the edge server, speaker output, button and touch sensor controls, Wi-Fi management via a captive portal, and over-the-air firmware updates. The README notes that the ESP32 does not require PSRAM for the speech-to-speech pipeline.
The edge server, in the server/ directory, handles the WebSocket connections from the ESP32 and proxies requests to the AI provider APIs. There are three server implementation options: Deno Edge Functions for low-latency global deployment, Cloudflare Workers for scale and their built-in Workers AI provider support, and a FastAPI server for teams who prefer Python. The FastAPI server integrates with Pipecat, which the README says supports over 100 speech-to-text, LLM, and text-to-speech pipeline combinations as of April 2026.
The frontend, in frontend-nextjs/, is a Next.js application deployed to Vercel. It provides the control interface for creating AI agent configurations, managing device registrations, viewing conversation transcripts, and adjusting voice settings including pitch factor for cartoon-like voices.
Voice AI Providers Supported on the Deno and Cloudflare Paths
The Deno Edge server path supports six speech-to-speech providers, each in its own model subdirectory. OpenAI's Realtime API and Gemini's Live API are the two largest providers documented. xAI's Grok Voice Agent API, Eleven Labs Conversational AI Agents, Hume AI EVI-4, and Boson Higgs Realtime are also implemented.
The README notes that OpenAI, Gemini, xAI, Hume AI, and EVI-4 demos are available as linked YouTube videos, which gives a practical reference for what the audio quality looks like on each provider before committing to a hardware build.
The Cloudflare Workers path takes a different approach. Rather than implementing specific integrations with each provider, it routes through Cloudflare's Workers AI, which in turn provides access to over 80 LLM choices, over 10 text-to-speech models including Deepgram and MeloTTS, and 5 speech-to-text models including Whisper and Deepgram. Under this path, only an LLM API key is needed: Cloudflare's platform handles the STT and TTS layers natively.
The FastAPI server option connects to Pipecat, which supports a different set of provider combinations. Choosing between Deno, Cloudflare Workers, and FastAPI depends on which providers are needed and whether the team prefers TypeScript or Python for the server layer.
Key Features in the Firmware and Application Layer
The README lists 24 features across the platform. Several are notable for hardware builders.
Server VAD (voice activity detection) turn detection handles conversation flow: the system detects when the user has stopped speaking and transitions to AI response without requiring a physical button press. This is distinct from button-based or touch-based triggering, which are also supported as alternative input methods.
Opus audio compression reduces the bandwidth required for continuous audio streaming over Wi-Fi, which matters for devices deployed on slower networks or connecting to distant edge servers.
OTA (over-the-air) updates allow firmware changes to be pushed to deployed devices from the NextJS frontend, avoiding the need to physically connect each device for updates.
Tool calling enables the ESP32 firmware to invoke server-side functions, allowing the voice AI agent to take actions beyond conversation: a user could ask the device to check external data or trigger integrations that are implemented as tools on the edge function.
Conversation history is stored in Supabase, making past transcripts available through the frontend app. Volume control, pitch factor for voice modulation, and a factory reset option are all accessible from the NextJS frontend without physical interaction with the device.
DIY Hardware and Network Requirements
The README includes a DIY hardware design section with a reference to schematics and hardware layout, though the specific content was not visible in the excerpt. The project page at elatoai.com links to a Kickstarter campaign for a ready-made devkit, suggesting a hardware product is available for builders who do not want to source and assemble components individually.
The README does not specify minimum Wi-Fi speed requirements, but the combination of Opus-compressed audio and a WebSocket connection to a Deno Edge Function implies any stable broadband connection should be sufficient. The README explicitly notes that all data stays with the user in the local-first and privacy-first framing in the feature list, though in practice the audio is sent to cloud provider APIs for processing.
The Wi-Fi management with captive portal feature allows a user to configure the device's network credentials without a separate programming tool: the ESP32 broadcasts a hotspot, the user connects to it and selects their Wi-Fi network through a browser, and the credentials are saved to the device.
Limitations and Cases Where ElatoAI Is Not the Right Fit
Every voice pipeline in ElatoAI's main branch requires active cloud API credentials. OpenAI Realtime, Gemini Live, xAI, Eleven Labs, Hume AI, and the Cloudflare Workers AI path all involve sending audio data to cloud services and receiving responses over the internet. The README mentions a separate companion repository called local-ai-toys that supports local AI models on a Raspberry Pi using MLX, covering models like Qwen and Mistral, but that is a different project with different hardware requirements.
The Firebase or Supabase dependency for conversation history and device management means teams deploying ElatoAI at scale need to provision a Supabase instance, handle authentication, and maintain the database.
Builders who need a voice AI device that works without internet access, or who have strict data residency requirements that prevent audio from leaving their network, will find ElatoAI's standard configuration unsuitable. The local-ai-toys companion is an alternative for the former, but it requires Pi hardware and a local model, not an ESP32.
Starmoon AI is another open-source ESP32-based voice AI companion project. Like ElatoAI, Starmoon AI connects the microcontroller to cloud AI providers via a server layer. The two projects target the same hardware category and maker use case but represent independent implementations with different provider integrations and application layers.
Editorial conclusion
Hardware builders and developers who want to put a real-time conversational AI into a physical device will find ElatoAI's full-stack approach, from ESP32 firmware through edge functions to a web management app, covers most of the infrastructure work. The platform requires cloud API credentials for every voice pipeline, so there is no fully offline operation unless the local-ai-toys companion project is used. Before building hardware, verify the license terms in the repository's LICENSE file: the license type is listed as NOASSERTION by GitHub's automated detection.
Frequently asked questions
What hardware does ElatoAI require to run?
ElatoAI requires an ESP32 microcontroller built with the Arduino framework via PlatformIO. The README notes that PSRAM is not required for the speech-to-speech AI pipeline. A DIY hardware design section and a Kickstarter devkit option are referenced in the README and at elatoai.com.
Does ElatoAI work without an internet connection?
No. All voice pipelines in the main ElatoAI repository require active cloud API credentials from providers such as OpenAI, Gemini, or Cloudflare Workers AI. A separate companion repository called local-ai-toys supports local AI models on Raspberry Pi hardware, but that uses different hardware than the ESP32.
Which AI providers does ElatoAI support?
The Deno Edge path supports OpenAI Realtime API, Gemini Live API, xAI Grok Voice Agent API, Eleven Labs Conversational AI Agents, Hume AI EVI-4, and Boson Higgs Realtime. The Cloudflare Workers path provides access to 80-plus LLM choices and separate STT and TTS models through Cloudflare Workers AI. The FastAPI server integrates with Pipecat, which supports over 100 pipeline combinations.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/akdeb-elatoai)