Library / SDK
earlephilhower/ESP8266Audio avatar
earlephilhower/ESP8266Audio

ESP8266Audio: MP3, FLAC, MIDI and MOD playback on ESP8266, ESP32 and Pico

Arduino library to play MOD, WAV, FLAC, MIDI, RTTTL, OGG/Opus, MP3, and AAC files on I2S DACs or with a software emulated delta-sigma DAC on the ESP8266 and ESP32 and Pico

2,401 stars471 forksCGPL-3.0

At a glance

What is it?
ESP8266Audio is an Arduino library that decodes MOD, WAV, MP3, FLAC, MIDI, AAC, RTTTL and OGG/Opus and pushes PCM to an I2S DAC or a software delta-sigma DAC. It is a good fit for microcontrollers with a speaker attached and a bad fit for anything that needs glitch-free audio while your loop() is busy.
Who is it for?
Adopt ESP8266Audio if you are writing Arduino sketches for an ESP8266, ESP32 or Pico and you want an MP3, MOD, MIDI or FLAC file to come out of a DAC without writing a decoder. Do not adopt it if your loop() blocks for tens of milliseconds at a time, if you need AAC on an ESP8266 with SBR, or if you are building a commercial product around AAC without checking the Helix RSPL terms and the Via Licensing requirement the README names.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 59 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ESP8266Audio actually replaces

On a microcontroller, playing an MP3 from flash normally means porting a decoder, wiring it to a DMA-capable output, and handling the fact that the file may arrive over SPIFFS, LittleFS, SD, PROGMEM or an HTTP socket. ESP8266Audio splits that into three objects per stream, and the split is the whole design. An AudioFileSource produces bytes, an AudioGenerator decodes them, and an AudioOutput consumes PCM. The README describes exactly this: "Create an AudioInputXXX source pointing to your input file, an AudioOutputXXX sink ... and an AudioGeneratorXXX to actually take that input and decode it and send to the output."

The intended audience is small. Someone building an alarm clock, a model train sound unit, a web radio or a talking toy on an ESP8266 or ESP32. The README's own list of projects is all of that shape: an MP3 alarm clock, a word clock with an MP3 alarm, an MQTT model train controller, a light-and-sound device. Nobody is building a streaming appliance with this. It is for sketches where one file or one HTTP stream has to make noise.

The loop() contract and why audio hiccups

The mechanism is cooperative, not interrupt-driven. After you construct the three objects, you call AudioGeneratorXXX::loop() from your own loop(). That call reads as much of the file as the output needs, fills the I2S buffers, and returns immediately. The README is explicit about the consequence: "Since this is not interrupt driven, if you have large delay()s in your code, you may end up with hiccups in playback. Either break large delays into very small ones with calls to AudioGenerator::loop(), or reduce the sampling rate to require fewer samples per second."

That is the central trade-off of the library and it is worth stating plainly. You get a simple, portable API that works the same on ESP8266, ESP32 and RP2040. You pay for it with a hard coupling between your application's timing and the audio stream. Any blocking call, a long sensor read, a synchronous HTTP request, a delay() of a few hundred milliseconds, shows up as an audible gap. The README also points ESP32 and Pico users at a separate library, BackgroundAudio, which it says "can provide a simpler usage model and better results and performance on the Pico by using an interrupt-based, frame-aligned output model." That recommendation comes from the author of this library. Treat it as a signal about where the loop-driven model stops being the right answer.

Installing ESP8266Audio and playing an MP3 from SPIFFS

The README gives a git clone into the Arduino libraries directory rather than a Library Manager install. For ESP8266 it also asks for two IDE settings: lwIP Variant set to v1.4 Open Source or V2 Higher Bandwidth, and CPU Frequency set to 160MHz. The prerequisites section asks for Arduino ESP8266 core 2.6.3 or later, or the latest ESP32 SDK from Espressif.

sh
mkdir -p ~/Arduino/libraries
cd ~/Arduino/libraries
git clone https://github.com/earlephilhower/ESP8266Audio

After the clone, restart the IDE so the library is indexed. The README's example constructs an MP3 generator, a SPIFFS file source and a no-DAC I2S output, which is the software delta-sigma path rather than a real I2S DAC:

cpp
#include <Arduino.h>
#include "AudioFileSourceSPIFFS.h"
#include "AudioGeneratorMP3.h"
#include "AudioOutputI2SNoDAC.h"

AudioGeneratorMP3 *mp3;
AudioFileSourceSPIFFS *file;
AudioOutputI2SNoDAC *out;
void setup()
{
  Serial.begin(115200);
  delay(1000);
  SPIFFS.begin();
  file = new AudioFileSourceSPIFFS("/jamonit.mp3");

The README snippet stops there, so the missing pieces are the generator and output construction and the loop() call. What the README does state is the calling pattern: you must call AudioGeneratorXXX::loop() from inside your own main loop() one or more times. If you call it once per iteration and keep the rest of the iteration short, playback continues. If setup() or loop() blocks, it does not.

The examples directory is the practical starting point. It contains PlayMP3FromSPIFFS, StreamMP3FromHTTP, WebRadio, PlayFLAC-SD-SPDIF, PlayMIDIFromROM, PlayRTTTLToI2SDAC, PlayOpusFromLittleFS and others. Pick the one matching your storage and output combination rather than assembling the objects from scratch.

PlatformIO on ESP32 needs a different platform package

This is the most likely place to lose an afternoon. The README states that Espressif discontinued official PlatformIO integration of their Arduino core several releases ago, that the platformio/framework-arduinoespressif32 package is at IDF 4.x, and that this library needs IDF 5.x and the new I2S APIs. The failure is a compile error naming a missing header:

ini
; Pioarduino Arduino-ESP32 Latest
platform = https://github.com/pioarduino/platform-espressif32/releases/download/stable/platform-espressif32.zip

The README's fix is to point platform in platformio.ini at the community pioarduino core, which it describes as built from the current Espressif Arduino with IDF 5.x. The error the README quotes is `cannot open source file "driver/i2s_std.h"`. If you see that, you are on the old framework, not on a bug in your sketch.

Note what this costs you. You are now depending on a community redistribution of the Espressif core rather than the one PlatformIO ships. That is a supply chain decision, not just a config line, and it is worth recording in your project notes.

Codec licensing is not uniform across the decoders

The library is GPL-3.0, and the README opens with a disclaimer that all the code is released under the GPL and used at your own risk. That much is simple. What is not simple is that the decoders inside it come from several upstreams with different terms, and the README lists them.

MP3 and MOD come from libMAD and StellarPlayer. MIDI is a port of MIDITONES combined with a memory-optimized TinySoundFont. Opus comes from Xiph.org under the Xiph license with patents described in src/{opusfile,libogg,libopus}/COPYING. The one that matters commercially is AAC: the README states the AAC decode code is from the Helix project under RealNetwork's RSPL license, and adds that "for commercial use you're still going to need the usual AAC licensing from Via Licensing." If your product ships AAC, the GPL-3.0 on the library is not the only term you are agreeing to. Read the COPYING files under src/ before you commit to a codec, and take your own advice on licensing rather than mine.

There is also a capability split worth knowing before you pick AAC. The README states that AAC-SBR, which many web radio stations use to cut bandwidth, is supported on the ESP32 but not on the ESP8266, because the ESP8266 lacks onboard RAM. An ESP8266 web radio pointed at an SBR stream will not decode it.

Where ESP8266Audio is the wrong tool

Three cases stand out from what the README and the repository layout say.

First, anything with a busy main loop. If your sketch spends 200ms talking to a sensor, a display library or a blocking network call, the loop-driven model breaks and you get gaps. The README's own remedies are to shorten delays or lower the sample rate, and neither is free: lowering the sample rate degrades the audio, and shortening delays means restructuring application code around the audio library rather than the other way round.

Second, ESP32 and Pico projects that want the better output model. The README recommends BackgroundAudio for those targets, describing it as interrupt-based and frame-aligned with better results and performance on the Pico. Choosing ESP8266Audio there means choosing the older model deliberately.

Third, ESP8266 projects that need AAC. The README says SBR is unsupported on that chip for RAM reasons, so if your source is an SBR stream the ESP8266 is the wrong board regardless of the library. The same RAM constraint is why the README points at ESP8266SAM for speech synthesis rather than expecting this library to do it: formant synthesis with low memory and no network is a separate project that uses this one.

The repository also carries a tests/ directory and a tools/ directory, but the README does not document what they contain, so do not assume a documented test suite you can run against your own board.

How it compares with BackgroundAudio

The honest alternative here is not a different vendor's library. It is the author's own BackgroundAudio, which the README recommends to ESP32 and Pico users in a section near the top. The difference is the output model. ESP8266Audio is cooperative: your loop() drives decoding, and audio continuity is a property of how fast you return to it. BackgroundAudio is described as using an interrupt-based, frame-aligned output model, which decouples playback from your application's timing.

If you are on an ESP32 or a Pico and your sketch does anything else at all, that difference decides the choice. If you are on an ESP8266, BackgroundAudio is not offered as an option in the README, and the loop-driven model is what you have. That makes ESP8266Audio less a general-purpose audio stack and more the ESP8266 answer, with ESP32 and Pico support carried along.

Editorial conclusion

Adopt ESP8266Audio if you are writing Arduino sketches for an ESP8266, ESP32 or Pico and you want an MP3, MOD, MIDI or FLAC file to come out of a DAC without writing a decoder. Do not adopt it if your loop() blocks for tens of milliseconds at a time, if you need AAC on an ESP8266 with SBR, or if you are building a commercial product around AAC without checking the Helix RSPL terms and the Via Licensing requirement the README names. Before wiring anything, verify your toolchain: the README asks for Arduino ESP8266 core 2.6.3 or later, and on ESP32 it states the library needs IDF 5.x, which the stock platformio/framework-arduinoespressif32 package does not provide.

Frequently asked questions

What is ESP8266Audio used for?

It is an Arduino library that parses and decodes MOD, WAV, MP3, FLAC, MIDI, AAC, RTTTL and OGG/Opus files and plays them through an I2S DAC or a software delta-sigma DAC. The README lists uses such as MP3 alarm clocks, a word clock with an MP3 alarm, model train sound controllers and web radio.

Is there an audio library for ESP32?

Yes, this library supports the ESP32 in addition to the ESP8266 and the Raspberry Pi Pico RP2040 and Pico 2 RP2350. The README notes that ESP32 and Pico users should consider BackgroundAudio instead, since it uses an interrupt-based, frame-aligned output model with better results and performance on the Pico.

How do I install the ESP8266Audio library?

The README clones the repository into your Arduino libraries directory with git clone https://github.com/earlephilhower/ESP8266Audio. For ESP8266 it also asks you to set Tools->lwIP Variant to v1.4 Open Source or V2 Higher Bandwidth and Tools->CPU Frequency to 160MHz.

Why does ESP8266Audio fail to build on ESP32 with PlatformIO?

The README states that the platformio/framework-arduinoespressif32 package is much older than the current Espressif Arduino core and sits at IDF 4.x, while this library needs IDF 5.x and the new I2S APIs. The symptom it quotes is cannot open source file "driver/i2s_std.h", and the fix is to point platform at the pioarduino community core.

Does ESP8266Audio support AAC on the ESP8266?

The README states that AAC-SBR is supported on the ESP32 but not on the ESP8266, because the ESP8266 lacks the onboard RAM for it. It also notes that commercial use of AAC still requires the usual licensing from Via Licensing.

Official sources

  1. earlephilhower/ESP8266Audio on GitHub
  2. Issues
  3. License: GPL-3.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/earlephilhower-esp8266audio.svg)](https://hysenlabs.com/projects/earlephilhower-esp8266audio)