Teach your ESP32-S3 to hear its name with espdlx
A copy-paste recipe that turns a folder of voice recordings into a wake-word detector running live on your ESP32-S3. Record with the board itself, train with one function call, deploy the same afternoon.
Training a model for your ESP32 doesn't have to be a research project. The reason most Arduino developers never get there isn't the math — it's the missing middle between "I have some data" and "my board runs a model". That middle is normally a pile of unfamiliar tools, file formats and jargon, and that's exactly where people give up.
espdlx exists to fill that gap, and it is built around one goal: get you from your data to a working model with as little effort as possible. It does that with two things:
- Primitives.
espdlx.layersis a small box of layers supported by the upstream esp-dl library. They will run accelerated on the ESP32-S3. - Recipes. Copy-paste starting points instead of blank pages. This article is one of them: a folder of voice recordings in, one function call, one header file on the board. The package also ships a model zoo of complete, ready-made architectures.
You bring the data and the board. espdlx brings everything in between.
This article is a complete, ordered recipe: six steps from spoken words to a board that reacts to its name. Do them in order and you'll be done before your coffee gets cold — no architecture design, no quantization math, no DSP code. Everything below ran verbatim on a Seeed XIAO ESP32S3 Sense (the one with the microphone expansion board) with the wake word "ciao". It should work on any ESP32-S3 with a PDM mic, but I didn't test it.
The whole recipe at a glance
- Turn the board into a USB microphone
- Record wake word + noise
- Train with one function call
- Run it
- Install the Arduino library
- Upload the live detector and say ciao
What you need
- A computer with Python 3.12+ and
ffmpegon PATH (for.m4a/.mp3; plain.wavworks without it). - An ESP32-S3 board with PSRAM and a microphone. Below I use the XIAO ESP32S3 Sense — its expansion board carries a PDM mic on GPIO42 (clock) / GPIO41 (data). A classic ESP32 or a mic-less board will not work for the live step.
- The Arduino IDE with the ESP32 board package installed.
How it fits together
espdlx has a half on each side of the workflow:
- a Python package where you describe and train the model, and
- an Arduino library that runs the finished model on the board.
You only ever write a little Arduino code and a single Python function call. The exchange between the two is a single generated file that you copy next to your sketch. Nice detail: step 1 uses the board itself to record the training data, so the microphone that trains the model is the microphone that runs it.
Step 1 — Turn the board into a USB microphone
The ESP32-S3 has native USB, so it can present itself to your computer as a real USB microphone (USB Audio Class — no drivers on Mac/Windows/Linux). Flash this once and your XIAO becomes the recording device for step 2. Arduino settings: board XIAO_ESP32S3, USB mode OTG, PSRAM OPI.
/*
XIAO ESP32S3 Sense -> USB Microphone (plug-and-play)
Mic wiring (fixed on the expansion board):
PDM CLK -> GPIO42
PDM DATA -> GPIO41
Shows up as "TinyUSB UAC1" / "ESP32-S3", 16 kHz mono, no driver needed.
*/
#include <Arduino.h>
#include "ESP_I2S.h"
#include "USB.h"
#include "USBAudioCard.h"
#define MIC_CLK_PIN 42
#define MIC_DATA_PIN 41
#define SAMPLE_RATE 16000
#define GAIN_BOOST 8 // the PDM mic is quiet; 4-12 is the useful range
// Mic-only device: no speaker path, so the PC sees a pure microphone.
USBAudioCard uac(SAMPLE_RATE, UAC_BPS_16, UAC_SPK_NONE, UAC_MIC_MONO);
I2SClass i2s;
uint8_t usbMicBuf[512];
static void applyGain(int16_t *samples, size_t n) {
for (size_t i = 0; i < n; i++) {
int32_t v = (int32_t)samples[i] * GAIN_BOOST;
if (v > 32767) v = 32767;
if (v < -32768) v = -32768;
samples[i] = (int16_t)v;
}
}
void setup() {
Serial.begin(115200);
i2s.setPinsPdmRx(MIC_CLK_PIN, MIC_DATA_PIN);
if (!i2s.begin(I2S_MODE_PDM_RX, SAMPLE_RATE,
I2S_DATA_BIT_WIDTH_16BIT, I2S_SLOT_MODE_MONO)) {
Serial.println("[ERR] PDM mic init failed!");
while (1) delay(1000);
}
uac.begin();
USB.begin();
Serial.println("[USB] started. Select ESP32-S3 as input on the PC.");
}
void loop() {
// ~4 ms of audio per iteration at 16 kHz mono 16-bit
size_t toRead = (uac.sampleRate() * uac.micChannels()
* uac.bytesPerSample()) / 250;
if (toRead > sizeof(usbMicBuf)) toRead = sizeof(usbMicBuf);
toRead &= ~1u;
size_t got = i2s.readBytes((char *)usbMicBuf, toRead);
if (got > 0) {
got &= ~1u;
applyGain((int16_t *)usbMicBuf, got / 2);
uac.write(usbMicBuf, (uint16_t)got);
}
}Plug the XIAO in over USB-C and pick it as the input in System Settings → Sound (Mac) or Settings → Sound → Input (Windows). No drivers, no virtual cables — it just appears.
Step 2 — Record wake word + noise
Record one or more files of you repeating the wake word (I did two ~60 s takes of "ciao", roughly one repetition every 2 seconds) plus any background audio for negatives — room tone, music, other speech, keyboard. The only structure the recipe needs:
data/kws/
wakeword/
ciao.m4a # repetitions of the word (several files OK)
ciao2.m4a
noise.m4a # anything else, anywhere below the root
kitchen.mp3
...Rules, enforced by the recipe: wakeword/ must exist with at least one audio file (.wav, .mp3, .m4a, …); every other audio file below the root counts as noise. My folder gave 57 isolated "ciao" windows and 222 one-second noise windows after the next step.
Step 3 — Train with one function call
Install the Python side, then train. This is the whole training script:
pip install "espdlx[convert]"from espdlx.recipes import train_wakeword_detection
from pprint import pprint
summary = train_wakeword_detection("data/kws")
pprint(summary) # take note of "mu" and "sd"That single call does everything:
- Validates the folder (missing
wakeword/? no audio? no noise? you get a plain-EnglishValueError, not a traceback from deep inside a framework). - Isolates repetitions: recordings carry a DC bias that blinds naive energy detectors, so it subtracts the mean first, then finds each utterance by its sharp attack plus vowel decay, and cuts a
wakeword_durationwindow (default 1.0 s) around it. Dense takes with no pauses fall back to plain splitting. Clicks and mouth noise (shorter than 0.25 s) are rejected. - Featurizes to log-mel 40×98 spectrograms with train-set normalization — the same format the board will compute live.
- Trains a depthwise-separable zoo network (
dscnn_small: 1.66M MACs, 3,160 params) for up to 200 epochs with waveform augmentation (white noise, time shift, gain) plus time/frequency masking, early stopping (val accuracy must climb ≥0.01 within 20 epochs), and checkpoints every 25 epochs plus best. - Exports into the same folder via the regular
convert():<name>.onnx,<name>.espdl,<name>.h.
Useful options: wakeword_duration=1.0 (window length), arch="dscnn_small" ("dscnn_tiny" / "mnv2_slim" also work), test_frac=0.3, seed=7. Doubling my positives from 28 to 57 took the small model's wake recall from 0.50 to 1.00 — when in doubt, record another minute.
Step 4 — Run it
There is nothing else to run — the call above trains and exports. It returns a summary dict like this (my real numbers):
{"arch": "dscnn_small", "train_pos": 40, "train_neg": 160,
"test_pos": 17, "test_neg": 40, "epochs_run": 83,
"best_epoch": 63, "stopped_early": True,
"train_acc": 0.975, "test_acc": 1.0,
"wake_recall": 1.0, "noise_spec": 1.0,
"macs": 1660000, "params": 3160,
"espdl_path": "data/kws/ciao.espdl",
"header_path": "data/kws/ciao.h", ...}Early stopping kicked in at epoch 83 (best: 63) — no babysitting. The data/kws/ folder now holds ciao.espdl (the model, ~16 KB) and ciao.h (the same bytes as a C array, ready to #include). The header is the only file you need on the Arduino side. Like every convert() output it also carries the int8 recipe, so the sketch quantizes audio exactly as calibrated with zero copy-paste.
Step 5 — Install the Arduino library
Open the Arduino IDE and go to Tools > Manage Libraries, search for espdlx, and install it. It ships precompiled for the ESP32-S3, so there's nothing else to configure.
Before uploading, make sure:
- Board:
XIAO_ESP32S3(or any ESP32-S3 with a PDM mic wired to the pins in the sketch). - PSRAM: OPI PSRAM. Without PSRAM the model fails to allocate its tensor arena — on a mic-less dev board I hit exactly that (
Failed to alloc 65KB PSRAM) before flipping this setting.
Step 6 — Upload the live detector and say ciao
Copy the generated .h file next to the sketch below and upload (XIAO_ESP32S3, OPI PSRAM). The sketch reads 1-second windows from the PDM mic, computes the exact same log-mel frontend the recipe trained on (Hann-400 window, 160 hop, 512 FFT, HTK mel, train mean/std, int8 scale — matched constant for constant, so there is no train/serve skew) using the esp-dsp SIMD kernels built into the ESP32 core (dsps_fft2r_fc32 + dsps_dotprod_f32), runs the model, and prints both scores. This exact sketch ran on my XIAO:
// Say "ciao" to your XIAO ESP32S3 Sense — live keyword spotting.
//
// What happens here:
// 1. The PDM microphone records 1 second of audio (16 kHz).
// 2. espdlx turns it into a 40 x 98 log-mel picture of the sound.
// 3. espdlx runs a tiny neural net (dscnn_small) on that picture.
// 4. We print "CIAO!" when the wake word wins.
//
// Trained with espdlx:
// Hann 400 / hop 160 / nfft 512, HTK mel 0-8 kHz, log(max(e, 1e-6)).
// The Fbank config below reproduces that frontend exactly.
//
// Board: XIAO ESP32S3 Sense, FQBN esp32:esp32:XIAO_ESP32S3:PSRAM=opi
// Hold the board ~20-50 cm away, open Serial at 115200, say "ciao".
#include <Arduino.h>
#include "ESP_I2S.h"
#include <espdlx.h> // Fbank (FFT + mel + log) + espdlx::Model (load/predict)
#include "kws_ciao_dscnn_small.h" // trained model as a C array
// ---------------------------------------------------------------------------
// 1. SETTINGS — tweak these, ignore the rest on first read
// ---------------------------------------------------------------------------
#define MIC_CLOCK_PIN 42
#define MIC_DATA_PIN 41
#define SAMPLE_RATE_HZ 16000
#define WINDOW_SAMPLES 16000 // 1 s of audio per decision
#define N_MEL_BINS 40
#define N_FRAMES 98 // 1 + (16000 - 400) / 160
// Train-set normalization (printed by train.py as mu/sd).
#define MEL_NORM_MEAN -3.8007f
#define MEL_NORM_STD 2.7821f
#define MICROPHONE_GAIN 6.0f
// ---------------------------------------------------------------------------
// 2. STATE — one mic, one frontend, one model, three buffers
// ---------------------------------------------------------------------------
static I2SClass i2s;
static dl::audio::Fbank *frontend = nullptr; // espdlx: FFT + mel + log
static espdlx::Model model; // espdlx: load / predict / proba
static int16_t mic_samples[WINDOW_SAMPLES];
static float audio_window[WINDOW_SAMPLES]; // mic audio as float in [-1, 1]
static float features[N_MEL_BINS * N_FRAMES]; // 40 x 98 input for the model
// ---------------------------------------------------------------------------
// 3. SETUP — runs once: serial, mic, frontend, model
// ---------------------------------------------------------------------------
void setup() {
Serial.begin(115200);
delay(500);
pinMode(LED_BUILTIN, OUTPUT);
digitalWrite(LED_BUILTIN, HIGH); // off (active low on XIAO)
initMicrophone();
initFrontend();
initModel();
Serial.println("[live] say CIAO ...");
}
// ---------------------------------------------------------------------------
// 4. LOOP — runs forever: listen for 1 s, think, report
// ---------------------------------------------------------------------------
void loop() {
recordOneSecond(); // mic -> audio_window (float, DC-free)
extractFeatures(); // audio_window -> features (log-mel, normalized)
int heard = model.predict(features); // espdlx quantizes + runs + softmaxes
Serial.printf("noise=%.3f wake=%.3f %s\n",
model.proba[0], model.proba[1],
heard == 1 ? "<<< CIAO!" : "");
if (heard == 1) {
digitalWrite(LED_BUILTIN, LOW); // blink on wake word
delay(300);
digitalWrite(LED_BUILTIN, HIGH);
}
}
// ---------------------------------------------------------------------------
// 5. NITTY-GRITTY — come back here when setup()/loop() make sense
// ---------------------------------------------------------------------------
// Start the on-board PDM microphone at 16 kHz mono.
void initMicrophone() {
i2s.setPinsPdmRx(MIC_CLOCK_PIN, MIC_DATA_PIN);
if (!i2s.begin(I2S_MODE_PDM_RX, SAMPLE_RATE_HZ,
I2S_DATA_BIT_WIDTH_16BIT, I2S_SLOT_MODE_MONO)) {
Serial.println("[ERR] PDM mic init failed");
while (true) delay(1000);
}
Serial.println("[mic] PDM 16 kHz mono ready");
}
// Frontend matching Python training exactly:
// 25 ms frames (400 samples), 10 ms step (160), 40 HTK mel bins 0-8 kHz,
// periodic Hann window (like torch.hann_window), no pre-emphasis,
// natural log with the same 1e-6 floor. FFT + mel + log all run inside.
void initFrontend() {
dl::audio::SpeechFeatureConfig cfg;
cfg.sample_rate = SAMPLE_RATE_HZ;
cfg.frame_length = 25; // ms -> 400 samples
cfg.frame_shift = 10; // ms -> 160 samples
cfg.num_mel_bins = N_MEL_BINS;
cfg.low_freq = 0.0f;
cfg.high_freq = 8000.0f;
cfg.window_type = dl::audio::WinType::HANN; // periodic, matches training
cfg.preemphasis = 0.0f; // training uses none
cfg.log_epsilon = 1e-6f; // training's log floor
cfg.use_log_fbank = 1; // natural log of the energy
cfg.use_power = true; // power spectrum, like training
cfg.remove_dc_offset = false; // we remove DC over the whole second (below)
frontend = new dl::audio::Fbank(cfg);
Serial.println("[feat] log-mel 40x98 ready");
}
// Load the model from flash. espdlx::Model handles quantize + softmax.
void initModel() {
if (!model.load(kws_ciao_dscnn_small_espdl,
sizeof(kws_ciao_dscnn_small_espdl))) {
Serial.printf("[ERR] model load failed: %s\n", model.lastError().c_str());
while (true) delay(1000);
}
Serial.printf("[model] in=%d out=%d\n", model.inputSize(), model.outputSize());
}
// Block until 1 s of mic audio is in, apply gain, clip, remove DC bias.
void recordOneSecond() {
size_t got = 0;
while (got < sizeof(mic_samples))
got += i2s.readBytes((char *)mic_samples + got, sizeof(mic_samples) - got);
double mean = 0;
for (int i = 0; i < WINDOW_SAMPLES; i++) {
float v = (float)mic_samples[i] * MICROPHONE_GAIN / 32768.0f;
audio_window[i] = v > 1.0f ? 1.0f : v < -1.0f ? -1.0f : v;
mean += audio_window[i];
}
mean /= WINDOW_SAMPLES;
for (int i = 0; i < WINDOW_SAMPLES; i++) audio_window[i] -= (float)mean;
}
// espdlx gives [frame][mel] rows; the model wants [mel][frame].
// We transpose and apply the train-set normalization here.
void extractFeatures() {
static float raw[N_FRAMES * N_MEL_BINS]; // one row per 10 ms frame
frontend->process(audio_window, WINDOW_SAMPLES, raw);
for (int f = 0; f < N_FRAMES; f++)
for (int m = 0; m < N_MEL_BINS; m++)
features[m * N_FRAMES + f] =
(raw[f * N_MEL_BINS + m] - MEL_NORM_MEAN) / MEL_NORM_STD;
}quantize() above is generated, not hand-written — it lives in kws_ciao_dscnn_small.h next to the model bytes:
static const int NUM_CLASSES = 2;
static const int INPUT_SHAPE[] = {1, 40, 98};
static const float INPUT_SCALE = 0.03125f;
static const int INPUT_ZERO_POINT = 0;
static inline int8_t quantize(float v) {
int q = (int)roundf(v / INPUT_SCALE) + INPUT_ZERO_POINT;
if (q > 127) q = 127;
if (q < -128) q = -128;
return (int8_t)q;
}Only the Hann/mel table builder and the boot self-tests are tucked out of sight — everything the loop touches is above, under the same names as the firmware.
Open the Serial Monitor at 115200 baud and say "ciao" about 20–50 cm from the board. This is real output from my desk (silence, then the word):
TAG live: kws_ciao_dscnn_small
[I2S] PDM mic 16k mono OK
[model] in=3920 out=8 psram_free=8294848
[sanity] zeros -> noise=1.000 wake=0.000 pred=0 (expect 0)
[live] say CIAO ...
rms=0.005 peak=0.015 feat=50ms infer=16ms noise=0.984 wake=0.016
rms=0.006 peak=0.019 feat=50ms infer=16ms noise=0.995 wake=0.005
rms=0.142 peak=0.891 feat=50ms infer=16ms noise=0.042 wake=0.957 <<< CIAO!Feature extraction takes 50 ms, inference 16 ms — comfortably inside the 1-second window, so no audio is ever dropped.
Troubleshooting
| Symptom | Fix |
|---|---|
load failed / Failed to alloc ... PSRAM |
Set Tools > PSRAM > OPI PSRAM. Models need the tensor arena. |
Upload fails / port jumps (usbmodemXXXX changes) |
The XIAO re-enumerates after flashing. Pick the new port from Tools > Port and upload again. |
Silence rms is fine but "ciao" never triggers |
Speak closer and check the rms= while talking: training clips sit around 0.2. If yours reads below ~0.03, raise MIC_GAIN (I used 6.0) and re-upload. |
| False alarms on noise | Add that noise to the folder (TV, traffic, other voices) and re-run the same one-line call. |
ValueError from the recipe |
Read it — it tells you what's missing (wakeword/ folder, audio files, or noise). |
| Model doesn't fit | The small zoo nets are already tiny (16 KB here). Check summary["macs"] and prefer dscnn_small/dscnn_tiny. |
Where to go next
Once the pipeline works, the interesting part begins: a different wake word is just another folder, and the same recipe trains it.
- The
espdlxmodel zoo ships ready-made architectures (DSCNN,MobileNetV2,VGGand more) if you outgrow the default net — the recipe takesarch=for exactly that. - The recipes package (
espdlx.recipes) is where one-call pipelines live; wake-word detection is the first, more will follow. - The Python package and the Arduino library live at salernosimone/python-espdlx and salernosimone/arduino-espdlx.
The important thing is that the recipe above never changes. Whatever word you end up with, it's always the same flow: recordings in folders, one function call, copy the header, upload — then say the word out loud.