Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ The keypad sends notes `36..51` in the measured physical key order and uses the

At startup, the firmware scans `/samples`, loads saved assignments, and prepares eligible samples in RAM. Preloading is limited to supported files no longer than 5 seconds that fit within the configured RAM budget (1 MiB by default). Other supported files stream from SD; failed preloads fall back to streaming. Missing or unsupported files are marked unavailable when assignments are prepared.

Each trigger starts a voice. Retriggering the same sample fades out its older voices, and if all 32 slots are occupied, the oldest voice is replaced. RAM and SD playback use the same decoder and mixer, with float summation and a look-ahead peak limiter before PCM16 conversion. A single voice at `VOL=100` keeps its original digital level; overlapping voices are attenuated when their sum would exceed full scale. The audio task runs on core 1; the UI and sample loader run on core 0 and communicate with it through queues.
Each trigger starts a voice. Retriggering the same sample fades out its older voices, and if all 32 slots are occupied, the oldest voice is replaced. RAM and SD playback use the same decoder and mixer, with float summation and a look-ahead peak limiter before PCM16 conversion. Each voice has a short 35-frame (about 0.8 ms) ramp at the file boundaries. Outside these ramps, a single voice at `VOL=100` keeps its original digital level; overlapping voices are attenuated when their sum would exceed full scale. The audio task runs on core 1; the UI and sample loader run on core 0 and communicate with it through queues.

Saving stores assignments, assigned sample volumes, playback modes, and the panic note in `/sampler_config.json`, and refreshes RAM preparation. Save between performances: the save process waits for playback to finish and can stop running loops before rebuilding the sample pool.

Expand Down Expand Up @@ -86,6 +86,8 @@ Take tip (left) or ring (right), plus sleeve (headphone ground), from the headph
4. Open `LIB`, preview a sample, then hold the right button and send a MIDI note or press a keypad key to assign it.
5. Trigger the sound, set `VOL` and `SHOT/LOOP`, then select `SAVE` before powering off.

During boot, every loaded library entry (up to 32, including unassigned samples) is checked for the supported WAV format and valid file structure. The display shows progress and the rejected count. Rejected entries remain visible in `LIB` with `!` and a reason; they cannot play. Results are cached in RAM: restart after changing files on the SD card.

No configuration file is required for first boot; without one, the device starts with no assignments, one-shot playback, and the default RAM budget.

The repository includes a batch conversion command, requiring `ffmpeg` and Make:
Expand Down Expand Up @@ -133,7 +135,7 @@ Run the native tests without an ESP32 connected:
pio test -e native
```

`make test` runs the same command. The tests in `test/` cover UI navigation, sample and panic assignment, keypad mapping, saving state, and routing playback requests to RAM or SD, including fallback and loop controls. The audio mixer regression tests exercise the production Samplotron mixer with a simulated output: unity solo playback, summation, 32 full-scale voices, linked limiting, look-ahead/release, backpressure, tail draining, idle silence and fade retries. Run it alone with `pio test -e native -f test_audio_mixer`. Shared hardware stubs live in `test/support/`; audio timing, SD throughput, and physical wiring require checks on the device.
`make test` runs the same command. The tests in `test/` cover UI navigation, sample and panic assignment, keypad mapping, saving state, and routing playback requests to RAM or SD, including fallback and loop controls. The audio mixer regression tests exercise the production Samplotron mixer with a simulated output: unity solo playback, summation, 32 full-scale voices, linked limiting, look-ahead/release, backpressure, tail draining, idle silence and fade retries. Run it alone with `pio test -e native -f test_audio_mixer`. The full playback regression (`pio test -e native -f test_audio_playback`) also runs the production WAV decoder, voice engine, RAM/SD source adapters, fades and mixer; it checks simultaneous/staggered playback, 6 ms retriggers, nonzero endpoints and repeated slot reuse. The `test_i2s_transport` suite separately covers the ESP32 block writer and a continuously advancing simulated output clock. See [audio regression coverage](docs/audio-regression.md). Shared hardware stubs live in `test/support/`; audio timing, SD throughput, and physical wiring require checks on the device.

### Code structure

Expand Down
31 changes: 31 additions & 0 deletions docs/audio-regression.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Audio playback regression

Run `pio test -e native -f test_audio_playback` for playback continuity, or `pio test -e native` for the complete regression suite. No ESP32 is needed. These tests run on the development computer, not during firmware boot or performance.

The playback suite compiles the production `Audio` entry points and update loop, voice allocation/retrigger/stop code, ESP8266Audio 2.4.1 WAV decoder, RAM source, validated SD source view, stream manager, budgeted output, mixer and waveform capture. Only the SD filesystem, I2S device, clock and RTOS primitives are simulated. It therefore tests the actual decoding and scheduling interaction, not just arithmetic on an ideal array of simultaneous samples.

| Scenario | Assertion |
| --- | --- |
| Solo and 2/8 simultaneous distinct samples, RAM and SD | Output matches the independently summed source PCM within 1 PCM unit, including unequal file endings and trailing silence. |
| Mixed RAM/SD, offsets 1, 63, 96, 511, 977 and 3000 frames | Sample order, phase, duration and start offsets are preserved across decoder-budget and mixer-buffer boundaries. |
| Same sample triggered again before EOF | New occurrence plus the independently rendered 6 ms fade of the previous one matches the actual retrigger output. Both RAM and SD are covered. |
| Repeated source-slot reuse after EOF | No stale PCM, missing tail, extra I2S start/stop, or unintended voice stealing. |
| Nonzero start, EOF and retrigger | Boundary steps are bounded; an ongoing background voice continues unchanged. Sustain retains unity level. |
| Staggered 2/8/32 loud voices | Actual overload occurs; PCM stays bounded, follows the wide sum's polarity, and has continuous inferred limiter gain. |
| Output rejection and retry | A blocked sink accepts no frames and advances no simulated audio time. Once it accepts data again, the output matches the uninterrupted run. No claim about real hardware deadlines is made. |
| Very short files and positive/negative full-scale edges | 1–1501-frame fixtures stay bounded and end cleanly, identically under retries. |
| 6 ms explicit fade | Fade lasts 265 frames (rounding tolerance 1 frame) after queued PCM; every step of a 12000-level fixture is below 48 PCM units. |
| Deliberately damaged output | The same PCM comparison used by the continuity tests rejects a dropped frame, repeated frame, inserted zero and a 500-unit spike, even though those signals remain below full scale. |

Source tones have silent margins and smooth envelopes so an exact-reference failure identifies a playback defect rather than a discontinuity in the test recording. Separate constant-level fixtures deliberately have discontinuous endpoints to test boundary handling. The reference includes 64 frames of mixer look-ahead and the WAV decoder's initial pending zero. Changes to that decoder behavior must update the interface deliberately; the test should not silently realign broken output.

Before the boundary fix, the start and retrigger tests measured a 12000-unit jump in one frame; the EOF test measured 12045 including the continuing background tone. Each exceeded its 900/1000-unit fixture-specific limit. Playback now applies 35-frame (about 0.8 ms) smoothstep ramps at file boundaries. The 6 ms fade test also found a 108-unit final jump caused by discarded Q15 division remainder; the remainder is now distributed across the ramp. These are reproducible digital defects, not proof that every audible crackle has the same cause.

The existing `test_audio_mixer` suite separately checks unity gain, cancellation, linked limiting, 32 full-scale voices, aligned-sine shape against an independent unity sine, attack/release and sample acceptance. The playback suite's inferred-gain check bounds adjacent gain changes to 0.033 away from zero crossings (32-frame attack plus PCM rounding tolerance). It is a discontinuity check, not a perceptual transparency or distortion measurement.

Host tests cannot establish ESP32 CPU headroom, SD read latency, I2S underruns, transformer saturation or analog distortion. Confirm on the device with the same pair of samples played solo, simultaneously, with offsets, and retriggered, comparing RAM-loaded short samples against SD-streamed long ones. Capture the headphone-derived output if crackles remain and keep note of trigger timing and volume. Continuous loop crossfades and click-free stealing when all 32 slots are occupied remain separate work; the current suite does not certify them.


## I2S transport timing regression

`pio test -e native -f test_i2s_transport` tests the 128-frame staging buffer and the actual ESP32 branch of `StableAudioOutputI2S` against a driver fake. The ordinary playback suite still uses its immediate simulated I2S sink; it does not include this additional staging latency in its PCM timeline reference. Transport tests cover exact stereo packing and unity level, one driver call per complete DMA block, byte-granular partial writes and failed writes. A continuously advancing clock model also demonstrates underruns with a deliberately insufficient single-frame write budget; this corrects the blind spot of tests where simulated time stops whenever output is blocked. The timing assumptions are synthetic and do not benchmark the ESP32. See the [recording investigation](measurements/sample-test-2026-09-09.md).
28 changes: 22 additions & 6 deletions docs/documentation.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,9 +117,19 @@ Hardware verification (2026-09-09): `CHIPPOWER = 0xAA` introduced output noise o
- `.wav` and `.WAV` are recognized,
- file list is sorted alphabetically,
- UI sample limit: `32` (the first 32 matching files encountered are collected, then sorted).
- Assigned playback requires uncompressed PCM (`audioFormat = 1`), 16-bit, 44100 Hz, mono. Prepare every library sample in this format, including previews.
- All playback paths require uncompressed PCM (`audioFormat = 1`), 16-bit, 44100 Hz, mono, including previews and stream fallback.
- Use a FAT32 card. The conversion command in section 9 modifies files in place, including leading-silence trimming and gain adjustment.

### Boot-time WAV validation

Before playback tasks start, `SampleLibrary::loadFromSd()` validates every collected entry, including unassigned files. The boot screen displays `Checking WAV: n/N` and `Rejected`. Invalid entries stay in the library with `!` and a status (`BAD FORMAT`, `BAD WAV`, or `READ ERROR`); the playback router blocks them before enqueueing preview or MIDI triggers. Unchecked entries are also blocked. Valid entries continue to work.

`wav_validation.cpp` checks RIFF/WAVE identification, exact RIFF/file size agreement, chunk boundaries and odd-byte padding, one `fmt ` chunk before one nonempty `data` chunk, supported PCM parameters, byte rate, block alignment, and whole PCM16 frames. Extended `fmt ` chunks must have a consistent extension length. Unknown chunks are skipped, including metadata after PCM; duplicate or truncated chunks are rejected. Files must fit the signed 32-bit playback seek range. Chunk padding follows the [RIFF specification](https://learn.microsoft.com/en-us/windows/win32/xaudio2/resource-interchange-file-format--riff-).

Validation reads headers and seeks past PCM and metadata payloads; it does not scan audio content, detect clicks/clipping, or guarantee that every PCM sector is readable. Each catalog entry caches status, PCM offset, and length in RAM. Assignment classification and saving reuse this cache. RAM preload reads the cached PCM range directly. SD streaming uses `ValidatedWavSource` to supply a canonical 44-byte header from memory and expose only the cached PCM range, avoiding on-disk header parsing at each trigger or loop restart. Opening the file, checking that the cached range still fits, and reading PCM still involve SD access during playback.

Restart after changing files on the SD card. There is no runtime revalidation or hot-swap support; cache validity assumes files remain unchanged for the session. Validation covers the loaded library's existing 32-entry limit, not additional files outside that list.

## 5. `sampler_config.json` Configuration

Location and parser: `src/settings_store.cpp`.
Expand Down Expand Up @@ -163,7 +173,7 @@ Saving writes and parses the temporary JSON file before rotating the previous co

Sample preparation pipeline:

- MIDI-assigned samples are classified as `RAM` or `STREAM`,
- MIDI-assigned samples are classified from cached boot validation as `RAM` or `STREAM`,
- `RAM` is used only for WAV files that meet all conditions:
- PCM format (`audioFormat = 1`),
- `16-bit`,
Expand All @@ -178,10 +188,12 @@ Playback engine behavior:

- fixed `32`-voice playback pool (`Audio::kVoiceCount`),
- each trigger allocates a free voice slot when available,
- retriggering the same sample starts a new voice instance and requests short fade-out on already active voices in the same retrigger group,
- retriggering the same sample starts a new voice instance and requests a 6 ms fade-out on already active voices in the same retrigger group,
- both RAM and SD playback apply 35-frame (about 0.8 ms) smoothstep ramps at file start and natural EOF. These advance only on accepted samples, preserve the sustain level, and do not add SD reads. On extremely short samples the two ramps overlap and reduce the peak,
- if all voices are active, the incoming trigger steals the oldest active voice (deterministic `oldest-voice` policy),
- if incoming MIDI NOTE ON matches configured panic note, all currently active voices are quickly faded out and pending trigger backlog is cleared,
- works for both SD-streamed and RAM-backed sample playback,
- ESP32 I2S writes are staged in 128-frame blocks (512 bytes, up to 2.9 ms additional buffering); short writes retain their exact byte suffix and EOF tails complete with continuous idle silence,
- voice update loop applies bounded per-voice decode budget (`kVoiceLoopSampleBudget`) to keep scheduling predictable,
- per-voice gain follows `VOL / 100` in floating point. The 32-input `SamplerMixer` sums before limiting; it never narrows an overloaded sum to PCM16 first. A stereo-linked peak limiter uses 64 frames of look-ahead (1.45 ms at 44.1 kHz), linear predictive gain bounds reached within 32 frames and held for the remainder of the look-ahead, and a 50 ms peak-envelope decay constant. It reduces gain only around overloads, including the look-ahead and release intervals.
- trigger events are sent through a queue from UI/MIDI domain to dedicated audio task (no direct playback calls from UI code path).
Expand Down Expand Up @@ -272,10 +284,14 @@ Assignment rules:
### `include/audio_internal.h`

- Audio mixer buffer size: `kMixerBufferSamples = 512`
- Retrigger fade-in (new voice): `kRetriggerFadeInUs = 800`
- Retrigger fade-out (older voices in same group): `kRetriggerFadeOutUs = 6000`
- Default control stop fade-out: `kDefaultStopFadeOutUs = 9000`
- Decode budget per voice update: `kVoiceLoopSampleBudget = 96`
### `include/budgeted_audio_output.h`

- File boundary ramps: `kEdgeFrames = 35` (about 0.8 ms at 44.1 kHz), including the decoder’s initial pending zero in playback position accounting.
- Explicit stop fades distribute the Q15 division remainder over their full duration to avoid an extra final step. Queued audio is preserved, so the 6 ms retrigger fade starts at the voice’s next unqueued frame.

### `include/sampler_mixer.h`

- Mixer inputs: `kMaxInputs = 32`
Expand Down Expand Up @@ -363,9 +379,9 @@ The repository workflow runs native tests and builds the main firmware. Pushes t

### Test scope and diagnostics

`pio test -e native` covers UI navigation, sample/panic learning, keypad mapping, saving state, and playback routing, including RAM-to-stream fallback and loop control. These tests use stubs and do not exercise the real audio engine, SD hardware, or I2C wiring.
`pio test -e native` covers UI navigation, sample/panic learning, keypad mapping, saving state, and playback routing, including RAM-to-stream fallback and loop control. The `test_audio_playback` suite additionally runs the real WAV decoder, voice engine, source adapters, budgeted fades and mixer against simulated SD/I2S hardware. It compares sample timelines, tests 2/8/32 overlapping voices at high levels, and includes negative controls for missing, repeated, zeroed and spiked PCM. See [coverage and limits](audio-regression.md). Host tests do not measure ESP32 deadlines, actual SD throughput or analog output.

Main firmware serial output is limited to keypad initialization and key presses from `src/input.cpp`. For encoder or MIDI diagnostics, upload the corresponding debug environment and open the monitor at 115200 baud. These are separate applications; upload the main environment again to resume sampling.
Main firmware serial output includes keypad diagnostics, WAV rejection reasons and codec volume-register verification at boot. For encoder or MIDI diagnostics, upload the corresponding debug environment and open the monitor at 115200 baud. These are separate applications; upload the main environment again to resume sampling.

## 10. Module Map (Code Orientation)

Expand Down
5 changes: 4 additions & 1 deletion docs/manual.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Connect the isolated mono output jack to your mixer and start at a low monitorin

## 1. Power On and Wait for Ready

When you boot the device, you will first see loading status.
When you boot the device, `Checking WAV: n/N` shows validation progress for the loaded library (up to 32 samples). `Rejected` counts files that cannot play. Validation runs once at startup, before playback starts; valid samples remain available even if other files are rejected. The screenshots below show an earlier screen layout.

![Boot loading screen](screenshots/boot.jpg)

Expand Down Expand Up @@ -36,6 +36,9 @@ In `LIB`:

- Right rotate: browse samples.
- Right click: play preview.
- Entries marked `!` cannot play: `BAD FORMAT` means unsupported audio settings, `BAD WAV` means invalid file structure, and `READ ERROR` means the file could not be read. Convert or replace the file using PCM16, 44.1 kHz, mono, then restart the device.

Results remain in memory until power-off. Restart after changing SD files; saving assignments does not refresh validation.
- Left click: go back.

## 4. Assign a Sample to a MIDI Note
Expand Down
Loading
Loading