Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ For a walkthrough with screen photos, see the [musician's manual](docs/manual.md

## Technical Details

Samplotron uses an ESP32 with PSRAM, built with the `esp-wrover-kit` PlatformIO board configuration, and an ES8388 audio codec. It plays mono PCM16 WAV files at 44.1 kHz through a shared 32-voice engine. Short assigned samples can be preloaded into RAM; longer samples stream from SD.
Samplotron uses an ESP32 with PSRAM, built with the `esp-wrover-kit` PlatformIO board configuration, and an ES8388 audio codec. It plays mono PCM16 WAV files at 44.1 kHz through a shared 32-voice engine. Assigned samples are preloaded into PSRAM when they fit; the rest stream from SD.

### Controls

Expand Down Expand Up @@ -51,7 +51,7 @@ The keypad sends notes `36..51` in the measured physical key order and uses the

### How playback works

At startup, the firmware scans `/samples`, loads saved assignments, and prepares eligible samples in RAM. Preloading is limited to supported files no longer than 5 seconds that fit within the configured RAM budget (1 MiB by default). Other supported files stream from SD; failed preloads fall back to streaming. Missing or unsupported files are marked unavailable when assignments are prepared.
At startup, the firmware scans `/samples`, loads saved assignments, and prepares eligible samples in RAM. The RAM pool is sized from free PSRAM at startup (most of a 4 MB module, several tens of seconds of mono audio), and there is no per-sample length limit. If the assigned samples do not all fit, the shortest are preloaded first. Other supported files stream from SD through a separate reader task, with their first 16 KiB kept in RAM so they start without delay; failed preloads also fall back to streaming. Serial output reports how many samples were loaded and how long it took. Missing or unsupported files are marked unavailable when assignments are prepared.

Each trigger starts a voice. Retriggering the same sample fades out its older voices, and if all 32 slots are occupied, the oldest voice is replaced. RAM and SD playback use the same decoder and mixer, with float summation and a look-ahead peak limiter before PCM16 conversion. Each voice has a short 35-frame (about 0.8 ms) ramp at the file boundaries. Outside these ramps, a single voice at `VOL=100` keeps its original digital level; overlapping voices are attenuated when their sum would exceed full scale. The audio task runs on core 1; the UI and sample loader run on core 0 and communicate with it through queues.

Expand Down Expand Up @@ -90,7 +90,7 @@ The separate L/R speaker terminals carry a switching, speaker-level signal from

During boot, every loaded library entry (up to 32, including unassigned samples) is checked for the supported WAV format and valid file structure. The display shows progress and the rejected count. Rejected entries remain visible in `LIB` with `!` and a reason; they cannot play. Results are cached in RAM: restart after changing files on the SD card.

No configuration file is required for first boot; without one, the device starts with no assignments, one-shot playback, and the default RAM budget.
No configuration file is required for first boot; without one, the device starts with no assignments and one-shot playback.

On Linux, install FFmpeg using your distribution's package manager (for example, `sudo apt install ffmpeg` on Debian/Ubuntu). Run this in Bash, setting `samples_dir` to the directory containing your WAV files:

Expand Down Expand Up @@ -121,7 +121,7 @@ Run it on copies of your source recordings: it replaces WAV files in place, conv

Display orientation is configured by `DisplayConfig::ROTATE_180` in [include/display_config.h](include/display_config.h). It defaults to `true` (180° rotation); set it to `false` for the original orientation. Rebuild and upload the firmware after changing it. This applies to all screens in the main firmware and the input diagnostic firmware.

The on-device `SAVE` action writes `/sampler_config.json`. For manual configuration, including the RAM budget and panic note, see the [configuration format](docs/documentation.md#5-sampler_configjson-configuration). Changing the RAM budget requires a reboot.
The on-device `SAVE` action writes `/sampler_config.json`. For manual configuration, including the panic note, see the [configuration format](docs/documentation.md#5-sampler_configjson-configuration).

The main firmware prints keypad initialization and key-press diagnostics at 115200 baud. Dedicated firmware environments are available for testing encoders and MIDI input:

Expand Down
25 changes: 11 additions & 14 deletions docs/documentation.md
Original file line number Diff line number Diff line change
Expand Up @@ -141,7 +141,7 @@ Before playback tasks start, `SampleLibrary::loadFromSd()` validates every colle

`wav_validation.cpp` checks RIFF/WAVE identification, exact RIFF/file size agreement, chunk boundaries and odd-byte padding, one `fmt ` chunk before one nonempty `data` chunk, supported PCM parameters, byte rate, block alignment, and whole PCM16 frames. Extended `fmt ` chunks must have a consistent extension length. Unknown chunks are skipped, including metadata after PCM; duplicate or truncated chunks are rejected. Files must fit the signed 32-bit playback seek range. Chunk padding follows the [RIFF specification](https://learn.microsoft.com/en-us/windows/win32/xaudio2/resource-interchange-file-format--riff-).

Validation reads headers and seeks past PCM and metadata payloads; it does not scan audio content, detect clicks/clipping, or guarantee that every PCM sector is readable. Each catalog entry caches status, PCM offset, and length in RAM. Assignment classification and saving reuse this cache. RAM preload reads the cached PCM range directly. SD streaming uses `ValidatedWavSource` to supply a canonical 44-byte header from memory and expose only the cached PCM range, avoiding on-disk header parsing at each trigger or loop restart. Opening the file, checking that the cached range still fits, and reading PCM still involve SD access during playback.
Validation reads headers and seeks past PCM and metadata payloads; it does not scan audio content, detect clicks/clipping, or guarantee that every PCM sector is readable. Each catalog entry caches status, PCM offset, and length in RAM. Assignment classification and saving reuse this cache. RAM preload reads the cached PCM range directly. SD streaming (`StreamManager`) supplies a canonical 44-byte header from memory and exposes only the cached PCM range, avoiding on-disk header parsing at each trigger or loop restart. An `sd_reader` task on core 0 opens files, checks that the cached range still fits, and fills a 16 KiB ring buffer (about 185 ms) per stream in PSRAM, serving the emptiest stream first; the audio task on core 1 only copies buffered PCM. Up to 16 SD streams can play at once; a further SD trigger is dropped. The first 16 KiB (about 185 ms) of every assigned sample that streams is preloaded into the RAM pool, so it starts immediately while the reader opens the file and reads on from there; loops replay that start from RAM. Samples without a preloaded start (library previews, unsaved assignments) begin with silence until their first 2 KiB arrives (roughly 10–30 ms at 4 MHz). If a stream runs dry, only that voice receives silence for the affected update; other voices keep playing. Looped SD samples keep their file open and the reader continues into the next iteration, so restarts need no SD access.

Restart after changing files on the SD card. There is no runtime revalidation or hot-swap support; cache validity assumes files remain unchanged for the session. Validation covers the loaded library's existing 32-entry limit, not additional files outside that list.

Expand All @@ -155,7 +155,6 @@ Minimal format:
{
"version": "1.0",
"global_settings": {
"sample_ram_budget_bytes": 1048576,
"panic_note": 24
},
"midi_assignments": [
Expand All @@ -178,7 +177,8 @@ Notes:
- `volume = 100`: unity per-voice gain; a single sample keeps its original level (no automatic normalization of quiet WAV files)
- `sample_path`: full SD path, for example `/samples/snare.wav`
- maximum assignments in the settings structure: `128`; the UI catalog holds at most `32` samples and assigns each sample to one note
- without a readable configuration, loading begins from defaults: no assignments, no panic note, a 1 MiB RAM budget, and one-shot playback
- without a readable configuration, loading begins from defaults: no assignments, no panic note, and one-shot playback
- `sample_ram_budget_bytes`, written by older firmware, is ignored; the RAM pool is sized automatically (section 6)

The writer also saves `sample_playback_modes`, an array of `sample_path` / `playback_mode` objects. The UI save flow includes library samples set to `loop`, including those without note assignments; omitted unassigned samples use the default `shot` mode. Assigned sample volumes are saved in `midi_assignments`; unassigned preview volumes are not persisted.

Expand All @@ -194,8 +194,11 @@ Sample preparation pipeline:
- `16-bit`,
- `44100 Hz`,
- `mono`,
- duration `<= 5.0 s`,
- fit into the RAM budget,
- fit into the RAM budget; there is no per-sample duration limit,
- the budget is the largest free PSRAM block at the first preparation minus a 512 KiB reserve (`SampleRamManager::budgetBytes()`); without PSRAM it falls back to 1 MiB,
- the first 16 KiB of every assigned sample is reserved before packing (`SampleClassifier::kStreamHeadBytes`); samples that end up streaming keep that start in RAM so they begin without waiting for SD. If even these starts do not fit, none are reserved,
- when the assigned samples do not all fit, they are packed shortest first, so the fewest samples stream from SD,
- preload reads 32 KiB chunks; Serial reports the loaded count, size and time. At the 4 MHz SD fallback, a full pool takes several seconds to load at boot and on every `SAVE`,
- if preload fails, the entry falls back to `STREAM`.
- if assigned sample format is unsupported/missing, playback for that note is blocked (`UNAVAILABLE`) instead of trying to decode anyway.

Expand All @@ -215,9 +218,7 @@ Playback engine behavior:

Important behavior:

- RAM pool budget is "locked" after the first `prepare()` (`sample_ram_manager.cpp`),
- changing `sample_ram_budget_bytes` in the same runtime session is recorded in the preparation result as `fixedBudgetMismatch`,
- a real budget change requires a device reboot.
- the RAM pool is allocated once at the first `prepare()` and never resized (`sample_ram_manager.cpp`); later saves repack samples into the same pool.

### Mixing level policy

Expand Down Expand Up @@ -281,13 +282,9 @@ Assignment rules:
- Encoder detent: `4` ticks
- Long press (right encoder): `700 ms`

### `include/sample_classifier.h`

- RAM preload threshold: `kFixedPreloadThresholdSeconds = 5.0f`

### `include/settings_store.h`

- Default RAM budget: `kDefaultSampleRamBudgetBytes = 1 MB`
- Fallback RAM budget without PSRAM: `kDefaultSampleRamBudgetBytes = 1 MB`

### `include/ui.h`

Expand Down Expand Up @@ -396,7 +393,7 @@ The repository workflow runs native tests and builds the main firmware. Pushes t

`pio test -e native` covers UI navigation, sample/panic learning, keypad mapping, saving state, and playback routing, including RAM-to-stream fallback and loop control. The `test_audio_playback` suite additionally runs the real WAV decoder, voice engine, source adapters, budgeted fades and mixer against simulated SD/I2S hardware. It compares sample timelines, tests 2/8/32 overlapping voices at high levels, and includes negative controls for missing, repeated, zeroed and spiked PCM. See [coverage and limits](audio-regression.md). Host tests do not measure ESP32 deadlines, actual SD throughput or analog output.

Main firmware serial output includes keypad diagnostics, WAV rejection reasons and codec volume-register verification at boot. For encoder or MIDI diagnostics, upload the corresponding debug environment and open the monitor at 115200 baud. These are separate applications; upload the main environment again to resume sampling.
Main firmware serial output includes keypad diagnostics, WAV rejection reasons and codec volume-register verification at boot. It also reports the SD SPI clock the card mounted at (20 MHz, falling back to 10 or 4 MHz). During playback, a line starting with `Audio:` appears in any second with new I2S underruns or SD reads slower than 3 ms. Underruns make the DMA replay stale blocks, heard as stutter and stretched-sounding playback; a quiet log means playback kept up. For encoder or MIDI diagnostics, upload the corresponding debug environment and open the monitor at 115200 baud. These are separate applications; upload the main environment again to resume sampling.

## 10. Module Map (Code Orientation)

Expand Down
24 changes: 23 additions & 1 deletion include/audio.h
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,21 @@ class Audio {
uint32_t voiceStealCount = 0;
};

// Cumulative counters for diagnosing playback that cannot keep up. Safe to
// poll from another task; values may be a few updates stale.
struct StreamingDiagnostics {
uint32_t i2sUnderrunCount = 0;
// Voice updates that got silence because the SD reader fell behind.
uint32_t starvedUpdateCount = 0;
uint32_t sdReadCount = 0;
uint32_t sdBytesRead = 0;
uint32_t sdMaxReadUs = 0;
uint32_t sdMaxReadBytes = 0;
uint32_t sdOpenFailureCount = 0;
// SD triggers dropped because every stream buffer was in use.
uint32_t sdNoFreeStreamCount = 0;
};

struct WaveformSnapshot {
int8_t points[kWaveformPointCount] = {0};
uint16_t validPoints = 0;
Expand All @@ -28,11 +43,17 @@ class Audio {

void setSampleCatalog(const SampleLibrary::Catalog *catalog) { catalog_ = catalog; }
bool begin();
// Moves SD streaming reads to their own task. Until then, update() reads.
bool startStreamReader(uint8_t priority, int core);
void update();
// head/headBytes: optional preloaded start of the file's PCM, played from
// RAM while the SD reader catches up.
void playSamplePath(const String &samplePath,
uint8_t volume = 100,
int16_t retriggerGroupId = -1,
bool loopEnabled = false);
bool loopEnabled = false,
const uint8_t *head = nullptr,
uint32_t headBytes = 0);
void stopAllVoices();
void fadeOutAllVoices(uint32_t fadeOutUs);
void stopLoopingVoicesForGroup(int16_t retriggerGroupId);
Expand All @@ -47,6 +68,7 @@ class Audio {
bool loopEnabled = false);
RuntimeStats runtimeStats() const;
uint32_t voiceStealCount() const;
StreamingDiagnostics streamingDiagnostics() const;
bool waveformSnapshot(WaveformSnapshot &snapshot) const;

private:
Expand Down
8 changes: 8 additions & 0 deletions include/audio_internal.h
Original file line number Diff line number Diff line change
Expand Up @@ -79,19 +79,24 @@ class StableAudioOutputI2S : public AudioOutputI2S {
public:
StableAudioOutputI2S(int port, int outputMode, int dmaCount, int useApll);

bool begin() override;
bool SetRate(int hz) override;

bool ConsumeSample(int16_t sample[2]) override;

uint32_t rateSetCalls() const;
uint32_t skippedRateSetCalls() const;
uint32_t appliedRateSetCalls() const;
// DMA blocks replayed because no new PCM arrived in time (ESP32 only).
uint32_t underrunCount() const;

private:
#ifdef ESP32
static size_t writeBlock(void *context, const uint8_t *data, size_t bytes);
static bool onSendQueueOverflow(i2s_chan_handle_t handle, i2s_event_data_t *event, void *context);
PcmBlockBuffer block_;
#endif
volatile uint32_t underrunCount_ = 0;
int lastRateHz_ = -1;
uint32_t rateSetCalls_ = 0;
uint32_t skippedRateSetCalls_ = 0;
Expand Down Expand Up @@ -125,6 +130,7 @@ struct VoiceState {
FreshStartAudioGeneratorWAV *wav = nullptr;
AudioFileSourceRamWav *ramSource = nullptr;
AudioFileSource *activeSource = nullptr;
StreamManager::SdStream *stream = nullptr; // Set for StreamPath voices.
SamplerMixerInput *stub = nullptr;
BudgetedAudioOutput *budgetedOut = nullptr;
float targetGain = 0.0f; // Per-voice gain from sample volume (0..1), before playback fades.
Expand Down Expand Up @@ -174,6 +180,8 @@ int allocateVoiceSlot(EngineState *impl, int16_t retriggerGroupId, bool &voiceWa
bool beginVoiceFromPath(EngineState *impl,
int voiceIndex,
const String &samplePath,
const uint8_t *head,
uint32_t headBytes,
uint8_t volume,
int16_t retriggerGroupId,
bool loopEnabled,
Expand Down
5 changes: 4 additions & 1 deletion include/sample_classifier.h
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,9 @@ namespace SampleLibrary { struct Catalog; }

namespace SampleClassifier {

constexpr float kFixedPreloadThresholdSeconds = 5.0f;
// Start of each streamed sample kept in RAM (~185 ms), so playback begins
// at once while the SD reader opens the file and catches up.
constexpr uint32_t kStreamHeadBytes = 16 * 1024;
constexpr uint32_t kRequiredSampleRate = 44100;
constexpr uint16_t kRequiredChannelCount = 1;
constexpr uint16_t kRequiredBitsPerSample = 16;
Expand All @@ -33,6 +35,7 @@ struct AssignedSampleClassification {
uint32_t dataOffset = 0;
float durationSeconds = 0.0f;
StorageMode mode = StorageMode::ReadError;
uint32_t headBytes = 0; // Stream only: bytes preloaded from the start.
};

struct ClassificationReport {
Expand Down
8 changes: 8 additions & 0 deletions include/sample_ram_manager.h
Original file line number Diff line number Diff line change
Expand Up @@ -25,17 +25,25 @@ struct LoadReport {
uint32_t usedBytes = 0;
int requestedRamCount = 0;
int loadedRamCount = 0;
int requestedHeadCount = 0;
int loadedHeadCount = 0;
int fallbackToStreamCount = 0;
int readErrorCount = 0;
bool fixedBudgetMismatch = false;
};

// RAM pool size: free PSRAM at the first call minus a reserve, then fixed
// until release(). Falls back to kDefaultSampleRamBudgetBytes without PSRAM.
uint32_t budgetBytes();

bool prepare(const SettingsStore::SamplerSettings &settings,
const SampleClassifier::ClassificationReport &classification,
LoadReport &report);

bool getLoadedSampleByPath(const String &path, LoadedSampleInfo &info);
bool getLoadedSampleDataByPath(const String &path, LoadedSampleData &data);
// Start of a streamed sample, preloaded so playback begins before SD data.
bool getLoadedHeadByPath(const String &path, LoadedSampleData &data);
void release();

} // namespace SampleRamManager
3 changes: 3 additions & 0 deletions include/sampler_app.h
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ class SamplerApp {
void processLoaderCommand(const LoaderCommand &command);
void runLoaderTask();
void runUiTask();
void logStreamingDiagnostics();

Audio audio_;
Input input_;
Expand All @@ -55,4 +56,6 @@ class SamplerApp {
QueueHandle_t uiStatusQueue_ = nullptr;
TaskHandle_t loaderTaskHandle_ = nullptr;
TaskHandle_t uiTaskHandle_ = nullptr;
Audio::StreamingDiagnostics loggedDiagnostics_;
uint32_t lastSdBytesRead_ = 0;
};
6 changes: 5 additions & 1 deletion include/sampler_mixer.h
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@

#include "AudioOutput.h"
#include <cstddef>
#include <cstdint>

namespace AudioInternal {

Expand Down Expand Up @@ -45,14 +46,17 @@ class SamplerMixer {
bool start(int id);
bool consume(int id, float left, float right);
bool emit(float left, float right);
int queued(int id) const;
AudioOutput *sink_;
Frame *mix_ = nullptr;
int capacity_;
int read_ = 0;
bool sinkStarted_ = false;
bool allocated_[kMaxInputs] = {};
bool running_[kMaxInputs] = {};
int queued_[kMaxInputs] = {};
// Absolute frame positions; queued(id) = written_[id] - emitted_.
uint32_t written_[kMaxInputs] = {};
uint32_t emitted_ = 0;
Frame delay_[kLookaheadSamples];
int delayHead_ = 0;
float gain_ = 1;
Expand Down
4 changes: 4 additions & 0 deletions include/storage_sd.h
Original file line number Diff line number Diff line change
@@ -1,7 +1,11 @@
#pragma once

#include <stdint.h>

namespace StorageSD {

bool init();
// SPI clock the card was mounted at, or 0 when not mounted.
uint32_t spiFrequencyHz();

} // namespace StorageSD
Loading
Loading