Compare commits

..

6 Commits

Author SHA1 Message Date
bsncubed c0e0388ff0 CLAUDE.md: step 7 progress and open HLS items
HLS plays on AES67; list what to come back to (clock drift vs PTP,
download speed, untested cases) before moving on to cspot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 15:00:09 +10:00
bsncubed 3ebb27e4bb Step 7.4: HLS audio on AES67 (44.1 -> 48 kHz into the ring)
- decoder: s16 PCM -> stereo int32 -> esp_ae_rate_cvt 44.1 -> 48 kHz
  (32-bit, complexity 3; bypassed at 48 kHz; mono duplicated) -> ring.
  esp_audio_effects pinned to ~1.3.0 (1.4+ needs P4 rev >= 3).
- hls: each segment is downloaded completely into PSRAM (max 4 MB), then
  decoded; the connection is not held open while the decoder waits for
  ring space at playback speed.
- player: player_write() blocks while the ring is full and gives up when
  the source changes; the temporary 440 Hz producer is removed.
- Verified with Triple J Hottest: 441344 -> 480375 frames (10.008 s) and
  440320 -> 479260 (9.985 s) per segment; RTP 15000 packets, 0 gaps,
  peak -10 dBFS / RMS -22 dBFS; ring ~4.0 s, 0 underruns over ~50 s;
  heap 422 KB, PSRAM 27.5 MB free.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:52:19 +10:00
bsncubed 8a6cb53e3b Step 7.3: decode HLS TS segments to PCM (own TS demux + AAC decoder)
- main/decoder: MPEG-TS demux (packets reassembled across HTTP chunks,
  PAT -> PMT -> first ADTS-AAC PID, PES headers stripped) feeding the
  esp_audio_codec simple AAC decoder (ADTS, AAC-Plus enabled for HE-AAC
  v1/v2 variants). 7.3 only counts and logs PCM per segment.
- esp_audio_codec pinned to ~2.5.0: 2.6+ needs P4 rev >= 3 (this board
  is rev 1.3). Noted in CLAUDE.md, also for esp_audio_effects < 1.4.
- The library's combined TS decoder lost ~8% of the frames (segments
  decoded to 7.9-9.6 s, "decode error -1"); with the own demux every
  segment is sample exact: 441344 / 440320 frames = 10.008 / 9.985 s,
  matching EXTINF 10.0078 / 9.9846 (431 / 430 AAC frames), no errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:48:12 +10:00
bsncubed 852c0ffa0f Step 7.2: HLS client fetches live segments (log only)
- main/hls: task active while source.mode is hls/auto and hls_url is
  set. esp_http_client over HTTPS (IDF certificate bundle), redirects
  followed (max 5), streamed reads in 4 KB chunks to a sink callback.
- Master playlist: highest BANDWIDTH variant; URL resolution for
  absolute, host-relative and path-relative references; CRLF tolerant.
- Media playlist: TARGETDURATION, MEDIA-SEQUENCE, segments; start 3
  segments behind the live edge, fetch each new one in order, skip ahead
  if the window moved past us, reload after target/2 when nothing is new.
- Verified with Triple J Hottest (ABC, Akamai): 252 kbit/s AAC-LC
  variant chosen, 3 back-fill segments then one new ~10 s segment at a
  time, ~303 KB in ~1.25 s each; heap 439 KB, PSRAM 31.9 MB free.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:40:18 +10:00
bsncubed d3897f3179 Step 7.1: player plumbing (PSRAM ring -> AES67 TX), test tone mode
- main/audio_ring: SPSC ring of interleaved int32 frames in PSRAM
  (4 s at 48 kHz stereo), lock-free with acquire/release counters; the
  TX pull callback never blocks.
- main/player: source selection from source.mode and the pull callback
  for aes67_tx. tone -> the core's PTP-phased 1 kHz tone; off -> silence;
  hls/spotify -> ring with 1 s prefill (silence while buffering is not an
  underrun; running dry is, and prefills again). Status fields
  active_source, source_state, spotify_state, buffer_ms. source config
  applies live.
- New source.mode "tone" (validation, UI dropdown, doc) for commissioning.
- Temporary 440 Hz producer in hls mode (until the HLS player exists).
- Verified: tone 999.7 Hz -18 dBFS; off silent; hls 440.0 Hz with max
  sample step 60816 (ideal sine 60825, i.e. no discontinuities), buffer
  3983 ms, 0 underruns; spotify silent/not implemented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:36:15 +10:00
bsncubed 2a7ab7922d Step 7 prep: enable the 32 MB PSRAM
CONFIG_SPIRAM=y (hex mode, 200 MHz; 250 MHz needs chip rev >= 3).
malloc() places blocks above 16 KB in PSRAM, 32 KB internal RAM stays
reserved, DMA buffers remain internal. Needed for HLS segment buffers,
cspot and decoders.
Verified on board: psram_free 33,551,300 bytes, internal heap 463 KB,
PTP locked, AES67 stream unchanged (1000/s, no gaps, tone within 1 LSB).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:27:04 +10:00
17 changed files with 849 additions and 6 deletions
+9
View File
@@ -25,6 +25,7 @@ Repo: https://gitea.apointless.space/bsncubed/aes67-ESP32-P4
- Chip revision < v3.0 (engineering silicon, seen on some of these boards) needs `CONFIG_ESP32P4_SELECTS_REV_LESS_V3=y`. - Chip revision < v3.0 (engineering silicon, seen on some of these boards) needs `CONFIG_ESP32P4_SELECTS_REV_LESS_V3=y`.
Our board: **v1.3** (set in sdkconfig.defaults, min rev v1.0). Our board: **v1.3** (set in sdkconfig.defaults, min rev v1.0).
- Our board: **32 MB** flash (GigaDevice c8/4019). App slots must stay below 16 MB (cache mapping above 16 MB is experimental in IDF). - Our board: **32 MB** flash (GigaDevice c8/4019). App slots must stay below 16 MB (cache mapping above 16 MB is experimental in IDF).
- Rev < 3 also limits Espressif's prebuilt audio libraries: `esp_audio_codec` must stay < 2.6 and `esp_audio_effects` < 1.4 (newer versions use P4 assembly that needs rev >= 3; the build fails with a message saying so). Check this for any new Espressif binary component.
- Embed `web/index.html` via `EMBED_TXTFILES` in `aes67_web`. - Embed `web/index.html` via `EMBED_TXTFILES` in `aes67_web`.
- Flash over the network (normal way since step 2b; keep USB for recovery): - Flash over the network (normal way since step 2b; keep USB for recovery):
`curl -f --data-binary @build/aes67_p4.bin -H 'Content-Type: application/octet-stream' http://p4-aes67/api/ota` `curl -f --data-binary @build/aes67_p4.bin -H 'Content-Type: application/octet-stream' http://p4-aes67/api/ota`
@@ -51,6 +52,14 @@ Repo: https://gitea.apointless.space/bsncubed/aes67-ESP32-P4
- [x] 5. SAP discovery, then syslog, then health/temperatures (one at a time). VLAN split moved to phase 2. - [x] 5. SAP discovery, then syslog, then health/temperatures (one at a time). VLAN split moved to phase 2.
- [x] 6. PTP TimeTransmitter: BMCA roles (auto/master), hybrid mode. - [x] 6. PTP TimeTransmitter: BMCA roles (auto/master), hybrid mode.
- [ ] 7. Sources: HLS player, then cspot (Spotify Connect), then failover + /api/player. - [ ] 7. Sources: HLS player, then cspot (Spotify Connect), then failover + /api/player.
- [x] 7.1-7.4 HLS plays on AES67 (PSRAM ring, TS demux + AAC, 44.1 -> 48 kHz). Tested with Triple J Hottest (TS, AAC-LC 44.1k).
- [ ] Come back to HLS (open items):
- Clock drift: the station's encoder clock vs our PTP clock is not corrected. Starts 3 segments (~30 s) behind live, so it shows after hours/days (skip when falling out of the live window, or buffering). Fix: steer the 44.1 -> 48 kHz ratio by a few ppm from the distance to the live edge / ring level.
- Download speed ~1.4 Mbit/s over TLS (fine for ~250 kbit/s; tune buffer sizes / per-chunk overhead).
- Not yet tested: HE-AAC variant (140k), other stations, fMP4/ADTS-only playlists, discontinuities (#EXT-X-DISCONTINUITY), network loss and recovery, long runs.
- Audio starts only after PTP lock (~20 s after boot): intended, TX needs PTP.
- [ ] cspot (Spotify Connect)
- [ ] failover (auto mode) + /api/player
- [ ] 8. Mono sum, gain, polish. - [ ] 8. Mono sum, gain, polish.
## Phase 2 (parked) ## Phase 2 (parked)
+31 -1
View File
@@ -1,4 +1,32 @@
dependencies: dependencies:
espressif/esp_audio_codec:
component_hash: 16e2880dbdd5a72264051f750f33a6e5a38fd25621c83e3a32a85a152d2d2643
dependencies:
- name: idf
require: private
version: '>=4.4'
source:
registry_url: https://components.espressif.com/
type: service
version: 2.5.0
espressif/esp_audio_effects:
component_hash: 6ad72ee106e25b07c9ae5b8ec77b566ef8d697473095f1c982e13247005b9ba6
dependencies:
- name: espressif/gmf_fft
registry_url: https://components.espressif.com
require: private
version: ~1.0
source:
registry_url: https://components.espressif.com/
type: service
version: 1.3.0~1
espressif/gmf_fft:
component_hash: b59451ef2f96d22b80c27a162515b21f1b85954dda434c0b092172164028779d
dependencies: []
source:
registry_url: https://components.espressif.com
type: service
version: 1.0.0
espressif/mdns: espressif/mdns:
component_hash: b679eafd0acae2066e2645bd91d073e33ca5be515e2545cd3c4e55a3e3bce3cb component_hash: b679eafd0acae2066e2645bd91d073e33ca5be515e2545cd3c4e55a3e3bce3cb
dependencies: dependencies:
@@ -14,7 +42,9 @@ dependencies:
type: idf type: idf
version: 5.5.5 version: 5.5.5
direct_dependencies: direct_dependencies:
- espressif/esp_audio_codec
- espressif/esp_audio_effects
- espressif/mdns - espressif/mdns
manifest_hash: e5eba11c864076c14a58153c7f761bec2b064244e66ca3f76c0dc170db907329 manifest_hash: 95404bd3492778da35a8aeee53935ee8ad13b09c56ae5fcbe77a4e921dafbb23
target: esp32p4 target: esp32p4
version: 2.0.0 version: 2.0.0
+2 -1
View File
@@ -30,8 +30,9 @@ source (cspot Spotify Connect | HLS player) -> decode (Vorbis/AAC/MP3) -> SRC 44
- Firmware needs Spotify events: cspot connect, disconnect, play, pause. - Firmware needs Spotify events: cspot connect, disconnect, play, pause.
## Project config group ## Project config group
- source: {mode: spotify|hls|auto|off, spotify_name, spotify_bitrate, hls_url, autoplay, gain_db, failover_delay_s, failover_on_pause} - source: {mode: spotify|hls|auto|tone|off, spotify_name, spotify_bitrate, hls_url, autoplay, gain_db, failover_delay_s, failover_on_pause}
- Project status fields: active_source, source_state, spotify_state, buffer_ms - Project status fields: active_source, source_state, spotify_state, buffer_ms
- mode "tone": the core's 1 kHz / -18 dBFS test tone, phase-locked to PTP (commissioning, e.g. Riedel import tests). "off": silence.
## Player control API (project routes) ## Player control API (project routes)
- GET /api/player -> {source, forced, state(playing|paused|stopped|buffering|idle), artist, title, album, position_ms, duration_ms, volume(0-100), can:{pause,next,prev,seek}} - GET /api/player -> {source, forced, state(playing|paused|stopped|buffering|idle), artist, title, album, position_ms, duration_ms, volume(0-100), can:{pause,next,prev,seek}}
+2 -2
View File
@@ -1,5 +1,5 @@
idf_component_register(SRCS "main.c" "project_cfg.c" idf_component_register(SRCS "main.c" "project_cfg.c" "player.c" "audio_ring.c" "hls.c" "decoder.c"
INCLUDE_DIRS "." INCLUDE_DIRS "."
REQUIRES esp_app_format esp_hw_support REQUIRES esp_app_format esp_hw_support heap esp_http_client mbedtls esp_timer
aes67_board aes67_health aes67_net aes67_ota aes67_ptp aes67_sdp_sap aes67_board aes67_health aes67_net aes67_ota aes67_ptp aes67_sdp_sap
aes67_syslog aes67_tx aes67_web) aes67_syslog aes67_tx aes67_web)
+76
View File
@@ -0,0 +1,76 @@
#include "audio_ring.h"
#include <string.h>
#include "esp_heap_caps.h"
static int32_t *s_buf;
static size_t s_cap; // frames
static int s_ch;
// Monotonic frame counters; level = wr - rd. Written by one side each (acquire/release).
static uint32_t s_wr, s_rd;
esp_err_t audio_ring_init(size_t capacity_frames, int channels)
{
s_buf = heap_caps_malloc(capacity_frames * channels * sizeof(int32_t), MALLOC_CAP_SPIRAM);
if (!s_buf) {
return ESP_ERR_NO_MEM;
}
s_cap = capacity_frames;
s_ch = channels;
return ESP_OK;
}
size_t audio_ring_level(void)
{
return __atomic_load_n(&s_wr, __ATOMIC_ACQUIRE) - __atomic_load_n(&s_rd, __ATOMIC_ACQUIRE);
}
size_t audio_ring_space(void)
{
return s_cap - audio_ring_level();
}
// Copy n frames between the ring (starting at frame index pos) and a linear buffer.
static void copy(int32_t *ring_to_lin, const int32_t *lin_to_ring, uint32_t pos, size_t n)
{
size_t start = pos % s_cap;
size_t first = n < s_cap - start ? n : s_cap - start;
size_t fb = first * s_ch * sizeof(int32_t), rb = (n - first) * s_ch * sizeof(int32_t);
if (ring_to_lin) {
memcpy(ring_to_lin, s_buf + start * s_ch, fb);
memcpy(ring_to_lin + first * s_ch, s_buf, rb);
} else {
memcpy(s_buf + start * s_ch, lin_to_ring, fb);
memcpy(s_buf, lin_to_ring + first * s_ch, rb);
}
}
size_t audio_ring_write(const int32_t *frames, size_t n)
{
uint32_t wr = __atomic_load_n(&s_wr, __ATOMIC_RELAXED);
size_t space = s_cap - (wr - __atomic_load_n(&s_rd, __ATOMIC_ACQUIRE));
n = n < space ? n : space;
if (n) {
copy(NULL, frames, wr, n);
__atomic_store_n(&s_wr, wr + n, __ATOMIC_RELEASE);
}
return n;
}
size_t audio_ring_read(int32_t *frames, size_t n)
{
uint32_t rd = __atomic_load_n(&s_rd, __ATOMIC_RELAXED);
size_t level = __atomic_load_n(&s_wr, __ATOMIC_ACQUIRE) - rd;
n = n < level ? n : level;
if (n) {
copy(frames, NULL, rd, n);
__atomic_store_n(&s_rd, rd + n, __ATOMIC_RELEASE);
}
return n;
}
void audio_ring_flush(void)
{
__atomic_store_n(&s_rd, __atomic_load_n(&s_wr, __ATOMIC_ACQUIRE), __ATOMIC_RELEASE);
}
+15
View File
@@ -0,0 +1,15 @@
// Single-producer / single-consumer ring of interleaved int32 frames in PSRAM.
// The producer is a source task, the consumer the AES67 TX pull callback (must never block).
#pragma once
#include <stddef.h>
#include <stdint.h>
#include "esp_err.h"
esp_err_t audio_ring_init(size_t capacity_frames, int channels);
size_t audio_ring_level(void); // frames available to read
size_t audio_ring_space(void); // frames that can be written
size_t audio_ring_write(const int32_t *frames, size_t n); // producer; returns frames written
size_t audio_ring_read(int32_t *frames, size_t n); // consumer; returns frames read
void audio_ring_flush(void); // consumer side: drop everything buffered
+234
View File
@@ -0,0 +1,234 @@
#include "decoder.h"
#include <stdlib.h>
#include <string.h>
#include "esp_audio_dec_default.h"
#include "esp_audio_simple_dec.h"
#include "esp_audio_simple_dec_default.h"
#include "esp_ae_rate_cvt.h"
#include "player.h"
#include "esp_heap_caps.h"
#include "esp_log.h"
static const char *TAG = "decoder";
#define OUT_RATE 48000
#define CONV_FRAMES 4096 // per rate converter call
static esp_ae_rate_cvt_handle_t s_cvt; // NULL: no conversion (source is 48 kHz)
static uint32_t s_cvt_rate; // input rate the converter was opened for
static int32_t *s_in32, *s_out32; // stereo int32 work buffers (PSRAM)
static uint64_t s_seg_out; // 48 kHz frames written this segment
static esp_audio_simple_dec_handle_t s_dec;
static uint8_t *s_pcm;
static uint32_t s_pcm_size = 8192;
static esp_audio_simple_dec_info_t s_info;
static uint64_t s_seg_frames, s_seg_count, s_seg_es;
// MPEG-TS demux state: packets reassembled across HTTP chunks, PAT -> PMT -> ADTS-AAC PID.
static uint8_t s_pkt[188];
static size_t s_pkt_len;
static int s_pmt_pid = -1, s_audio_pid = -1;
static void report_segment(void)
{
if (s_seg_count && s_info.sample_rate) {
ESP_LOGI(TAG, "segment: %llu ES bytes -> %llu frames at %lu Hz = %.3f s -> %llu frames at 48 kHz = %.3f s",
s_seg_es, s_seg_frames, (unsigned long)s_info.sample_rate,
(double)s_seg_frames / s_info.sample_rate, s_seg_out, (double)s_seg_out / OUT_RATE);
}
s_seg_frames = 0;
s_seg_es = 0;
s_seg_out = 0;
s_seg_count++;
}
// Decoded PCM (s16, any channel count) -> stereo int32 -> 48 kHz -> ring.
static void output_pcm(const int16_t *pcm, uint32_t frames)
{
int ch = s_info.channel;
if (s_info.sample_rate != s_cvt_rate) {
if (s_cvt) {
esp_ae_rate_cvt_close(s_cvt);
s_cvt = NULL;
}
s_cvt_rate = s_info.sample_rate;
if (s_cvt_rate != OUT_RATE) {
esp_ae_rate_cvt_cfg_t cfg = {
.src_rate = s_cvt_rate, .dest_rate = OUT_RATE, .channel = 2, .bits_per_sample = 32,
.complexity = 3, .perf_type = ESP_AE_RATE_CVT_PERF_TYPE_SPEED,
};
if (esp_ae_rate_cvt_open(&cfg, &s_cvt) != ESP_AE_ERR_OK) {
ESP_LOGE(TAG, "rate converter %lu -> %d Hz: open failed", (unsigned long)s_cvt_rate, OUT_RATE);
} else {
ESP_LOGI(TAG, "rate converter %lu -> %d Hz", (unsigned long)s_cvt_rate, OUT_RATE);
}
}
}
while (frames) {
uint32_t n = frames < CONV_FRAMES ? frames : CONV_FRAMES;
for (uint32_t i = 0; i < n; i++) { // to stereo int32 (full scale = INT32_MAX)
int32_t l = (int32_t)pcm[i * ch] << 16;
s_in32[i * 2] = l;
s_in32[i * 2 + 1] = ch > 1 ? (int32_t)pcm[i * ch + 1] << 16 : l;
}
const int32_t *out = s_in32;
uint32_t out_n = n;
if (s_cvt) {
out_n = CONV_FRAMES * 2;
if (esp_ae_rate_cvt_process(s_cvt, s_in32, n, s_out32, &out_n) != ESP_AE_ERR_OK) {
ESP_LOGW(TAG, "rate conversion failed");
return;
}
out = s_out32;
}
s_seg_out += player_write(out, out_n);
pcm += n * ch;
frames -= n;
}
}
// Feed ADTS-AAC elementary stream bytes to the decoder.
static void decode_es(const uint8_t *data, size_t len)
{
s_seg_es += len;
esp_audio_simple_dec_raw_t raw = { .buffer = (uint8_t *)data, .len = len };
while (raw.len) {
esp_audio_simple_dec_out_t out = { .buffer = s_pcm, .len = s_pcm_size };
esp_audio_err_t err = esp_audio_simple_dec_process(s_dec, &raw, &out);
if (err == ESP_AUDIO_ERR_BUFF_NOT_ENOUGH) {
uint8_t *p = heap_caps_realloc(s_pcm, out.needed_size, MALLOC_CAP_SPIRAM);
if (!p) {
return;
}
s_pcm = p;
s_pcm_size = out.needed_size;
continue;
}
if (err != ESP_AUDIO_ERR_OK) {
ESP_LOGW(TAG, "decode error %d, resetting decoder", err);
esp_audio_simple_dec_reset(s_dec);
return;
}
if (out.decoded_size) {
if (!s_info.sample_rate) {
esp_audio_simple_dec_get_info(s_dec, &s_info);
ESP_LOGI(TAG, "stream: %lu Hz, %u ch, %u bit, %lu bit/s", (unsigned long)s_info.sample_rate,
s_info.channel, s_info.bits_per_sample, (unsigned long)s_info.bitrate);
}
uint32_t n = out.decoded_size / (s_info.channel * s_info.bits_per_sample / 8);
s_seg_frames += n;
if (s_info.bits_per_sample == 16) {
output_pcm((const int16_t *)out.buffer, n);
}
}
raw.buffer += raw.consumed;
raw.len -= raw.consumed;
}
}
static void ts_packet(const uint8_t *p)
{
if (p[0] != 0x47) {
return;
}
bool pusi = p[1] & 0x40;
int pid = ((p[1] & 0x1f) << 8) | p[2];
int afc = (p[3] >> 4) & 3;
if (!(afc & 1)) {
return; // no payload
}
size_t off = 4 + ((afc & 2) ? 1 + p[4] : 0);
if (off >= 188) {
return;
}
const uint8_t *pl = p + off;
size_t len = 188 - off;
if (pid == 0 && pusi) { // PAT: first program's PMT PID
const uint8_t *sec = pl + 1 + pl[0];
s_pmt_pid = ((sec[10] & 0x1f) << 8) | sec[11];
} else if (pid == s_pmt_pid && pusi && s_audio_pid < 0) { // PMT: first ADTS-AAC stream
const uint8_t *sec = pl + 1 + pl[0];
int slen = ((sec[1] & 0x0f) << 8) | sec[2];
int i = 12 + (((sec[10] & 0x0f) << 8) | sec[11]);
while (i + 5 <= 3 + slen - 4 && sec + i + 5 <= p + 188) {
int type = sec[i], epid = ((sec[i + 1] & 0x1f) << 8) | sec[i + 2];
if (type == 0x0F) {
s_audio_pid = epid;
ESP_LOGI(TAG, "TS: PMT PID 0x%x, ADTS-AAC on PID 0x%x", s_pmt_pid, epid);
break;
}
i += 5 + (((sec[i + 3] & 0x0f) << 8) | sec[i + 4]);
}
} else if (pid == s_audio_pid) {
if (pusi) { // strip the PES header
if (len < 9 || pl[0] || pl[1] || pl[2] != 1 || 9u + pl[8] > len) {
return;
}
size_t h = 9 + pl[8];
pl += h;
len -= h;
}
decode_es(pl, len);
}
}
// hls_sink_t: MPEG-TS bytes in any chunking.
bool decoder_feed(const uint8_t *data, size_t len, bool segment_start)
{
if (segment_start) {
report_segment();
s_pkt_len = 0; // segments start on a packet boundary
}
while (len) {
if (!s_pkt_len) {
// resync on 0x47 if needed, then take whole packets straight from the input
while (len && data[0] != 0x47) {
data++;
len--;
}
while (len >= 188 && data[0] == 0x47) {
ts_packet(data);
data += 188;
len -= 188;
}
if (len && data[0] != 0x47) {
continue;
}
}
size_t n = 188 - s_pkt_len < len ? 188 - s_pkt_len : len;
memcpy(s_pkt + s_pkt_len, data, n);
s_pkt_len += n;
data += n;
len -= n;
if (s_pkt_len == 188) {
ts_packet(s_pkt);
s_pkt_len = 0;
}
}
return true;
}
esp_err_t decoder_init(void)
{
esp_audio_dec_register_default();
esp_audio_simple_dec_register_default();
// Own TS demux (the library's TS decoder lost ~8% of the frames); AAC with ADTS headers,
// AAC-Plus on so HE-AAC v1/v2 variants work too.
esp_aac_dec_cfg_t aac_cfg = ESP_AAC_DEC_CONFIG_DEFAULT();
aac_cfg.aac_plus_enable = true;
esp_audio_simple_dec_cfg_t cfg = {
.dec_type = ESP_AUDIO_SIMPLE_DEC_TYPE_AAC, .dec_cfg = &aac_cfg, .cfg_size = sizeof(aac_cfg),
};
if (esp_audio_simple_dec_open(&cfg, &s_dec) != ESP_AUDIO_ERR_OK) {
ESP_LOGE(TAG, "AAC decoder open failed");
return ESP_FAIL;
}
s_pcm = heap_caps_malloc(s_pcm_size, MALLOC_CAP_SPIRAM);
s_in32 = heap_caps_malloc(CONV_FRAMES * 2 * sizeof(int32_t), MALLOC_CAP_SPIRAM);
s_out32 = heap_caps_malloc(CONV_FRAMES * 2 * 2 * sizeof(int32_t), MALLOC_CAP_SPIRAM);
return s_pcm && s_in32 && s_out32 ? ESP_OK : ESP_ERR_NO_MEM;
}
+12
View File
@@ -0,0 +1,12 @@
// Decoder for HLS segments (MPEG-TS with AAC-LC / HE-AAC), fed by the HLS client's sink.
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include "esp_err.h"
esp_err_t decoder_init(void);
// hls_sink_t: segment bytes in; PCM out (7.3: counted and logged).
bool decoder_feed(const uint8_t *data, size_t len, bool segment_start);
+305
View File
@@ -0,0 +1,305 @@
#include "hls.h"
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include "aes67_cfg.h"
#include "esp_crt_bundle.h"
#include "esp_heap_caps.h"
#include "esp_http_client.h"
#include "esp_log.h"
#include "esp_timer.h"
#include "freertos/FreeRTOS.h"
#include "freertos/task.h"
#define URL_MAX 384
#define PLAYLIST_MAX (64 * 1024)
#define MAX_SEGMENTS 64
#define LIVE_BACK 3 // start this many segments behind the live edge
#define CHUNK 4096
#define MAX_REDIRECTS 5
#define SEG_MAX (4 * 1024 * 1024) // largest segment we accept (PSRAM)
static const char *TAG = "hls";
typedef struct {
int target_s;
long long media_seq;
int count;
char uri[MAX_SEGMENTS][URL_MAX];
} media_pl_t;
static hls_sink_t s_sink;
static char *s_text; // playlist buffer (PSRAM)
static media_pl_t *s_pl; // parsed media playlist (PSRAM)
static char s_url[URL_MAX]; // configured URL
static char s_media_url[URL_MAX];
static uint8_t *s_seg; // current segment (PSRAM)
// hls_url when source.mode needs HLS, else "".
static void wanted_url(char *out, size_t n)
{
cJSON *src = cfg_get("source");
const char *mode = cJSON_GetObjectItemCaseSensitive(src, "mode")->valuestring;
const char *url = cJSON_GetObjectItemCaseSensitive(src, "hls_url")->valuestring;
bool on = !strcmp(mode, "hls") || !strcmp(mode, "auto");
strlcpy(out, on ? url : "", n);
cJSON_Delete(src);
}
static bool still_wanted(void)
{
char u[URL_MAX];
wanted_url(u, sizeof(u));
return strcmp(u, s_url) == 0;
}
/* ----- HTTP ----- */
// Open a GET and follow redirects. Returns the client with headers read, or NULL.
static esp_http_client_handle_t http_open(const char *url)
{
esp_http_client_config_t cfg = {
.url = url, .crt_bundle_attach = esp_crt_bundle_attach, .timeout_ms = 10000,
.buffer_size = CHUNK, .buffer_size_tx = 1024, .user_agent = "p4-aes67",
.disable_auto_redirect = true, .keep_alive_enable = true,
};
esp_http_client_handle_t c = esp_http_client_init(&cfg);
if (!c) {
return NULL;
}
for (int i = 0; i <= MAX_REDIRECTS; i++) {
esp_err_t err = esp_http_client_open(c, 0);
if (err != ESP_OK) {
ESP_LOGW(TAG, "open %s: %s", url, esp_err_to_name(err));
break;
}
esp_http_client_fetch_headers(c);
int st = esp_http_client_get_status_code(c);
if (st == 301 || st == 302 || st == 303 || st == 307 || st == 308) {
esp_http_client_flush_response(c, NULL);
esp_http_client_set_redirection(c);
esp_http_client_close(c);
continue;
}
if (st == 200) {
return c;
}
ESP_LOGW(TAG, "GET %s: HTTP %d", url, st);
break;
}
esp_http_client_cleanup(c);
return NULL;
}
// GET into s_text (NUL-terminated). Returns length or -1.
static int http_get_text(const char *url)
{
esp_http_client_handle_t c = http_open(url);
if (!c) {
return -1;
}
int n = 0, r;
while (n < PLAYLIST_MAX - 1 && (r = esp_http_client_read(c, s_text + n, PLAYLIST_MAX - 1 - n)) > 0) {
n += r;
}
s_text[n] = 0;
esp_http_client_close(c);
esp_http_client_cleanup(c);
return n;
}
/* ----- playlists ----- */
// Absolute, host-relative ("/x") or path-relative ("x") reference against base.
static void resolve(char *out, size_t n, const char *base, const char *ref)
{
if (strstr(ref, "://")) {
strlcpy(out, ref, n);
return;
}
char b[URL_MAX];
strlcpy(b, base, sizeof(b));
char *q = strchr(b, '?');
if (q) {
*q = 0;
}
if (ref[0] == '/') {
char *host_end = strchr(strstr(b, "://") + 3, '/');
if (host_end) {
*host_end = 0;
}
} else {
char *slash = strrchr(b, '/');
if (slash) {
slash[1] = 0;
}
}
snprintf(out, n, "%s%s", b, ref);
}
// Next line of s_text (CRLF tolerant); NULL at the end.
static char *next_line(char **p)
{
if (!*p || !**p) {
return NULL;
}
char *line = *p, *e = strchr(line, '\n');
*p = e ? e + 1 : NULL;
if (e) {
*e = 0;
}
size_t l = strlen(line);
while (l && (line[l - 1] == '\r' || line[l - 1] == ' ')) {
line[--l] = 0;
}
return line;
}
// Master playlist: pick the variant with the highest BANDWIDTH. false if s_text is a media playlist.
static bool pick_variant(const char *base, char *out, size_t n)
{
if (!strstr(s_text, "#EXT-X-STREAM-INF")) {
return false;
}
long best = -1, bw = -1;
char *p = s_text, *line;
while ((line = next_line(&p))) {
if (!strncmp(line, "#EXT-X-STREAM-INF:", 18)) {
const char *b = strstr(line, "BANDWIDTH=");
bw = b ? atol(b + 10) : 0;
} else if (line[0] && line[0] != '#' && bw >= 0) {
if (bw > best) {
best = bw;
resolve(out, n, base, line);
}
bw = -1;
}
}
ESP_LOGI(TAG, "variant %ld bit/s: %s", best, out);
return best >= 0;
}
static bool parse_media(const char *base)
{
s_pl->target_s = 10;
s_pl->media_seq = 0;
s_pl->count = 0;
char *p = s_text, *line;
bool is_m3u = false;
while ((line = next_line(&p))) {
if (!strcmp(line, "#EXTM3U")) {
is_m3u = true;
} else if (!strncmp(line, "#EXT-X-TARGETDURATION:", 22)) {
s_pl->target_s = atoi(line + 22);
} else if (!strncmp(line, "#EXT-X-MEDIA-SEQUENCE:", 22)) {
s_pl->media_seq = atoll(line + 22);
} else if (line[0] && line[0] != '#' && s_pl->count < MAX_SEGMENTS) {
resolve(s_pl->uri[s_pl->count++], URL_MAX, base, line);
}
}
return is_m3u && s_pl->count > 0;
}
/* ----- segments ----- */
static bool fetch_segment(long long seq, const char *url)
{
int64_t t0 = esp_timer_get_time();
esp_http_client_handle_t c = http_open(url);
if (!c) {
return false;
}
// Download completely first, then decode: the connection is not held open while the
// decoder waits for space in the ring (a stalled TCP stream for ~10 s may get cut by the CDN).
size_t total = 0;
int r = 0;
bool ok = true;
while (total < SEG_MAX && (r = esp_http_client_read(c, (char *)s_seg + total, SEG_MAX - total)) > 0) {
total += r;
}
if (r < 0 || total == SEG_MAX) {
ok = false;
}
esp_http_client_close(c);
esp_http_client_cleanup(c);
int64_t us = esp_timer_get_time() - t0;
ESP_LOGD(TAG, "segment %lld: %u bytes in %.2f s (%.1f Mbit/s)%s", seq, (unsigned)total, us / 1e6,
us ? total * 8.0 / us : 0.0, ok ? "" : ", failed");
if (ok && s_sink) {
ok = s_sink(s_seg, total, true); // blocks at playback speed while the ring is full
}
return ok;
}
static void hls_task(void *arg)
{
while (1) {
wanted_url(s_url, sizeof(s_url));
if (!s_url[0]) {
vTaskDelay(pdMS_TO_TICKS(1000));
continue;
}
ESP_LOGI(TAG, "start: %s", s_url);
// Master -> media playlist
if (http_get_text(s_url) < 0) {
vTaskDelay(pdMS_TO_TICKS(5000));
continue;
}
if (!pick_variant(s_url, s_media_url, sizeof(s_media_url))) {
strlcpy(s_media_url, s_url, sizeof(s_media_url));
}
long long next = -1;
int errors = 0;
while (still_wanted() && errors < 5) {
int64_t t0 = esp_timer_get_time();
if (http_get_text(s_media_url) < 0 || !parse_media(s_media_url)) {
errors++;
vTaskDelay(pdMS_TO_TICKS(2000));
continue;
}
long long first = s_pl->media_seq, last = first + s_pl->count - 1;
if (next < 0) {
next = last - LIVE_BACK + 1 > first ? last - LIVE_BACK + 1 : first;
ESP_LOGI(TAG, "live playlist: seq %lld..%lld, target %d s, starting at %lld",
first, last, s_pl->target_s, next);
} else if (next < first) {
ESP_LOGW(TAG, "fell behind the live window (%lld < %lld), skipping ahead", next, first);
next = first;
}
bool got_new = false;
for (; next <= last && still_wanted(); next++) {
if (!fetch_segment(next, s_pl->uri[next - first])) {
errors++;
break;
}
errors = 0;
got_new = true;
}
// No new segment: wait half the target duration before reloading (RFC 8216 6.3.4).
if (!got_new) {
int64_t wait = s_pl->target_s * 500000LL - (esp_timer_get_time() - t0);
if (wait > 0) {
vTaskDelay(pdMS_TO_TICKS(wait / 1000));
}
}
}
if (errors >= 5) {
ESP_LOGW(TAG, "too many errors, restarting from the playlist in 5 s");
vTaskDelay(pdMS_TO_TICKS(5000));
}
}
}
esp_err_t hls_start(hls_sink_t sink)
{
s_sink = sink;
s_text = heap_caps_malloc(PLAYLIST_MAX, MALLOC_CAP_SPIRAM);
s_pl = heap_caps_malloc(sizeof(media_pl_t), MALLOC_CAP_SPIRAM);
s_seg = heap_caps_malloc(SEG_MAX, MALLOC_CAP_SPIRAM);
if (!s_text || !s_pl || !s_seg) {
return ESP_ERR_NO_MEM;
}
return xTaskCreate(hls_task, "hls", 10240, NULL, 4, NULL) == pdPASS ? ESP_OK : ESP_ERR_NO_MEM;
}
+13
View File
@@ -0,0 +1,13 @@
// HLS client: playlists, variant choice, live segment loop. Segment bytes go to a sink callback.
#pragma once
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include "esp_err.h"
// Receives segment data in chunks (MPEG-TS etc.). Return false to abort the segment.
typedef bool (*hls_sink_t)(const uint8_t *data, size_t len, bool segment_start);
esp_err_t hls_start(hls_sink_t sink);
+4
View File
@@ -0,0 +1,4 @@
dependencies:
# Newer versions use P4 assembly that needs chip rev >= 3.0; this board is rev 1.3 (see CLAUDE.md).
espressif/esp_audio_codec: "~2.5.0"
espressif/esp_audio_effects: "~1.3.0"
+2
View File
@@ -10,6 +10,7 @@
#include "esp_app_desc.h" #include "esp_app_desc.h"
#include "esp_chip_info.h" #include "esp_chip_info.h"
#include "esp_log.h" #include "esp_log.h"
#include "player.h"
#include "project_cfg.h" #include "project_cfg.h"
static const char *TAG = "main"; static const char *TAG = "main";
@@ -39,6 +40,7 @@ void app_main(void)
project_cfg_register(); project_cfg_register();
ESP_ERROR_CHECK(aes67_ota_init()); ESP_ERROR_CHECK(aes67_ota_init());
ESP_ERROR_CHECK(aes67_sdp_sap_init()); ESP_ERROR_CHECK(aes67_sdp_sap_init());
ESP_ERROR_CHECK(player_init());
ESP_ERROR_CHECK(aes67_tx_start()); ESP_ERROR_CHECK(aes67_tx_start());
ESP_ERROR_CHECK(aes67_web_start()); ESP_ERROR_CHECK(aes67_web_start());
} }
+122
View File
@@ -0,0 +1,122 @@
#include "player.h"
#include <string.h>
#include "aes67_cfg.h"
#include "aes67_tx.h"
#include "aes67_web.h"
#include "audio_ring.h"
#include "decoder.h"
#include "hls.h"
#include "esp_log.h"
#include "freertos/FreeRTOS.h"
#include "freertos/task.h"
#define RATE 48000
#define CHANNELS 2
#define RING_FRAMES (4 * RATE) // 4 s in PSRAM
#define PREFILL_FRAMES (RATE) // 1 s before (re)starting output
static const char *TAG = "player";
typedef enum { SRC_TONE, SRC_OFF, SRC_HLS, SRC_SPOTIFY } src_t;
static const char *const SRC_NAME[] = { "tone", "off", "hls", "spotify" };
static volatile src_t s_src = SRC_OFF;
static volatile bool s_playing; // ring output running (after prefill)
static volatile bool s_flush; // consumer drops buffered audio on the next read
static src_t mode_of(const char *m)
{
return !strcmp(m, "tone") ? SRC_TONE : !strcmp(m, "hls") ? SRC_HLS :
!strcmp(m, "spotify") ? SRC_SPOTIFY : !strcmp(m, "auto") ? SRC_HLS /* failover: step 7 */ : SRC_OFF;
}
// AES67 TX pull callback (TX task, must not block). Silence while buffering or off: not an underrun.
static size_t player_read(int32_t *buf, size_t frames)
{
if (s_flush) {
audio_ring_flush();
s_flush = false;
s_playing = false;
}
if (s_src == SRC_OFF || s_src == SRC_SPOTIFY) { // Spotify: step 7 (cspot)
memset(buf, 0, frames * CHANNELS * sizeof(int32_t));
return frames;
}
if (!s_playing) {
if (audio_ring_level() < PREFILL_FRAMES) {
memset(buf, 0, frames * CHANNELS * sizeof(int32_t));
return frames;
}
s_playing = true;
}
size_t got = audio_ring_read(buf, frames);
if (got < frames) {
s_playing = false; // ran dry: underrun (counted by TX), prefill again
}
return got;
}
void player_apply(const cJSON *source)
{
src_t src = mode_of(cJSON_GetObjectItemCaseSensitive(source, "mode")->valuestring);
if (src == s_src) {
return;
}
s_src = src;
s_flush = true;
// The test tone is TX's own (phase-locked to PTP); everything else goes through player_read.
aes67_tx_set_source(src == SRC_TONE ? NULL : player_read);
ESP_LOGI(TAG, "source: %s", SRC_NAME[src]);
}
// Source side (HLS task): write converted 48 kHz frames, waiting for space at playback speed.
// Gives up when the source is no longer HLS.
size_t player_write(const int32_t *frames, size_t n)
{
size_t done = 0;
while (done < n) {
if (s_src != SRC_HLS) {
return done;
}
size_t w = audio_ring_write(frames + done * CHANNELS, n - done);
if (!w) {
vTaskDelay(pdMS_TO_TICKS(20));
}
done += w;
}
return done;
}
static void player_status(cJSON *st)
{
const char *state = s_src == SRC_TONE ? "playing" : s_src == SRC_OFF ? "idle" :
s_src == SRC_SPOTIFY ? "not implemented" : s_playing ? "playing" : "buffering";
cJSON_AddStringToObject(st, "active_source", SRC_NAME[s_src]);
cJSON_AddStringToObject(st, "source_state", state);
cJSON_AddStringToObject(st, "spotify_state", "not implemented");
cJSON_AddNumberToObject(st, "buffer_ms", (double)(audio_ring_level() * 1000 / RATE));
}
esp_err_t player_init(void)
{
esp_err_t err = audio_ring_init(RING_FRAMES, CHANNELS);
if (err != ESP_OK) {
ESP_LOGE(TAG, "ring buffer: %s", esp_err_to_name(err));
return err;
}
cJSON *src = cfg_get("source");
s_src = (src_t)-1;
player_apply(src);
cJSON_Delete(src);
err = decoder_init();
if (err != ESP_OK) {
return err;
}
err = hls_start(decoder_feed);
if (err != ESP_OK) {
return err;
}
return status_register(player_status);
}
+14
View File
@@ -0,0 +1,14 @@
// Project player: picks the audio source for AES67 TX (test tone, silence, HLS, Spotify).
#pragma once
#include <stddef.h>
#include <stdint.h>
#include "cJSON.h"
#include "esp_err.h"
esp_err_t player_init(void);
// Apply the "source" config group (mode etc.).
void player_apply(const cJSON *source);
// Source side: write 48 kHz stereo int32 frames, blocking while the ring is full.
size_t player_write(const int32_t *frames, size_t n);
+3 -2
View File
@@ -3,6 +3,7 @@
#include "aes67_cfg.h" #include "aes67_cfg.h"
#include "esp_err.h" #include "esp_err.h"
#include "player.h"
static const char SOURCE_DEFAULTS[] = static const char SOURCE_DEFAULTS[] =
"{\"mode\":\"spotify\",\"spotify_name\":\"P4 AES67\",\"spotify_bitrate\":320,\"hls_url\":\"\"," "{\"mode\":\"spotify\",\"spotify_name\":\"P4 AES67\",\"spotify_bitrate\":320,\"hls_url\":\"\","
@@ -10,7 +11,7 @@ static const char SOURCE_DEFAULTS[] =
static bool source_validate(const cJSON *g, char *err, size_t n) static bool source_validate(const cJSON *g, char *err, size_t n)
{ {
static const char *const modes[] = { "spotify", "hls", "auto", "off", NULL }; static const char *const modes[] = { "spotify", "hls", "auto", "tone", "off", NULL };
static const double bitrates[] = { 96, 160, 320 }; static const double bitrates[] = { 96, 160, 320 };
return cfg_check_enum(g, "mode", modes, err, n) && return cfg_check_enum(g, "mode", modes, err, n) &&
cfg_check_str(g, "spotify_name", 1, 63, err, n) && cfg_check_str(g, "spotify_name", 1, 63, err, n) &&
@@ -29,5 +30,5 @@ void project_cfg_defaults(void)
void project_cfg_register(void) void project_cfg_register(void)
{ {
ESP_ERROR_CHECK(cfg_register("source", SOURCE_DEFAULTS, source_validate, NULL)); ESP_ERROR_CHECK(cfg_register("source", SOURCE_DEFAULTS, source_validate, player_apply));
} }
+4
View File
@@ -24,3 +24,7 @@ CONFIG_ETH_TRANSMIT_MUTEX=y
# use 7, leaving ~3 for web clients (httpd allows 7): a browser's keep-alive connections then # use 7, leaving ~3 for web clients (httpd allows 7): a browser's keep-alive connections then
# starve new requests. 16 = 7 others + 7 web clients + OTA self-test client + 1 spare. # starve new requests. 16 = 7 others + 7 web clients + OTA self-test client + 1 spare.
CONFIG_LWIP_MAX_SOCKETS=16 CONFIG_LWIP_MAX_SOCKETS=16
# 32 MB PSRAM in the P4 package (hex mode, 200 MHz; 250 MHz needs rev >= 3). Needed for step 7:
# HLS segment buffers, cspot, decoders. malloc() puts blocks >16 KB in PSRAM; DMA buffers stay internal.
CONFIG_SPIRAM=y
+1
View File
@@ -64,6 +64,7 @@ e.g. curl -X POST http://p4-aes67.local/api/player/next
<option value="spotify">Spotify Connect</option> <option value="spotify">Spotify Connect</option>
<option value="hls">HLS / M3U8 URL</option> <option value="hls">HLS / M3U8 URL</option>
<option value="auto">Spotify, fail over to HLS</option> <option value="auto">Spotify, fail over to HLS</option>
<option value="tone">Test tone (1 kHz, -18 dBFS)</option>
<option value="off">Off (silence)</option></select></label> <option value="off">Off (silence)</option></select></label>
<label><span>Failover delay (s)</span><input name="source.failover_delay_s" type="number" min="0" max="3600"></label> <label><span>Failover delay (s)</span><input name="source.failover_delay_s" type="number" min="0" max="3600"></label>
<label><span>Also fail over when paused</span><input name="source.failover_on_pause" type="checkbox"></label> <label><span>Also fail over when paused</span><input name="source.failover_on_pause" type="checkbox"></label>