Commit Graph

3 Commits

Author SHA1 Message Date
bsncubed 5b0d7c9ba4 Step 7.5b3a: per-source PCM converter (audio_out), shared by HLS and Spotify
- main/audio_out: one converter instance per source (own rate converter
  state): s16 any rate/channels -> 48 kHz stereo int32 -> ring.
- player_write(src, ...) only writes for the active source, drops the
  rest; player_src_t is public in player.h.
- HLS decoder uses audio_out (no behaviour change).
- Verified: HLS still sample exact (441344 -> 480375 frames per
  10.008 s segment), buffer ~4 s, 0 underruns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 18:34:15 +10:00
bsncubed 3ebb27e4bb Step 7.4: HLS audio on AES67 (44.1 -> 48 kHz into the ring)
- decoder: s16 PCM -> stereo int32 -> esp_ae_rate_cvt 44.1 -> 48 kHz
  (32-bit, complexity 3; bypassed at 48 kHz; mono duplicated) -> ring.
  esp_audio_effects pinned to ~1.3.0 (1.4+ needs P4 rev >= 3).
- hls: each segment is downloaded completely into PSRAM (max 4 MB), then
  decoded; the connection is not held open while the decoder waits for
  ring space at playback speed.
- player: player_write() blocks while the ring is full and gives up when
  the source changes; the temporary 440 Hz producer is removed.
- Verified with Triple J Hottest: 441344 -> 480375 frames (10.008 s) and
  440320 -> 479260 (9.985 s) per segment; RTP 15000 packets, 0 gaps,
  peak -10 dBFS / RMS -22 dBFS; ring ~4.0 s, 0 underruns over ~50 s;
  heap 422 KB, PSRAM 27.5 MB free.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:52:19 +10:00
bsncubed 8a6cb53e3b Step 7.3: decode HLS TS segments to PCM (own TS demux + AAC decoder)
- main/decoder: MPEG-TS demux (packets reassembled across HTTP chunks,
  PAT -> PMT -> first ADTS-AAC PID, PES headers stripped) feeding the
  esp_audio_codec simple AAC decoder (ADTS, AAC-Plus enabled for HE-AAC
  v1/v2 variants). 7.3 only counts and logs PCM per segment.
- esp_audio_codec pinned to ~2.5.0: 2.6+ needs P4 rev >= 3 (this board
  is rev 1.3). Noted in CLAUDE.md, also for esp_audio_effects < 1.4.
- The library's combined TS decoder lost ~8% of the frames (segments
  decoded to 7.9-9.6 s, "decode error -1"); with the own demux every
  segment is sample exact: 441344 / 440320 frames = 10.008 / 9.985 s,
  matching EXTINF 10.0078 / 9.9846 (431 / 430 AAC frames), no errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 14:48:12 +10:00