feat: add AI_VOICE_AI_TIMEOUT_SECONDS for configurable voice timeout
deploy / deploy (push) Successful in 30s
deploy / deploy (push) Successful in 30s
This commit is contained in:
@@ -0,0 +1,47 @@
|
||||
# Streaming TTS produces choppy, syllable-by-syllable audio
|
||||
|
||||
**Status:** open, not started. `AI_VOICE_V2_STREAMING_TTS` is `0` (disabled) in
|
||||
`deployment/aimaq.env.production` until this is fixed — see commit `26d718b`
|
||||
(enabled) and `b8c922c` (reverted after live testing on the Creator plan).
|
||||
|
||||
## Symptom
|
||||
|
||||
With `AI_VOICE_V2_STREAMING_TTS=1`, live calls sound robotic / read
|
||||
syllable-by-syllable ("роботизированно, читает по слогам"), reported by
|
||||
Didar 2026-08-29 after testing on a paid ElevenLabs Creator plan (so it is
|
||||
not a quota/concurrency artifact — that was ruled out separately the same
|
||||
week).
|
||||
|
||||
## Root cause
|
||||
|
||||
`services/ai_voice_runtime_service/media_runtime.py`, `_speak_reply`
|
||||
(~line 1969) only buffers once, at the very start of a reply:
|
||||
|
||||
```python
|
||||
if prebuffered:
|
||||
interrupted = await _write_pcm_frames(pcm_8k) # fed straight through
|
||||
...
|
||||
```
|
||||
|
||||
`_tts_stream_prebuffer_ms` (200ms) absorbs jitter only until the first
|
||||
`_tts_stream_prebuffer_bytes` have arrived. After that, every network chunk
|
||||
from ElevenLabs is written to the AudioSocket the moment it arrives, with no
|
||||
ongoing cushion. `eleven_turbo_v2_5` (and flash models generally) deliver
|
||||
audio over the wire in uneven bursts, not a smooth constant stream — any
|
||||
gap between bursts mid-utterance becomes literal dead air in the outbound
|
||||
audio, which is what reads as "robotic"/"syllable by syllable".
|
||||
|
||||
## Fix direction
|
||||
|
||||
Replace the one-shot prebuffer with a rolling buffer maintained for the
|
||||
whole utterance: keep ~150-200ms of decoded PCM queued ahead of what's
|
||||
being paced out via `_FramePacer`, refilling from the producer queue
|
||||
continuously, instead of switching to pass-through after the first fill.
|
||||
|
||||
## Before re-enabling
|
||||
|
||||
1. Implement the rolling buffer above.
|
||||
2. Re-test live with `AI_VOICE_V2_STREAMING_TTS=1` on the current Creator
|
||||
plan and confirm no gaps/choppiness across a few real calls.
|
||||
3. Only then flip `AI_VOICE_V2_STREAMING_TTS` back to `1` in
|
||||
`deployment/aimaq.env.production`.
|
||||
Reference in New Issue
Block a user