The speculative "early plan" turn (computed on partial ASR, before the
caller finishes talking) can win the race and get spoken as the actual
reply, but it unconditionally skipped KB search and answered common
questions (schedule/address/price/status/problem) with a hardcoded
clarifying question even when the FAQ already had the answer.
KB search is a cheap in-memory lexical scan over a DB-cached row set,
so it fits the early-plan latency budget unlike a real LLM call. Now
early-plan runs it and, on a match, answers from the KB snippet
(intent resolved via normalize_intent) instead of guessing a generic
clarifying question; with no match it falls back to the prior
behavior unchanged. operator_request is unaffected.
Centralizes fixed control intents and adds a data-driven intent_code
field on kb_articles so many phrasings of the same FAQ question
resolve to one stable code (e.g. VOUCHER_ACTIVATION) instead of a
free-form, unvalidated string the LLM invented on the fly.
- services/shared/intents.py: CONTROL_INTENTS + normalize_intent()
- kb_articles.intent_code column (ORM + dev/sqlite runtime compat +
migrations/sql/0034_* for postgres/sqlite)
- kb_service CRUD exposes intent_code
- orchestrator surfaces intent_code to the LLM and validates its
intent output against control intents + the KB codes shown that turn
- voice.py: _voice_early_intent_bucket renamed to _voice_ack_topic_bucket
to stop it being conflated with the canonical FAQ intent
Both services/ai_orchestrator_service/voice.py and the aimaq persona/DOMAIN
SCOPE prompt were bind-mounted straight from the host on the aimaq stack,
bypassing git and CI entirely since they were first hand-edited in prod.
voice.py: merged the host's live business logic (gas/aimaq domain keyword
list, off-domain Kazakh/Russian replies, disabled re-correction of an already
obtained name) with the timeout_seconds fix from 09bcf74 that never reached
aimaq because the bind mount blocked it.
aimaq.env.production: replaced the AI_OPERATOR_* env values (which were never
interpolated -- {agent_name}/{company_name} would have been read literally)
with the final resolved Zhanna/Kazakgaz Aimaq text including the DOMAIN SCOPE
clause, matching what was actually live on the host.
RuntimeConfiguredTTSProvider (the live 'dynamic' TTS provider, resolves
voice from DB-backed config) never overrode cache_fingerprint(), so it
fell back to the base class's empty string. PrebakedAckBank keys its
on-disk cache on that fingerprint, so short filler phrases like
'Секунду' kept serving audio baked with the previous ElevenLabs voice
even after a voice change, while full LLM replies (cached inside the
resolved provider itself, keyed on its own voice id) already used the
new voice — explaining why callers heard two different voices in the
same call. Fix: delegate cache_fingerprint() to the resolved provider.
Also extend the voice delivery_hint so the model transliterates website
addresses and English words/abbreviations into spoken Cyrillic instead
of leaving raw Latin text for the TTS engine to mangle (egov.kz was
coming out as 'эговкз').
TTS reads raw text digit-by-digit with no number/date normalization layer.
Short numbers like 1414 were read as a single 4-digit number (tysyacha
chetyresta chetyrnadtsat) instead of a spoken code, and dates like '25
chisla' were read in the wrong grammatical case (dvadtsat pyat chislo
instead of dvadtsat pyatogo chisla). Extend the voice delivery_hint with
explicit spell-out rules so the model itself produces already-correct
spoken-form text.
Add sync_ai_operator_config_from_code() which overwrites the DB-cached
ai_operator_settings row from ai_operator_default_config() on startup
of ai_orchestrator_service and ai_voice_runtime_service. Greeting and
system prompt changes now go through git + deploy instead of manual
psql/API edits to prod. Also adds a root README pointing to existing docs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two independent latency fixes for the voice-assistant reply pipeline,
both scoped to the parts of the flow that run regardless of whether
voice_v2 is enabled for a queue:
1. media_runtime._process_utterance: the v1/fallback turn path (used by
any queue not covered by AI_VOICE_V2_QUEUE_CODES) silently awaited the
full LLM decision with no audio playing at all, unlike the v2 path
which already has a decision-timeout ack. Give v1 the same behavior:
wait up to 600ms (_v1_ack_wait_seconds) for the decision, and if it's
still not ready, play a short "Секунду." filler via the existing
_emit_early_ack before the real reply, instead of leaving the caller
in silence for the full LLM+TTS round trip. Reuses the same ack
selection/playback code path v2 already exercises, so no new failure
modes - just an added timeout branch mirroring the existing v2 one.
2. ai_orchestrator_service._kb_search: every voice/chat turn re-ran a
full-table scan of kb_articles (all columns, including body text) and
rescored every row in Python, even though the KB rarely changes
mid-conversation. Added an in-process cache keyed by language, gated
on a cheap content fingerprint (row count + max id + max updated_at +
summed title/body/tags length, all computed server-side without
transferring the text columns). A fingerprint mismatch always
triggers a fresh fetch, so this can never serve stale results after
an insert/update/delete - unlike a naive TTL cache, which would have
been be wrong the moment a test (or a real KB edit) changed the table
within the cache window.
Note the first fingerprint design (count + max id + max updated_at
only) was insufficient: utc_now_iso() truncates to whole seconds and
SQLite reuses primary keys after a full-table delete, so two
different row sets written in the same wall-clock second could share
a fingerprint. Caught this via a real test failure
(test_ai_whatsapp_relaxed_kb_search_answers_phrase_query breaking
only when run after test_ai_orchestrator_service.py in the same
process) before it could reach production; the summed content-length
term closes the gap.
Added test_media_runtime_plays_filler_ack_when_v1_decision_is_slow
(asserts greeting -> ack -> reply delivery order when process_turn is
slow) and verified the KB cache against the full
test_ai_orchestrator_service.py + test_ai_whatsapp_orchestrator_service.py
suite plus a wider kb/orchestrator/whatsapp/telegram/voice-filtered run:
only the same pre-existing, already-documented failures remain (unrelated
sales_service test-isolation ordering, one known persona-prompt
assertion) - no new failures from either change.
Streaming the LLM decision itself (start speaking reply_text before the
full structured JSON response finishes generating) was scoped but
deliberately deferred: it needs incremental JSON parsing on top of SSE
streaming to detect when just the reply_text field is complete, shared
across both voice and text-channel decision paths - a separate,
higher-risk change that deserves its own PR and testing pass rather than
being bundled here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- docs: longread.md — deep architecture/flow review of the whole
platform (services, event bus reality vs docs, AI/ML stack honesty
check, tech debt inventory)
- ai_orchestrator_service/voice.py: drop _voice_decision_legacy
(unreferenced) and the shadowed first _voice_decision definition
(silently overwritten by the real one, dead code)
- ui/analyst/app.js: drop duplicate dead definitions of
loadSavedAnalyticsViews/saveAnalyticsView/deleteAnalyticsView and
the first loadAnalyticsTrend implementation, all shadowed by later
declarations in the same file; kept the intentional AI-mode
drilldown wrapper layer (openAnalyticsDrilldown/exportAnalyticsDrilldownCsv/etc.)
since that duplication is deliberate delegation, not dead code
- ui/operator/vendor/sip-0.21.2.min.js: remove byte-identical orphaned
duplicate of ui/operator/sip-0.21.2.min.js (unreferenced anywhere)
Verified via full pytest run: identical set of 97 pre-existing
failures before and after (sales_* test-isolation ordering issue and
one known persona-prompt test), no new regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Remove duplicate function definitions with hardcoded "AI-оператор" strings
(ai_voice_runtime, ai_orchestrator, voice_name_config, voice.py)
- Remove unreachable dead code after return in ai_voice_runtime
- Add SQL LIMIT to 17 unbounded queries across 12 services to prevent OOM
- Move Python-side filtering to SQL WHERE in reporting_service
- Downgrade 19 logger.warning to logger.info for normal-flow events in media_runtime
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When a new voice session starts and the customer already has a trusted
display_name on file (e.g. from a previous call), automatically set
name_status to 'name_obtained' so the AI greets them by name instead
of asking again.
- Block old _voice_downstream_name_update from overwriting name_obtained status
unless user explicitly says 'меня зовут X' (explicit_candidate only)
- Move _persist_voice_name_state to run AFTER LLM extraction, not before,
so LLM always gets a chance to override the regex result
- Fixes bug where phrases like 'можешь назвать их' were incorrectly
captured as customer name, overwriting the real name