Author SHA1 Message Date
arystanbek 9d918af12a fix: bring aimaq host-override files (voice.py, persona env) back into git
Both services/ai_orchestrator_service/voice.py and the aimaq persona/DOMAIN
SCOPE prompt were bind-mounted straight from the host on the aimaq stack,
bypassing git and CI entirely since they were first hand-edited in prod.

voice.py: merged the host's live business logic (gas/aimaq domain keyword
list, off-domain Kazakh/Russian replies, disabled re-correction of an already
obtained name) with the timeout_seconds fix from 09bcf74 that never reached
aimaq because the bind mount blocked it.

aimaq.env.production: replaced the AI_OPERATOR_* env values (which were never
interpolated -- {agent_name}/{company_name} would have been read literally)
with the final resolved Zhanna/Kazakgaz Aimaq text including the DOMAIN SCOPE
clause, matching what was actually live on the host.
2026-08-30 07:45:35 +00:00
arystanbek ff85fa27fb Merge pull request 'feat: aimaq voice AI persona Zhanna + language-choice greeting' (#7) from feature/aimaq-zhanna-persona-language-choice into main
deploy / deploy (push) Successful in 30s
2026-08-30 07:22:16 +00:00
arys 82c89eff5a feat: give aimaq voice AI a Qazaqgaz Aimaq persona (Zhanna) and language-choice greeting
Parametrize ai_operator_default_config() via AI_OPERATOR_* env vars,
falling back to the existing hardcoded defaults so the other two stacks
(call-center, sales-call-center) that share this code are unaffected.

Set aimaq-only overrides matching the client-provided script (Скрипт
ии-оператора КГА.docx): agent renamed to Жанна, company to Казакгаз
Аймак, the voice greeting now asks the caller whether Russian or
Kazakh is more convenient before anything else, and the base system
prompt instructs the model to commit to whichever language the caller
picks for the rest of the call, ask how to address them, and confirm
there's nothing else before saying goodbye.
2026-08-30 12:22:00 +05:00
didar 36c2acb011 test: add test for emotive ack rotation not gated on v2 queue eligibility
deploy / deploy (push) Successful in 31s
2026-08-30 12:03:37 +05:00
didar fe9b3f3a80 feat: enhance voice reply logic to prevent duplicate name addressing and improve greeting handling
deploy / deploy (push) Successful in 32s
2026-08-30 11:53:56 +05:00
arystanbek 0a829d26c7 Merge pull request 'fix: TTS ack-bank stale voice cache + website/English pronunciation' (#6) from fix/voice-tts-ack-cache-and-pronunciation into main
deploy / deploy (push) Successful in 30s
2026-08-30 06:39:32 +00:00
arys fe8598e09f fix: TTS ack-bank served stale voice after voice-config changes; add website/English pronunciation rules
RuntimeConfiguredTTSProvider (the live 'dynamic' TTS provider, resolves
voice from DB-backed config) never overrode cache_fingerprint(), so it
fell back to the base class's empty string. PrebakedAckBank keys its
on-disk cache on that fingerprint, so short filler phrases like
'Секунду' kept serving audio baked with the previous ElevenLabs voice
even after a voice change, while full LLM replies (cached inside the
resolved provider itself, keyed on its own voice id) already used the
new voice — explaining why callers heard two different voices in the
same call. Fix: delegate cache_fingerprint() to the resolved provider.

Also extend the voice delivery_hint so the model transliterates website
addresses and English words/abbreviations into spoken Cyrillic instead
of leaving raw Latin text for the TTS engine to mangle (egov.kz was
coming out as 'эговкз').
2026-08-30 11:39:14 +05:00
didar 9e4c47eddd feat: enhance AudioSocketMediaRuntime to skip filler acks for closing intents and throttle repeated filler acks
deploy / deploy (push) Successful in 30s
2026-08-30 11:30:55 +05:00
arystanbek a013ac95e8 Merge pull request 'fix: voice AI spells out numbers/dates for TTS' (#5) from fix/voice-number-pronunciation into main
deploy / deploy (push) Successful in 31s
2026-08-30 06:22:05 +00:00
10 changed files with 591 additions and 38 deletions
+7
View File
@@ -185,3 +185,10 @@ EVENT_BUS_EXCHANGE=mvpcc.domain.events
ERROR_TELEGRAM_ALERTS=0
ERROR_TELEGRAM_ALERT_BOT_TOKEN=8734026216:AAEK-xdgVR5al6rfKp_DIb-HSJZXGUG2Tvg
ERROR_TELEGRAM_ALERT_CHANNEL_ID=t.me/konturaitelecom
AI_OPERATOR_AGENT_NAME=Жанна
AI_OPERATOR_COMPANY_NAME=Казакгаз Аймак
AI_OPERATOR_VOICE_GREETING_RU=Добрый день! Меня зовут Жанна, я ИИ-оператор контакт-центра Казакгаз Аймак. Вам удобнее получить консультацию на казахском или на русском языке?
AI_OPERATOR_VOICE_GREETING_KZ=Қайырлы күн! Менің атым Жанна, мен «Қазақгаз Аймақ» байланыс орталығында жасанды интеллект операторымын. Кеңесті қазақ немесе орыс тілінде алғыңыз келе ме?
AI_OPERATOR_IDENTITY_REPLY_RU=Я Жанна, ИИ-оператор контакт-центра Казакгаз Аймак. Чем могу помочь?
AI_OPERATOR_IDENTITY_REPLY_KZ=Мен Жаннамын, «Қазақгаз Аймақ» байланыс орталығының жасанды интеллект операторымын. Қалай көмектесе аламын?
AI_OPERATOR_BASE_SYSTEM_PROMPT=Ты Жанна, единый ИИ-оператор контакт-центра Казакгаз Аймак для звонков, Telegram и других каналов. Всегда сохраняй одну и ту же личность: тебя зовут Жанна. Если клиент спрашивает, кто ты или как тебя зовут, отвечай, что ты Жанна. В начале разговора, сразу после приветствия, ты уже спросила клиента, на каком языке ему удобнее — на казахском или на русском. Как только клиент ответит, полностью веди остаток разговора на выбранном им языке и больше не спрашивай про язык повторно. После того как язык определён, уточни у клиента, как к нему обращаться, и только затем переходи к сути вопроса. Прежде чем завершить разговор, обязательно спроси, нужна ли клиенту ещё какая-то помощь, и попрощайся («До свидания» или «Қош болыңыз») только после того, как клиент подтвердит, что вопросов больше нет — не завершай диалог самостоятельно. Отвечай естественно, кратко и по делу. Когда говоришь о себе, используй женский род: могла, смогла, сделала, готова, проверила, нашла. Телефонные номера читай по цифрам. Не используй Markdown, URL или таблицы. DOMAIN SCOPE: ты отвечаешь ТОЛЬКО на вопросы, связанные с услугами газоснабжения: оплата газа, счётчики, ваучеры, технические условия, отключение газа, мобильное приложение, квитанции, передача показаний, контакты филиалов, безопасность газа. Если вопрос НЕ связан с газом (погода, спорт, политика, другие компании, личные советы, кулинария, рецепты, анекдоты, новости, фильмы), вежливо откажи БЕЗ поиска в базе знаний и БЕЗ попытки ответить: RU «Я могу ответить только на вопросы, связанные с услугами газоснабжения. Пожалуйста, задайте вопрос по теме газа.» KK «Мен тек газмен қамтамасыз ету қызметтеріне байланысты сұрақтарға жауап беремін. Газ тақырыбында сұрақ қойыңызшы.» Не пытайся ответить на вопрос не по теме, даже если знаешь ответ.
@@ -143,7 +143,12 @@ def operator_system_prompt(*, language: str, channel_label: str, is_voice: bool,
"Short hotline or service numbers (e.g. 1414, 109) must be spelled out the way people say them as a code, "
"grouped and read naturally (\"1414\" as \"четырнадцать четырнадцать\", not \"тысяча четыреста четырнадцать\"). "
"Calendar dates must use the correct spoken grammatical case (\"25 числа\" as \"двадцать пятого числа\", "
"not \"двадцать пять число\"; \"14 марта\" as \"четырнадцатого марта\")."
"not \"двадцать пять число\"; \"14 марта\" as \"четырнадцатого марта\"). "
"Never leave a website address, domain, or English word/abbreviation in raw Latin script — the TTS engine "
"slurs it into gibberish (e.g. \"egov.kz\" comes out as \"эговкз\"). Transliterate it into how a person "
"actually pronounces it aloud, with an explicit pause word for punctuation: write \"egov.kz\" as "
"\"игов точка кэ-зэт\", write \".kz\"/\".com\" as \"точка кэ-зэт\"/\"точка ком\", spell out an acronym or "
"English word phonetically in Cyrillic (\"IT\" as \"ай-ти\", \"email\" as \"имейл\")."
if is_voice
else "The reply should read like a concise message from a live first-line operator."
)
+114 -20
View File
@@ -603,20 +603,42 @@ def _voice_has_name_correction(text: str | None) -> bool:
return any(marker in normalized for marker in correction_markers)
def _voice_reply_with_name(language: str, reply_text: str, name: str | None) -> str:
def _voice_reply_already_names_customer(reply_text: str, normalized_short: str) -> bool:
tokens = _voice_text_key(reply_text).split(" ")
if normalized_short in tokens:
return True
stem_len = max(len(normalized_short) - 2, 3)
stem = normalized_short[:stem_len]
return any(len(token) >= stem_len and token.startswith(stem) for token in tokens)
def _voice_greeting_word(language: str) -> str:
if str(language or "").strip().lower() == "kz":
return "Сәлеметсіз бе"
return "Здравствуйте"
def _voice_reply_with_name(
language: str,
reply_text: str,
name: str | None,
*,
greet: bool = False,
) -> str:
short_name = _voice_short_name(name)
if not short_name:
return reply_text
prefix = _voice_disclosure_prefix(language)
normalized_short = _voice_text_key(short_name)
lead = f"{_voice_greeting_word(language)}, {short_name}" if greet else short_name
if reply_text.startswith(prefix):
rest = reply_text[len(prefix) :].lstrip()
if normalized_short in _voice_text_key(rest).split(" "):
if _voice_reply_already_names_customer(rest, normalized_short):
return reply_text
return f"{prefix}{short_name}, {rest}"
if normalized_short in _voice_text_key(reply_text).split(" "):
return f"{prefix}{lead}, {rest}"
if _voice_reply_already_names_customer(reply_text, normalized_short):
return reply_text
return f"{short_name}, {reply_text}"
return f"{lead}, {reply_text}"
def _voice_name_metadata(
@@ -746,13 +768,7 @@ def _voice_downstream_name_update(
action = "provide" if not name_value else "confirm"
candidate = candidate or name_value
elif status == "name_obtained":
if (
candidate
and candidate_key
and candidate_key != current_key
and _voice_has_name_correction(transcript_text)
):
action = "correct"
pass
else:
if explicit_candidate:
action = "provide"
@@ -1039,7 +1055,6 @@ def _voice_has_service_topic(text: str | None) -> bool:
"ошибк",
"проблем",
"сбой",
"интернет",
"связь",
"оператор",
"менеджер",
@@ -1047,6 +1062,40 @@ def _voice_has_service_topic(text: str | None) -> bool:
"компан",
"подключ",
"доставк",
"газ",
"счётчик",
"счетчик",
"ваучер",
"отключ",
"квитанц",
"показан",
"приложен",
"безопасн",
"техническ",
"плит",
"котл",
"труб",
"утечк",
"запах",
"абонент",
"договор",
"поверк",
"монтаж",
"счёт",
"счет",
"долг",
"задолжен",
"перерасчёт",
"перерасчет",
"регион",
"аимак",
"aimaq",
"qazaqgaz",
"казахгаз",
"есептегіш",
"төлеу",
"өтінім",
"шарт",
)
return any(marker in normalized for marker in service_markers)
@@ -1080,6 +1129,37 @@ def _voice_is_off_domain_request(text: str | None) -> bool:
"java",
"javascript",
"погод",
"плов",
"приготов",
"кулинар",
"блюдо",
"фильм",
"кино",
"спорт",
"футбол",
"хоккей",
"теннис",
"баскетбол",
"президент",
"политик",
"выборы",
"правительств",
"курс валют",
"криптовалют",
"биткоин",
"песн",
"музык",
"танц",
"шутк",
"загадк",
"сериал",
"книга",
"стихотвор",
"чемпионат",
"ауа райы",
"аспаздық",
"кітап",
"ән айт",
)
if any(marker in normalized for marker in broad_markers):
return True
@@ -1089,7 +1169,14 @@ def _voice_is_off_domain_request(text: str | None) -> bool:
"что такое",
"объясни",
"расскажи про",
"почему",
"посоветуй",
"как приготовить",
"какая погода",
"кто президент",
"кто выиграл",
"какая команда",
"ауа райы қалай",
"әнді айт",
)
return any(normalized.startswith(prefix) for prefix in broad_openers)
@@ -1097,13 +1184,12 @@ def _voice_is_off_domain_request(text: str | None) -> bool:
def _voice_off_domain_reply(language: str) -> tuple[str, str]:
if language == "kz":
return (
"Men kompaniyamyzdyn qyzmetteri men otinishteri boiynsha komek bere alamyn. "
"Eger suraq bizdin qyzmetke qatysty bolsa, qysqasha naqtylaңыз. Qalasaңыз, operatorga qosamyn.",
"AI qongyraudyn taqyrybyn kompaniya qyzmetteri sheginde naqtylaudy usyndy.",
"Кешіріңіз, мен тек газ қызметтері бойынша көмектесемін: төлеу, есептегіш, ваучер. Қалай көмектесе аламын?",
"AI off-topic сұрауды газ қызметтері тақырыбына шектеді.",
)
return (
"Я помогу по вопросам наших услуг и обращений. Если вопрос связан с нашей компанией, скажите коротко, что именно нужно. Если хотите, сразу соединю с оператором.",
"AI мягко вернул разговор к вопросам компании и предложил перевод на оператора.",
"Извините, я помогаю только по вопросам газа: оплата, счётчики, ваучеры. Чем могу помочь?",
"AI отклонил off-topic вопрос и ограничил тему газоснабжением.",
)
@@ -1613,6 +1699,12 @@ def _voice_llm_prompt_messages(
is_voice=True,
config=operator_config,
)
system_prompt += (
" Do not open `reply_text` with a greeting or by addressing the customer by name "
"(e.g. do not write 'Здравствуйте, <имя>' or start with '<имя>,'). The system inserts "
"the customer's name into the spoken reply separately, so naming them yourself would "
"make it get said twice."
)
if name_status in ("name_not_obtained", "name_followup_required"):
system_prompt += (
" The user's name is not yet obtained. If the user explicitly provided their name in this turn, "
@@ -1810,7 +1902,7 @@ def _voice_decision(
"latency_ms": 1,
}
if not kb_results and _voice_is_off_domain_request(normalized):
if _voice_is_off_domain_request(normalized):
reply_text, summary_text = _voice_off_domain_reply(language)
decision = {
"language": language,
@@ -2500,10 +2592,12 @@ def turn_voice_session(session_id: str, payload: VoiceAITurnIn) -> VoiceAITurnDe
decision_metadata.update(decision.get("metadata") or {})
suppress_name_prefix = bool(request_metadata.get("suppress_name_prefix")) if isinstance(request_metadata, dict) else False
if effective_name_status == "name_obtained" and effective_name_value and not early_plan_only and not suppress_name_prefix:
just_learned_name = current_name_status != "name_obtained"
decision["reply_text"] = _voice_reply_with_name(
decision["language"],
decision["reply_text"],
effective_name_value,
greet=just_learned_name,
)
elif inline_name_followup and not decision["needs_handoff"] and not early_plan_only:
inline_followup = _voice_inline_name_followup(decision["language"], config)
+5 -1
View File
@@ -823,7 +823,11 @@ def _media_registration_from_row(row: VoiceAISessionRow, *, queue_code: str | No
voice_v2_duplex=bool(voice_v2_for_session and _voice_v2_duplex_enabled()),
voice_v2_streaming_asr_backend=streaming_backend,
voice_v2_prebaked_ack=bool(voice_v2_for_session and _voice_v2_prebaked_ack_enabled()),
voice_v2_emotive_ack=bool(voice_v2_for_session and _voice_v2_emotive_ack_enabled()),
# Deliberately NOT gated on voice_v2_for_session: phrase-variant rotation only
# needs a text pool + live TTS, not the v2 duplex/partial-ASR pipeline, so calls
# outside the v2 queue allowlist still get varied fillers instead of always the
# single fixed "Секунду." fallback string.
voice_v2_emotive_ack=bool(_voice_v2_emotive_ack_enabled()),
voice_v2_emotive_ack_ru_only=bool(_voice_v2_emotive_ack_ru_only()),
)
@@ -217,6 +217,10 @@ class AudioSocketMediaRuntime:
self._immediate_ack_min_ms = 700
self._v2_ack_post_gap_seconds = 0.10
self._v1_ack_wait_seconds = 0.6
# Skip a would-be filler if the previous one finished too recently, so rapid
# back-and-forth turns (e.g. a caller spelling out a phone number field by
# field) don't get a filler read before every single fragment.
self._ack_min_repeat_gap_seconds = 2.5
self._partial_poll_interval_seconds = 0.20
# Small startup cushion for streamed TTS playback: absorb ElevenLabs
# network delivery jitter before we start pacing frames out to the
@@ -314,10 +318,32 @@ class AudioSocketMediaRuntime:
return True
return normalized in cls._FINAL_LOW_SIGNAL_PHRASES
_CLOSING_INTENT_PHRASES = (
"до свидания",
"всего доброго",
"хорошего дня",
"хорошего вечера",
"прощайте",
"созвонимся",
"это все спасибо",
"это всё спасибо",
"у меня все спасибо",
"у меня всё спасибо",
"больше вопросов нет",
"вопросов больше нет",
"спасибо за помощь",
"спасибо большое до свидания",
"сау болыңыз",
"келесіге дейін",
"рахмет көп",
)
def _detect_early_intent(self, text: str) -> str:
normalized = self._normalize_intent_text(text)
if not normalized:
return "unknown"
if any(token in normalized for token in self._CLOSING_INTENT_PHRASES):
return "closing"
if any(token in normalized for token in ("оператор", "оператором", "человеком", "менеджер", "сотрудник")):
return "operator_request"
if any(token in normalized for token in ("график", "распис", "время работы", "work schedule", "жұмыс")):
@@ -347,6 +373,8 @@ class AudioSocketMediaRuntime:
@staticmethod
def _ack_kind_for_intent(intent: str) -> str:
if intent == "closing":
return "closing"
if intent == "operator_request":
return "handoff"
if intent in {"schedule", "address", "price", "status", "problem"}:
@@ -406,8 +434,18 @@ class AudioSocketMediaRuntime:
return True
return normalized_intent != "unknown"
def _should_emit_blind_ack(self, actor: MediaActor, pcm_bytes: bytes) -> bool:
def _should_emit_blind_ack(self, actor: MediaActor, pcm_bytes: bytes, partial_transcript: str) -> bool:
"""Duration-only fallback for when no usable partial transcript exists yet.
Must defer to the transcript when one *is* available: otherwise a caller
who already said a recognized filler-answer ("да"/"нет"/"хорошо") still
gets a blind ack just because the audio happened to cross the length
threshold, even though `_should_emit_partial_ack` correctly said no.
"""
del actor
transcript_text = str(partial_transcript or "").strip()
if transcript_text and self._is_low_signal_partial_transcript(transcript_text):
return False
return len(pcm_bytes) >= self._immediate_ack_min_bytes
@staticmethod
@@ -1116,10 +1154,22 @@ class AudioSocketMediaRuntime:
metadata: dict[str, Any],
ack_source: str,
ack_kind: str | None = None,
intent: str | None = None,
) -> None:
if actor.closed or actor.early_ack_started:
return
ack_kind = ack_kind or self._ack_kind_for_intent(actor.partial_intent or "unknown")
effective_intent = str(
intent or actor.stable_partial_intent or actor.partial_intent or "unknown"
).strip() or "unknown"
if effective_intent == "closing":
# The caller is wrapping up; a "thinking" filler right before the
# closing reply reads as robotic, so skip it and go straight to the reply.
return
if actor.last_ack_completed_monotonic and (
time.monotonic() - actor.last_ack_completed_monotonic
) < self._ack_min_repeat_gap_seconds:
return
ack_kind = ack_kind or self._ack_kind_for_intent(effective_intent)
ack_text, style_hints, ack_variant = self._select_ack_payload(
actor,
language=language,
@@ -1620,8 +1670,9 @@ class AudioSocketMediaRuntime:
language=actor.registration.language,
metadata=base_metadata,
ack_source="streaming_partial" if actor.asr_streaming_enabled else "precomputed_partial_asr",
intent=partial_intent,
)
elif self._should_emit_blind_ack(actor, pcm_bytes):
elif self._should_emit_blind_ack(actor, pcm_bytes, partial_transcript):
await self._emit_early_ack(
actor,
language=actor.registration.language,
@@ -114,6 +114,10 @@ class RuntimeConfiguredTTSProvider(TTSProvider):
self._provider_cache[cache_key] = provider
return provider
def cache_fingerprint(self, language: str | None, *, style_hints: dict[str, object] | None = None) -> str:
provider = self._provider_for_language(language)
return f"{provider.name}:{provider.cache_fingerprint(language, style_hints=style_hints)}"
def synthesize(self, text: str, *, language: str | None = None, style_hints: dict[str, object] | None = None):
provider = self._provider_for_language(language)
return provider.synthesize(text, language=language, style_hints=style_hints)
+40 -13
View File
@@ -1,6 +1,7 @@
from __future__ import annotations
import json
import os
from typing import Any
from sqlalchemy import select
@@ -13,22 +14,48 @@ from services.shared.sql_models import AIOperatorSettingsRow
AI_OPERATOR_SETTINGS_KEY = "global"
def _env_field(name: str, default: str) -> str:
value = os.getenv(name)
if value is None:
return default
normalized = value.strip()
return normalized or default
def ai_operator_default_config() -> AIOperatorConfig:
agent_name = _env_field("AI_OPERATOR_AGENT_NAME", "Айнур")
company_name = _env_field("AI_OPERATOR_COMPANY_NAME", "DigiOps")
return AIOperatorConfig(
agent_name="Айнур",
company_name="DigiOps",
base_system_prompt=(
"Ты Айнур, единый ИИ-оператор контакт-центра DigiOps для звонков, Telegram и других каналов. "
"Всегда сохраняй одну и ту же личность: тебя зовут Айнур. Если клиент спрашивает, кто ты или как "
"тебя зовут, отвечай, что ты Айнур. Отвечай естественно, кратко и по делу. Когда говоришь о себе, "
"используй женский род: могла, смогла, сделала, готова, проверила, нашла. Не завершай диалог "
"самостоятельно и говори «до свидания» только если клиент явно попрощался или попросил завершить "
"разговор. Телефонные номера читай по цифрам. Не используй Markdown, URL или таблицы."
agent_name=agent_name,
company_name=company_name,
base_system_prompt=_env_field(
"AI_OPERATOR_BASE_SYSTEM_PROMPT",
(
f"Ты {agent_name}, единый ИИ-оператор контакт-центра {company_name} для звонков, Telegram и других "
f"каналов. Всегда сохраняй одну и ту же личность: тебя зовут {agent_name}. Если клиент спрашивает, "
f"кто ты или как тебя зовут, отвечай, что ты {agent_name}. Отвечай естественно, кратко и по делу. "
"Когда говоришь о себе, используй женский род: могла, смогла, сделала, готова, проверила, нашла. "
"Не завершай диалог самостоятельно и говори «до свидания» только если клиент явно попрощался или "
"попросил завершить разговор. Телефонные номера читай по цифрам. Не используй Markdown, URL или "
"таблицы."
),
),
identity_reply_ru=_env_field(
"AI_OPERATOR_IDENTITY_REPLY_RU",
f"Я {agent_name}, оператор контакт-центра {company_name}. Чем могу помочь?",
),
identity_reply_kz=_env_field(
"AI_OPERATOR_IDENTITY_REPLY_KZ",
f"Мен {agent_name}мын, {company_name} байланыс орталығының операторымын. Қалай көмектесе аламын?",
),
voice_greeting_ru=_env_field(
"AI_OPERATOR_VOICE_GREETING_RU",
f"Здравствуйте. Я {agent_name}. Подскажите, пожалуйста, чем помочь. (тест деплоя)",
),
voice_greeting_kz=_env_field(
"AI_OPERATOR_VOICE_GREETING_KZ",
f"Сәлеметсіз бе. Мен {agent_name}мын. Қалай көмектесе аламын?",
),
identity_reply_ru="Я Айнур, оператор контакт-центра DigiOps. Чем могу помочь?",
identity_reply_kz="Мен Айнурмын, DigiOps байланыс орталығының операторымын. Қалай көмектесе аламын?",
voice_greeting_ru="Здравствуйте. Я Айнур. Подскажите, пожалуйста, чем помочь. (тест деплоя)",
voice_greeting_kz="Сәлеметсіз бе. Мен Айнурмын. Қалай көмектесе аламын?",
)
+55
View File
@@ -1681,6 +1681,61 @@ def test_voice_llm_prompt_includes_context_summary_and_uses_12_segments():
assert payload["history"][0]["sequence_no"] == 4
def test_voice_llm_prompt_instructs_model_not_to_self_name_customer():
messages = voice_module._voice_llm_prompt_messages(
language="ru",
customer=None,
interaction=SimpleNamespace(
interaction_id="int_voice_prompt_name",
status="open",
queue_id="que_voice",
subject="schedule",
customer_id="cus_voice_prompt_name",
),
transcript_text="Мне нужен график работы",
transcript_window=[],
conversation_summary_text="",
kb_results=[],
name_value="Ернур",
name_status="name_obtained",
)
system_prompt = messages[0]["content"]
assert "addressing the customer by name" in system_prompt
def test_voice_reply_with_name_does_not_duplicate_inflected_name_form():
# The model may address the customer using a grammatically declined form of
# their name ("Данияре" instead of "Данияр"); an exact-token dedup check
# would miss this and prepend the name a second time.
reply = voice_module._voice_reply_with_name(
"ru", "Здравствуйте, Данияре! Чем могу помочь?", "Данияр"
)
assert reply == "Здравствуйте, Данияре! Чем могу помочь?"
# A reply with no mention of the customer's name still gets it prefixed once.
reply = voice_module._voice_reply_with_name("ru", "Чем могу помочь?", "Данияр")
assert reply == "Данияр, Чем могу помочь?"
def test_voice_reply_with_name_greet_mode_uses_one_of_two_fixed_forms():
# Regular turns (name already known): just the name, never a greeting word.
reply = voice_module._voice_reply_with_name("ru", "Чем могу помочь?", "Данияр", greet=False)
assert reply == "Данияр, Чем могу помочь?"
# The turn the name is first learned: exactly "Здравствуйте, {name}, ...".
reply = voice_module._voice_reply_with_name("ru", "Чем могу помочь?", "Данияр", greet=True)
assert reply == "Здравствуйте, Данияр, Чем могу помочь?"
reply = voice_module._voice_reply_with_name("kz", "Немен көмектесе аламын?", "Ерлан", greet=True)
assert reply == "Сәлеметсіз бе, Ерлан, Немен көмектесе аламын?"
# Still deduplicates even in greet mode if the model already named the customer.
reply = voice_module._voice_reply_with_name(
"ru", "Здравствуйте, Данияре! Чем могу помочь?", "Данияр", greet=True
)
assert reply == "Здравствуйте, Данияре! Чем могу помочь?"
def test_voice_postprocess_reply_uses_summary_context_when_raw_window_lost_topic():
reply_text = voice_module._voice_postprocess_reply_text(
language="ru",
+282
View File
@@ -1249,6 +1249,288 @@ def test_media_runtime_voice_v2_emits_blind_ack_on_first_turn_without_partial_si
assert speak_events[1][0] == "Подскажите подробнее, пожалуйста."
def test_media_runtime_voice_v2_skips_filler_ack_when_caller_says_goodbye():
planned: list[tuple[str, str, str, dict | None]] = []
class _GoodbyeASRProvider(ASRProvider):
name = "goodbye-asr"
def transcribe(self, audio_bytes: bytes, *, language_hint: str | None = None) -> ASRTranscription:
assert audio_bytes
return ASRTranscription(text="Спасибо, до свидания", language=language_hint or "ru", confidence=0.9)
runtime = AudioSocketMediaRuntime(
enabled=True,
host="127.0.0.1",
port=0,
frame_ms=20,
idle_timeout_seconds=2.0,
registration_wait_timeout_seconds=0.5,
min_speech_ms=40,
trailing_silence_ms=40,
max_turn_ms=400,
asr_provider=_GoodbyeASRProvider(),
tts_provider=_StubTTSProvider(),
load_registration_by_media_uuid=lambda value: None,
mark_media_connected=lambda session_id, value: None,
mark_media_ended=lambda session_id, reason: None,
touch_media_frame=lambda session_id: None,
set_state=lambda session_id, state, handoff_reason, metadata: None,
get_pending_greeting=lambda session_id: None,
mark_reply_delivered=lambda session_id, text, is_greeting: None,
plan_reply=lambda session_id, text, metadata, kind: planned.append((session_id, text, kind, metadata)),
process_turn=lambda session_id, transcript_text, language, barge_in, metadata: (
time.sleep(0.25)
or VoiceAITurnDecisionOut(
language=language or "ru",
intent="closing",
reply_text="Хорошо, всего доброго!",
confidence=0.9,
needs_handoff=False,
handoff_reason=None,
case_action="close",
kb_refs=[],
summary_text="call wrapped up",
model="stub-voice",
latency_ms=1,
status="active",
)
),
request_handoff=lambda session_id, customer_request_text, decision: None,
handle_media_error=lambda session_id, message, metadata: None,
)
async def _fake_speak_text(
current_actor,
text: str,
*,
is_greeting: bool,
style_hints: dict[str, object] | None = None,
) -> None:
del current_actor, is_greeting, style_hints
await asyncio.sleep(0)
runtime._speak_text = _fake_speak_text # type: ignore[method-assign]
pcm_frame = (1000).to_bytes(2, "little", signed=True) * 160
async def _scenario() -> None:
actor = MediaActor(
registration=MediaRegistration(
voice_session_id="avs_media_runtime_v2_goodbye",
call_id="call_media_runtime_v2_goodbye",
interaction_id="int_media_runtime_v2_goodbye",
ai_session_id="ais_media_runtime_v2_goodbye",
language="ru",
media_uuid=str(uuid.uuid4()),
queue_code="voice_lab_ai",
queue_id="que_voice_lab_ai",
agent_profile="voice_support",
voice_v2_enabled=True,
voice_v2_ack_mode="immediate_short",
voice_v2_streaming_tts=True,
voice_v2_partial_asr=False,
),
reader=asyncio.StreamReader(),
writer=None, # type: ignore[arg-type]
vad=EnergyVAD(frame_ms=20, min_speech_ms=40, trailing_silence_ms=40, max_turn_ms=400),
frame_ms=20,
frame_bytes=320,
)
actor.finalized_caller_turn_count = 1
await runtime._process_utterance(actor, pcm_frame, False)
asyncio.run(_scenario())
assert [item[2] for item in planned] == ["reply"]
assert planned[0][1] == "Хорошо, всего доброго!"
def test_media_runtime_voice_v2_throttles_repeated_filler_ack_within_gap():
planned: list[tuple[str, str, str, dict | None]] = []
class _SlowASRProvider(ASRProvider):
name = "slow-asr"
def transcribe(self, audio_bytes: bytes, *, language_hint: str | None = None) -> ASRTranscription:
assert audio_bytes
return ASRTranscription(text="Хочу узнать график работы", language=language_hint or "ru", confidence=0.9)
runtime = AudioSocketMediaRuntime(
enabled=True,
host="127.0.0.1",
port=0,
frame_ms=20,
idle_timeout_seconds=2.0,
registration_wait_timeout_seconds=0.5,
min_speech_ms=40,
trailing_silence_ms=40,
max_turn_ms=400,
asr_provider=_SlowASRProvider(),
tts_provider=_StubTTSProvider(),
load_registration_by_media_uuid=lambda value: None,
mark_media_connected=lambda session_id, value: None,
mark_media_ended=lambda session_id, reason: None,
touch_media_frame=lambda session_id: None,
set_state=lambda session_id, state, handoff_reason, metadata: None,
get_pending_greeting=lambda session_id: None,
mark_reply_delivered=lambda session_id, text, is_greeting: None,
plan_reply=lambda session_id, text, metadata, kind: planned.append((session_id, text, kind, metadata)),
process_turn=lambda session_id, transcript_text, language, barge_in, metadata: (
time.sleep(0.25)
or VoiceAITurnDecisionOut(
language=language or "ru",
intent="clarification",
reply_text="Подскажите, пожалуйста, какой город вас интересует?",
confidence=0.9,
needs_handoff=False,
handoff_reason=None,
case_action="keep_open",
kb_refs=[],
summary_text="reply ready",
model="stub-voice",
latency_ms=1,
status="active",
)
),
request_handoff=lambda session_id, customer_request_text, decision: None,
handle_media_error=lambda session_id, message, metadata: None,
)
async def _fake_speak_text(
current_actor,
text: str,
*,
is_greeting: bool,
style_hints: dict[str, object] | None = None,
) -> None:
del current_actor, is_greeting, style_hints
await asyncio.sleep(0)
runtime._speak_text = _fake_speak_text # type: ignore[method-assign]
pcm_frame = (1000).to_bytes(2, "little", signed=True) * 160
async def _scenario() -> None:
actor = MediaActor(
registration=MediaRegistration(
voice_session_id="avs_media_runtime_v2_throttle",
call_id="call_media_runtime_v2_throttle",
interaction_id="int_media_runtime_v2_throttle",
ai_session_id="ais_media_runtime_v2_throttle",
language="ru",
media_uuid=str(uuid.uuid4()),
queue_code="voice_lab_ai",
queue_id="que_voice_lab_ai",
agent_profile="voice_support",
voice_v2_enabled=True,
voice_v2_ack_mode="immediate_short",
voice_v2_streaming_tts=True,
voice_v2_partial_asr=False,
),
reader=asyncio.StreamReader(),
writer=None, # type: ignore[arg-type]
vad=EnergyVAD(frame_ms=20, min_speech_ms=40, trailing_silence_ms=40, max_turn_ms=400),
frame_ms=20,
frame_bytes=320,
)
actor.finalized_caller_turn_count = 1
await runtime._process_utterance(actor, pcm_frame, False)
runtime._reset_live_turn_state(actor)
await runtime._process_utterance(actor, pcm_frame, False)
asyncio.run(_scenario())
assert [item[2] for item in planned] == ["ack", "reply", "reply"]
def test_media_runtime_voice_v2_blind_ack_defers_to_known_low_signal_partial_transcript():
speak_events: list[str] = []
runtime = AudioSocketMediaRuntime(
enabled=True,
host="127.0.0.1",
port=0,
frame_ms=20,
idle_timeout_seconds=2.0,
registration_wait_timeout_seconds=0.5,
min_speech_ms=40,
trailing_silence_ms=40,
max_turn_ms=2000,
asr_provider=_StubASRProvider(),
tts_provider=_StubTTSProvider(),
load_registration_by_media_uuid=lambda value: None,
mark_media_connected=lambda session_id, value: None,
mark_media_ended=lambda session_id, reason: None,
touch_media_frame=lambda session_id: None,
set_state=lambda session_id, state, handoff_reason, metadata: None,
get_pending_greeting=lambda session_id: None,
mark_reply_delivered=lambda session_id, text, is_greeting: None,
plan_reply=lambda session_id, text, metadata, kind: None,
process_turn=lambda session_id, transcript_text, language, barge_in, metadata: VoiceAITurnDecisionOut(
language=language or "ru",
intent="clarification",
reply_text="Подскажите подробнее, пожалуйста.",
confidence=0.9,
needs_handoff=False,
handoff_reason=None,
case_action="keep_open",
kb_refs=[],
summary_text="reply ready",
model="stub-voice",
latency_ms=1,
status="active",
),
request_handoff=lambda session_id, customer_request_text, decision: None,
handle_media_error=lambda session_id, message, metadata: None,
)
async def _fake_speak_text(
current_actor,
text: str,
*,
is_greeting: bool,
style_hints: dict[str, object] | None = None,
) -> None:
del current_actor, is_greeting, style_hints
speak_events.append(text)
await asyncio.sleep(0)
runtime._speak_text = _fake_speak_text # type: ignore[method-assign]
pcm_frame = (1000).to_bytes(2, "little", signed=True) * 160
async def _scenario() -> None:
actor = MediaActor(
registration=MediaRegistration(
voice_session_id="avs_media_runtime_v2_low_signal_blind",
call_id="call_media_runtime_v2_low_signal_blind",
interaction_id="int_media_runtime_v2_low_signal_blind",
ai_session_id="ais_media_runtime_v2_low_signal_blind",
language="ru",
media_uuid=str(uuid.uuid4()),
queue_code="voice_lab_ai",
queue_id="que_voice_lab_ai",
agent_profile="voice_support",
voice_v2_enabled=True,
voice_v2_ack_mode="immediate_short",
voice_v2_streaming_tts=True,
voice_v2_partial_asr=True,
),
reader=asyncio.StreamReader(),
writer=None, # type: ignore[arg-type]
vad=EnergyVAD(frame_ms=20, min_speech_ms=40, trailing_silence_ms=40, max_turn_ms=2000),
frame_ms=20,
frame_bytes=320,
)
actor.finalized_caller_turn_count = 1
# A caller who already said a recognized filler-answer ("да") should not
# get a blind ack just because the audio clip crossed the length threshold.
actor.stable_partial_transcript = "да"
await runtime._process_utterance(actor, pcm_frame * 40, False)
asyncio.run(_scenario())
assert speak_events == ["Подскажите подробнее, пожалуйста."]
def test_media_runtime_low_signal_filter_catches_short_asr_noise():
assert AudioSocketMediaRuntime._is_low_signal_partial_transcript("Давай")
assert AudioSocketMediaRuntime._is_low_signal_partial_transcript("твой")
+24
View File
@@ -855,3 +855,27 @@ def test_push_voice_ai_telephony_event_call_ended_returns_detached_safe_payload(
}
assert closed == [("avs_runtime_call_ended", "call_ended")]
assert orchestrator_calls == [("POST", "/ai/voice/sessions/avs_runtime_call_ended/close")]
def test_media_registration_emotive_ack_rotation_is_not_gated_on_v2_queue_eligibility(monkeypatch):
# AI_VOICE_V2_QUEUE_CODES defaults to "voice_lab_ai" only, so a queue outside
# that allowlist runs the plain v1 ack path. Emotive-ack rotation must still
# apply there — otherwise every filler collapses to the single fixed
# "Секунду." fallback string instead of rotating through phrase variants.
monkeypatch.delenv("AI_VOICE_V2_QUEUE_CODES", raising=False)
monkeypatch.delenv("AI_VOICE_V2_EMOTIVE_ACK_ENABLED", raising=False)
row = VoiceAISessionRow(
session_id="avs_runtime_emotive_ack_v1",
call_id="call_runtime_emotive_ack_v1",
interaction_id="int_runtime_emotive_ack_v1",
ai_session_id="ais_runtime_emotive_ack_v1",
language="ru",
media_uuid="media-emotive-ack-v1",
queue_id="que_not_v2_eligible",
agent_profile="voice_support",
)
registration = runtime_module._media_registration_from_row(row, queue_code="queue_outside_v2_allowlist")
assert registration.voice_v2_enabled is False
assert registration.voice_v2_emotive_ack is True