VoiceFlux
Overview
VoiceFlux converts a persisted AI Employee text reply into contextual synthesized audio for WhatsApp, Baileys, Telegram, Messenger, and Instagram. The original text remains the recovery response whenever synthesis, allowance, storage, or delivery cannot complete safely. VoiceFlux requires a Pro plan or higher, an active VoiceFlux add-on, organization activation, AI Employee activation, and conversation activation. Usage is measured in generated audio seconds. The implementation follows the validated architecture and evidence indocs/voiceflux/01-current-audio-pipeline-audit.md through docs/voiceflux/12-orchestrator-implementation-prompt.md.
Activation Controls
All three controls must be enabled:
The conversation control is unavailable when AI is disabled, the assigned AI Employee has voice replies disabled, the channel is unsupported, or the commercial entitlement is inactive.
Dedicated VoiceFlux Page
Open VoiceFlux from the main navigation. Its button is directly below Livia and before ZappFlux in both expanded and collapsed navigation. The dedicated page brings the operational and commercial controls together:- current plan/add-on status and a direct Billing action
- included, consumed, and remaining audio allowance
- reusable voice characters for all or selected AI Employees
- voice, locale, tone, intensity, pace, and formality
- organization activation, default character, safe text fallback, audio retention, and retention period
agents.read can view the page. Profile changes require agents.write, while organization policy changes require billing.manage.
Voice Profiles
Use the dedicated VoiceFlux page to manage reusable organization and AI Employee characters, or open an AI Employee’s settings to configure its assigned VoiceFlux profile:- voice and locale
- tone and emotional intensity
- pace and formality
Endpoint Reference
Payloads
Update organization settings:Examples
Model and Reliability Policy
gemini-3.1-flash-tts-preview is the stable, sole operational and commercial VoiceFlux model. A job can call it at most twice. Only after both calls fail may the backend make one final recovery call to gemini-2.5-flash-tts. There is no model selector in the UI, tenant settings, or environment variables, and recovery-model pricing is not used as the commercial baseline.
VoiceFlux reaches Google directly through the official Gemini API. The server uses its existing GOOGLE_AI_API_KEY; it does not require Google Cloud project/location settings and does not route synthesis through OpenRouter.
Best Practices
- Enable VoiceFlux first at organization level, then for the AI Employee, then for each intended conversation.
- Keep profiles concise and use structured style options rather than embedding delivery instructions in AI text.
- Review remaining audio seconds on Billing and recent job diagnostics in Conversation Logs.
- Use the shortest retention period compatible with operational requirements.

