Feature Enhancement Request: Customer-Controlled Azure Text-to-Speech (TTS) Voice Selection via SSML in RingCX Workflows
Feature Enhancement Request: Customer-Controlled Azure Text-to-Speech (TTS) Voice Selection via SSML in RingCX Workflows
Title:
Enable SSML Support and Customer-Controlled Azure Text-to-Speech Voice Selection for IVR Workflows
Product:
RingCX Workflow / IVR Text-to-Speech (Azure TTS)
Business Justification:
Following the migration from Poly TTS to Azure Text-to-Speech, customers have lost the ability to control which voice is used for IVR prompts. The platform currently assigns a default Azure voice with no option for customers to select a preferred voice or configure speech characteristics.
This has resulted in:
• Unexpected voice changes without prior customer notification.
• Customer dissatisfaction due to inconsistent IVR branding and user experience.
• Reduced flexibility for multilingual deployments and accessibility requirements.
• Lack of documentation explaining the change, supported voices, and current platform limitations.
Several enterprise customers have requested the ability to choose Azure voices directly, particularly through Speech Synthesis Markup Language (SSML), allowing customization of voice, language, speaking style, pronunciation, and other speech attributes.
Current Limitation
• Workflow Text-to-Speech uses a fixed Azure voice.
• Customers cannot specify:
Azure voice name
Language or locale
Neural voice
Speaking style
Prosody (rate, pitch, volume)
• No SSML input is supported.
• No fallback or notification mechanism exists when the Azure TTS backend experiences failures.
• Documentation does not clearly describe supported voices or service limitations.
Requested Enhancement
• Implement SSML support within the Workflow Text-to-Speech function to allow customers to fully leverage Azure Speech capabilities while maintaining compatibility with the existing Azure backend.
The enhancement should allow customers to:
- Enter SSML directly within the TTS node.
- Select supported Azure voices.
- Configure: • Voice • Language • Speaking style • Rate • Pitch • Volume • Pronunciation
- Support multilingual voice selection where available.
- Validate SSML before execution.
- Maintain backward compatibility for customers who continue using plain text.
Additional Requested Improvements
- Supported Voice Documentation
Provide documentation that includes:
• List of supported Azure voices
• Supported locales
• Neural voice availability
• SSML examples
• Feature limitations
• Usage guidelines
2. Advanced Voice Clarification
Clarify whether premium Azure/OpenAI voices (for example, Nova Multilingual Neural HD) will be supported and whether their usage incurs additional licensing or operational costs.
- Backend Monitoring and Failure Handling
Improve resiliency of the TTS service by providing:
• Detection of Azure TTS service failures.
• Workflow error handling when synthesis fails.
• Optional fallback logic.
• Monitoring and alerting for backend outages.
• Customer-facing notifications for service-impacting events.
- Customer Change Communication
Establish a standardized process for notifying customers whenever backend speech engines or default voices change to avoid unexpected IVR behavior and preserve customer experience.
Expected Benefits
• Restores customer control over IVR voice experience.
• Enables richer, more natural speech using Azure SSML.
• Improves multilingual support.
• Reduces customer complaints following backend changes.
• Improves documentation and product transparency.
• Enhances reliability through better monitoring and failure handling.
Aligns RingCX with Azure Speech capabilities and customer expectations.
Priority
High
This enhancement addresses a customer-facing functionality gap introduced during the migration from Poly TTS to Azure TTS and impacts IVR branding, accessibility, multilingual deployments, and overall customer experience.
Customer Impact
• Customers currently cannot customize Azure voices.
• Existing IVRs may sound different after backend changes without customer awareness.
• Organizations relying on branded voice experiences cannot maintain consistency.
• Lack of failure visibility can result in silent IVR issues during Azure TTS outages.
Proposed Next Steps
• Design and implement SSML support within the Workflow TTS node.
• Confirm the list of Azure/OpenAI voices that will be supported.
• Determine any licensing or pricing implications for premium voices.
• Publish comprehensive product documentation.
• Implement monitoring, fallback handling, and customer notification mechanisms.
• Communicate the implementation timeline to affected customers.
This enhancement was discussed jointly with Engineering and Support, with agreement to pursue SSML support as the long-term solution for restoring customer control over Azure Text-to-Speech voice selection.