Observed Signal · May 31, 2026 · Technical Implementation · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Inworld TTS Ignores Paralinguistic Tags — Working Alternatives
A HoneyChat Engineering case study found that inline paralinguistic tags (e.g., [laugh], [sigh]) are not supported by Inworld TTS and either produce silence or are spoken literally. HoneyChat integrated Inworld TTS-1.5 Max into its Telegram voice-first companion and tested many tag syntaxes across 26 archetypes and 15 languages. Instead of meta-tags, reliable ways to alter prosody and expressivity were: asterisks for emphasis, ellipses for pause-with-mood, SSML <break> for hard pauses, and onomatopoeia (e.g., "ha-ha", "mmm") for breath/laughter. Inworld exposes request parameters such as temperature and speakingRate plus a subset of SSML. HoneyChat implemented an enrich_for_tts() preprocessor that strips fake tags, inserts SSML breaks, wraps text in <speak>, and adjusts parameters to produce more expressive TTS output. Published 2026-05-31.
Practical implementation guidance for using Inworld TTS affects developers building voice-first conversational apps; clarifies unsupported paralinguistic tags and documents reliable alternatives and API parameters that improve expressive TTS output.
Track Google Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- HoneyChat integrated Inworld TTS-1.5 Max as its primary TTS engine; Inworld TTS-1.5 Max is listed at $10 per 1M characters, 1259 ELO on TTS Arena, and supports 15 languages.
- Inline paralinguistic tags like [laugh] and [sigh] are not supported by the Inworld API — they produced silence or the literal token in audio.
- Inworld TTS request parameters include temperature and speakingRate, and the API accepts a subset of SSML including the <break> tag for pauses.
- Four patterns reliably changed output: asterisks for emphasis, ellipsis for prosodic pauses, SSML <break> for hard pauses, and onomatopoeia (e.g., 'ha-ha', 'mmm') for laughs/breaths.
- HoneyChat uses a preprocessor function enrich_for_tts() to strip fake tags, convert ellipses to <break>, wrap text in <speak>, and adjust TTS parameters before sending requests.
Connected Companies & Entities
1 Entity mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Hermes Voice Control From Your Phone
A developer guide explains how to add two‑way voice control to the Hermes self‑hosted agent using a three‑stage pipeline: STT (speech-to-text), reasoning (Hermes handles requests like typed input), and TTS (text-to-speech). The post recommends a free local stack—faster‑whisper for on-device transcription and Edge TTS for spoken output—and enumerates cloud STT/TTS alternatives (Groq, OpenAI, Mistral, ElevenLabs, MiniMax, NeuTTS). It covers configuration snippets, platform workflows (Telegram, Discord, Signal, WhatsApp), required dependencies (ffmpeg + OGG/Opus for Telegram voice bubbles), mobile permissions, reliable speaking patterns, troubleshooting tips, and a quick-start recap including pip install and gateway startup commands. The guide emphasizes starting with local models, upgrading STT first if accuracy/latency require it, and tuning TTS later for quality.
ElevenLabs TTS: SDK Is Simple, Billing and Voices Aren't
A developer who runs a production pipeline for short-form video describes practical pitfalls when integrating the ElevenLabs text-to-speech API. While the SDK quickstart is minimal and easy to use, three operational issues caused production problems: library voice IDs can be removed (they are not stable), the API provides separate streaming and batch methods (streaming plus low-latency models are required for interactive agents), and ElevenLabs bills per character/credit for every generation including regenerations. The author recommends pinning or cloning voices to obtain stable IDs, using the streaming endpoint with low-latency models for agents, aggressively caching outputs keyed by (text, voice_id, model_id, settings), and proofing text before generation to avoid repeat charges. Overall the API is judged high-quality for narration, dubbing and voice agents but engineering for billing and voice stability is essential for production use.
Instagram adds voice filters to DMs
Instagram is testing AI-powered voice effects for Direct Messages, enabling users to make voice notes sound like characters such as a chipmunk, alien, demon, robot, underwater, or stadium announcer. In 1:1 and group chats, users record voice messages and apply effects via an in-app editor; recipients can see the effect used and may apply it in their replies. The voice remains bassed on the original tone and emotion while the vocal timbre changes. The process is designed to be simple: open a chat, record, choose an effect, preview, and send. The feature aligns with Meta’s ongoing expansion of messaging capabilities and follows other recent updates in Meta’s ecosystem, though it’s not yet clear which users or markets have access. The article notes that the deployment details are still uncertain.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
