Skip to content

#Voice

Oct 8

TodayOct 8Thu10 items
  1. Xiaomi MiMo63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    Xiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    Why it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  2. SiliconANGLE · AI38

    Automation Anywhere to acquire Boost.ai to expand customer-facing voice AI

    Automation Anywhere Inc. announced an agreement to acquire Boost.ai Inc., a conversational voice AI company, from Nordic Capital, to extend its autonomous enterprise platform into customer experience. Boost.ai supports more than 36 languages, serves hundreds of customers in regulated industries and Europe, and maintains more than 650 deployments and about 600 live AI agents. The deal follows Automation Anywhere's late 2025 acquisition of Aisera Inc.

  3. Sierra Blog62

    Sierra launches fleming-1 to detect AI agents calling by phone

    Sierra has launched fleming-1, a model that analyzes caller speech in real time and scores audio for signs it was generated by AI. It flags likely AI callers while keeping real people unflagged by default, and companies decide how to handle those calls. The model works with any voice agent built on Sierra, and Sierra also announced Personal Agent Protocol, an open standard for authorized agent-to-business interactions.

    Why it matters: The post explains why companies need to know when a caller is an AI agent, which frames the detection model as a business decision rather than an automatic block.

  4. Elvis Saravia34

    WER dropped from 85% to 7.65% on Bengali with one fine-tune. Numbers like that are why 80 labs asked to license Monsoon in a week. What I respect most is the rule behind it: @voicearena_ai only builds a dataset if it moves the needle. Most data vendors can't say that.

    WER dropped from 85% to 7.65% on Bengali with one fine-tune. Numbers like that are why 80 labs asked to license Monsoon in a week. What I respect most is the rule behind it: @voicearena_ai only builds a dataset if it moves the needle. Most data vendors can't say that.

  5. SiliconANGLE · AI22

    Willow picks CoreWeave for AI model training and forward-deployed support

    Willow Care Inc., maker of the AI dictation app Willow Voice, chose CoreWeave for its forward-deployed support rather than compute alone, according to co-founder and CTO Lawrence Liu. Liu said CoreWeave's reinforcement learning infrastructure lets Willow focus on eval alignment, while Willow fine-tunes its own speech recognition model and pairs it with a compact post-processing LLM. He said inference demand is growing faster than training as dictation use climbs.

  6. Elvis Saravia32

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

    I’ve never seen anything like this. AI voice technology is getting out of hand. Drama 3 is the most control a model has shown over the direction of language. It can change tone in the same sentence and has vocal control similar to human speech. I tried testing this feature and couldn’t believe how well it worked. Do yourself a favor and try @FishAudio out.

  7. The Verge · AI52

    Google's experimental AI Edge Foresight transcribes meetings fully offline on Mac

    Google has released AI Edge Foresight, a free experimental note-taking app that transcribes meetings and audio files entirely offline on macOS. It runs on the on-device EmbeddingGemma 2 model and turns shorthand notes into polished notes based on the transcript. Google says files, meeting audio, and notes never leave the computer, and the app is currently optimized only for Macs with Apple Silicon.

  8. ElevenLabs Blog39

    ElevenReader Launches in Brazil With Fábio Porchat Narration and 70,000 Portuguese Books

    ElevenReader, ElevenLabs' consumer audio platform, launches in Brazil with narration by actor and comedian Fábio Porchat. Through a partnership with Bookwire Brasil, the app offers 70,000 licensed Brazilian Portuguese titles, most of which have no audio edition. The app is free on iOS and Android, with the full catalog available through ElevenReader Ultra.

  9. ElevenLabs Blog33

    ElevenLabs opens Singapore office and launches local data residency for Asia-Pacific

    ElevenLabs is officially launching in Singapore, opening an office to establish its Southeast Asia hub. In July it launched Singapore data residency, offering enterprises local hosting, zero-retention options, lower latency across Singapore and East Asia, and SOC 2 compliance. Customers named include Funding Societies, Atome, and Rezonate.

Oct 7

Oct 7Wed
  1. fal34

    Vidu Q4 is now live on fal - Image-to-video from a single first frame, with native audio - Reference-to-video with up to 12 reference images and 3 voice clips for consistent characters and voices - 3 to 16 second clips from 540p up to 4K

    Vidu Q4 is now live on fal - Image-to-video from a single first frame, with native audio - Reference-to-video with up to 12 reference images and 3 voice clips for consistent characters and voices - 3 to 16 second clips from 540p up to 4K

  2. Liquid AI36

    d1-omni-600M is an experimental 600M-parameter model for text + image or text + audio. It combines LFM2.5-Encoder-350M with vision and audio encoders, and leads our text benchmark comparison in toxicity detection and paraphrase identification. Use it for voice-command routing, on-device moderation, and intent classification. 3/

    d1-omni-600M is an experimental 600M-parameter model for text + image or text + audio. It combines LFM2.5-Encoder-350M with vision and audio encoders, and leads our text benchmark comparison in toxicity detection and paraphrase identification. Use it for voice-command routing, on-device moderation, and intent classification. 3/

  3. OpenRouter42

    Eleven v4 is their most expressive voice model, with audio tags like [whispering]. v4 Turbo keeps that quality for real-time agents, Flash v2.5 is built for ultra-low latency, and Scribe v2 Medical is tuned for clinical conversations.

    Eleven v4 is their most expressive voice model, with audio tags like [whispering]. v4 Turbo keeps that quality for real-time agents, Flash v2.5 is built for ultra-low latency, and Scribe v2 Medical is tuned for clinical conversations.

  4. OpenRouter34

    ElevenLabs is officially on OpenRouter. You can now call 9 Text to Speech models, with up to 90+ languages, and 2 Speech to Text models with speaker labels and word-level timestamps. All @ElevenLabs models are 50% off, only on OpenRouter, through October 19, 8am PT.

    ElevenLabs is officially on OpenRouter. You can now call 9 Text to Speech models, with up to 90+ languages, and 2 Speech to Text models with speaker labels and word-level timestamps. All @ElevenLabs models are 50% off, only on OpenRouter, through October 19, 8am PT.

  5. Testing Catalog25

    GOOGLE 🔥: Besides preparation for Gemini 4 release on Antigravity, Google is also prototyping its own voice agent, internally called "Concierge". Yet, it looks like a very early version, and it is hard to say if it will ever see the light of day in its current form.

    GOOGLE 🔥: Besides preparation for Gemini 4 release on Antigravity, Google is also prototyping its own voice agent, internally called "Concierge". Yet, it looks like a very early version, and it is hard to say if it will ever see the light of day in its current form.

  6. ElevenLabs Blog14

    Contact center automation guide explains AI tools for faster customer support

    Contact center automation uses AI to handle customer support workflows with little or no human intervention, including voice, chat, and email. Unlike traditional IVR systems, AI contact center software understands intent, retrieves customer data, and routes complex cases to human agents. The guide cites Klarna, Rohlik, and Getmobil deployments of ElevenAgents, with Klarna offering voice support to 35 million US customers.

Oct 6

Oct 6Tue
  1. ElevenLabs34

    At the ElevenLabs Summit in Bengaluru, we signed our first MOU with an Indian state government to pilot voice AI. Over 70% of the 100M+ ElevenAgents conversations in India over the past year were in Hindi, Kannada, Tamil, and Telugu. Voice AI in India is multilingual, and the Summit brought together 500+ enterprise leaders, builders, and partners who are building for this future.

    At the ElevenLabs Summit in Bengaluru, we signed our first MOU with an Indian state government to pilot voice AI. Over 70% of the 100M+ ElevenAgents conversations in India over the past year were in Hindi, Kannada, Tamil, and Telugu. Voice AI in India is multilingual, and the Summit brought together 500+ enterprise leaders, builders, and partners who are building for this future.

  2. meng shao52

    xAI Cookbook adds five apps, expanding Grok API examples to ten

    The xAI Cookbook now has ten runnable Grok API examples across three tracks: real-time voice agents, multimodal generation, and live X data analysis. The author says four voice examples show the same Realtime Voice API across WebSocket, WebRTC, Twilio phone, and mobile transports. The four multimodal examples chain understanding, image generation or editing, video, and TTS, with Grok making creative decisions and Imagine models executing them.

  3. OpenRouter Blog62

    ElevenLabs text-to-speech and speech-to-text models now available on OpenRouter

    ElevenLabs now offers nine Text to Speech models and two Speech to Text models through OpenRouter, callable with an OpenRouter API key and no separate ElevenLabs plan. All ElevenLabs models are 50% off OpenRouter's list price through October 19, 8am PT, and Eleven v4, v4 Turbo, and Scribe v2 are recommended as starting points for narration, voice agents, and transcription.

    Why it matters: The source gives a concrete three-step build path and model selection guidance, showing how speech models plug into an existing text API for voice agents and transcription.

  4. Google Gemini29

    When you share your camera in Gemini Live with Guided Vision turned on, Gemini acts as a conversational partner. If your camera is pointed too high, too close, or slightly off to the side, Gemini provides natural verbal prompts to help you reframe, asking you to pan slowly to the right, tilt downward, or step back so it can get the clear visual context needed to answer your questions.

    When you share your camera in Gemini Live with Guided Vision turned on, Gemini acts as a conversational partner. If your camera is pointed too high, too close, or slightly off to the side, Gemini provides natural verbal prompts to help you reframe, asking you to pan slowly to the right, tilt downward, or step back so it can get the clear visual context needed to answer your questions.

  5. Google Gemini46

    Built alongside the blind and low-vision community, Guided Vision in Gemini Live offers conversational, real-time visual assistance. Now you can share your camera to receive dynamic audio descriptions and natural verbal reframing cues to explore your environment. 🧵

    Built alongside the blind and low-vision community, Guided Vision in Gemini Live offers conversational, real-time visual assistance. Now you can share your camera to receive dynamic audio descriptions and natural verbal reframing cues to explore your environment. 🧵

  6. eric zakariasson22

    5. podcast from a link paste an article or a pdf and two hosts talk it through. each line gets voiced as soon as grok writes it, so the episode starts playing before the script is finished. turn the sound on for this one. https://github.com/xai-org/xai-cookbook/tree/main/examples/podcast-from-a-link

    5. podcast from a link paste an article or a pdf and two hosts talk it through. each line gets voiced as soon as grok writes it, so the episode starts playing before the script is finished. turn the sound on for this one. https://github.com/xai-org/xai-cookbook/tree/main/examples/podcast-from-a-link

  7. eric zakariasson36

    1. storyboard to short film give it a one-line premise. grok-4.7 plans four shots, grok imagine draws and animates each one, and text-to-speech reads the narration. every shot is an edit of the first keyframe, which is how the cat stays the same cat. https://github.com/xai-org/xai-cookbook/tree/main/examples/storyboard-to-film

    1. storyboard to short film give it a one-line premise. grok-4.7 plans four shots, grok imagine draws and animates each one, and text-to-speech reads the narration. every shot is an edit of the first keyframe, which is how the cat stays the same cat. https://github.com/xai-org/xai-cookbook/tree/main/examples/storyboard-to-film

  8. ElevenLabs20

    Ask ElevenAgents Architect: - Why customers asked for a human agent on refund calls - How to improve resolution rates on account queries - To build a test set that keeps your agent on brand It knows best practices, analyzes your transcripts, then suggests and implements improvements.

    Ask ElevenAgents Architect: - Why customers asked for a human agent on refund calls - How to improve resolution rates on account queries - To build a test set that keeps your agent on brand It knows best practices, analyzes your transcripts, then suggests and implements improvements.

Oct 5

Oct 5Mon
  1. Liquid AI · new models on Hugging Face44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    LiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.

  2. LiveKit30

    Listen to this support call using @MicrosoftAI's speech models and LiveKit Agents. • MAI-Transcribe-2-Streaming hears you • Gemma 4 on LiveKit Inference reasons and calls tools • MAI-Voice-2.1-Flash answers What do you think? STT: https://livek.it/E4qoZvU TTS: https://livek.it/RsCc37l

    Listen to this support call using @MicrosoftAI's speech models and LiveKit Agents. • MAI-Transcribe-2-Streaming hears you • Gemma 4 on LiveKit Inference reasons and calls tools • MAI-Voice-2.1-Flash answers What do you think? STT: https://livek.it/E4qoZvU TTS: https://livek.it/RsCc37l

  3. ElevenLabs8

    Eric Glyman is coming to ElevenLabs Summit NYC. The co-founder and co-CEO of @tryramp will join @Mati on stage to talk about building one of fintech's fastest-growing companies and what happens when AI accelerates every function. @eglyman takes the stage November 11.

    Eric Glyman is coming to ElevenLabs Summit NYC. The co-founder and co-CEO of @tryramp will join @Mati on stage to talk about building one of fintech's fastest-growing companies and what happens when AI accelerates every function. @eglyman takes the stage November 11.

  4. ElevenLabs22

    The best ads have a jingle that you're still humming the next day. @ElevenCreative is launching The Search, a $100,000 competition to find the world's catchiest ad. $50,000 for first place. 11 winners… and each gets a 1:1 session with the ElevenLabs Creative Production Team.

    The best ads have a jingle that you're still humming the next day. @ElevenCreative is launching The Search, a $100,000 competition to find the world's catchiest ad. $50,000 for first place. 11 winners… and each gets a 1:1 session with the ElevenLabs Creative Production Team.

  5. vLLM23

    Great work from @fractalyze_io optimizing Qwen3-Omni on vLLM-Omni for a single RTX 5090. Their AWQ-4bit, batch-1 text-prompt tests cut time to first audio from 213ms to 23ms vs. stock vLLM-Omni. We’d love to see these optimizations contributed upstream to vLLM-Omni so more users can benefit!😄

    Great work from @fractalyze_io optimizing Qwen3-Omni on vLLM-Omni for a single RTX 5090. Their AWQ-4bit, batch-1 text-prompt tests cut time to first audio from 213ms to 23ms vs. stock vLLM-Omni. We’d love to see these optimizations contributed upstream to vLLM-Omni so more users can benefit!😄

Oct 4

Oct 4Sun
  1. IT之家 · 人工智能38

    SKF uses AI to recreate late actress Greta Garbo in advertisement

    Swedish bearing maker SKF has used AI to recreate Hollywood star Greta Garbo, who died in 1990, in an advertisement. The AI-generated figure says it returns for one final work, with its image built from text prompts and its voice trained on audio from one of Garbo's early films. SKF said the project was approved by Garbo's estate and family.