Skip to contentSkip to stories
Updated

#Voice

Oct 8

  1. Xiaomi MiMoAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    AIXiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    Why it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  2. Sierra BlogAI score62

    Sierra launches fleming-1 to detect AI agents calling by phone

    AISierra has launched fleming-1, a model that analyzes caller speech in real time and scores audio for signs it was generated by AI. It flags likely AI callers while keeping real people unflagged by default, and companies decide how to handle those calls. The model works with any voice agent built on Sierra, and Sierra also announced Personal Agent Protocol, an open standard for authorized agent-to-business interactions.

    Why it matters: The post explains why companies need to know when a caller is an AI agent, which frames the detection model as a business decision rather than an automatic block.

Oct 6

  1. OpenRouter BlogAI score62

    ElevenLabs text-to-speech and speech-to-text models now available on OpenRouter

    AIElevenLabs now offers nine Text to Speech models and two Speech to Text models through OpenRouter, callable with an OpenRouter API key and no separate ElevenLabs plan. All ElevenLabs models are 50% off OpenRouter's list price through October 19, 8am PT, and Eleven v4, v4 Turbo, and Scribe v2 are recommended as starting points for narration, voice agents, and transcription.

    Why it matters: The source gives a concrete three-step build path and model selection guidance, showing how speech models plug into an existing text API for voice agents and transcription.

Oct 1

  1. Google · Gemini appAI score60

    Google launches Guided Vision in Gemini Live for blind and low-vision users

    AIGoogle is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.

    Why it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.

Sep 30

  1. Kling AI BlogAI score62

    Kling 4.0 enters early access with 30-second native video generation

    AIKling 4.0 is entering early access, with a wider rollout planned for October, and Kling 4.0 Flash opens to Ultra Yearly subscribers on September 28. The update generates videos up to 30 seconds in a single pass, accepts up to 15 reference assets, and supports up to 10 keyframe images. Upcoming features include 10-bit HDR output at 4K and 1080p and video extension up to 2 minutes.

    Why it matters: The post specifies concrete capability limits such as 30-second native generation, up to 15 references, and 10 keyframes, which help users judge fit for production workflows.

Sep 24

  1. Azure BlogAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  2. Google DeepMindAI score62

    Google DeepMind adds Live Avatar to Gemini 3.8 Live for enterprise

    AIGoogle DeepMind has launched Gemini 3.8 Live with Live Avatar, which adds near real-time visual presence to its native live dialogue models. The feature is available today in Gemini Enterprise, supports 97 languages with adaptive lip-sync, and allows custom avatars through enterprise allowlisting. All output carries an imperceptible SynthID watermark.

    Why it matters: The post specifies the new avatar capabilities, the Gemini Enterprise access path, and the SynthID watermark, which helps readers judge its enterprise deployment fit.

  3. Google · Gemini appAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

Sep 23

  1. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  2. Baseten BlogAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

Sep 22

  1. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 15

  1. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  2. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 9

  1. Microsoft Foundry BlogAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

May 10

  1. Thinking Machines LabAI score67

    Thinking Machines Lab previews interaction models for real-time human-AI collaboration

    AIThinking Machines Lab announced a research preview of interaction models that take in audio, video, and text continuously and respond in real time without external turn-detection harnesses. The model, TML-Interaction-Small, is a 276B-parameter MoE with 12B active parameters, paired with an asynchronous background model for sustained reasoning and tool use. The post reports competitive intelligence scores and lower turn-taking latency against GPT-realtime and Gemini Live models, along with new interactivity benchmarks where baseline models largely failed.

    Why it matters: The post explains a time-aligned, full-duplex design and benchmarks against turn-based models, showing how interaction and background reasoning can be split across two cooperating models.

Mar 17

  1. Xiaomi MiMoAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

Jan 21

  1. Mistral AI · new models on Hugging FaceAI score65

    Mistral releases open-weight Voxtral Mini 4B Realtime 2602 speech model

    AIMistral AI released Voxtral Mini 4B Realtime 2602, a multilingual realtime speech-transcription model with 13 supported languages under the Apache 2.0 license. The model has a configurable transcription delay from 240ms to 2.4s, and it matches leading offline open-source models at a 480ms delay. The source says it is optimized for on-device deployment and is currently supported only in vLLM.

    Why it matters: The source specifies the 480ms delay operating point, 4B size, Apache 2.0 license, and vLLM serving path, which matter for teams weighing realtime transcription deployment.

That’s everything