Skip to content

#Voice

Oct 4

Oct 4Sun

Oct 3

Oct 3Sat
  1. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Echo-S2ST-9B-FP4 released as NVFP4 quantized speech translation model

    IndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  2. IndexTeam (Bilibili) · new models on Hugging FaceAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    IndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  3. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

    IndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

Oct 2

Oct 2Fri
  1. LiveKitAI score23

    @AssemblyAI Universal 3.6 Pro is live in LiveKit Inference! - 45% fewer wrong yes/no confirmations - ~30% less background speech - 32 languages + code-switching - Endpointing that waits out phone numbers & emails Swap to universal-3-6-pro, same $0.45/hr Try it now > https://docs.livekit.io/agents/models/stt/assemblyai/

    @AssemblyAI Universal 3.6 Pro is live in LiveKit Inference! - 45% fewer wrong yes/no confirmations - ~30% less background speech - 32 languages + code-switching - Endpointing that waits out phone numbers & emails Swap to universal-3-6-pro, same $0.45/hr Try it now > https://docs.livekit.io/agents/models/stt/assemblyai/

  2. eric zakariassonAI score47

    we're releasing an experimental @SpaceXAI typescript sdk! npm install @xai-official/sdk get text, voice, image and video in one sdk, with the latest grok models and tools that run on our servers, like real-time X search, web search, code execution and remote mcp

    we're releasing an experimental @SpaceXAI typescript sdk! npm install @xai-official/sdk get text, voice, image and video in one sdk, with the latest grok models and tools that run on our servers, like real-time X search, web search, code execution and remote mcp

  3. NVIDIA AIAI score40

    A speech model can support Arabic and still struggle with local dialects. Fine-tuning NVIDIA Nemotron 3.5 ASR cut word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. Check out our new tutorial and learn how to adapt it for other dialects and languages: https://nvda.ws/4hYqGPM

    A speech model can support Arabic and still struggle with local dialects. Fine-tuning NVIDIA Nemotron 3.5 ASR cut word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. Check out our new tutorial and learn how to adapt it for other dialects and languages: https://nvda.ws/4hYqGPM

  4. ElevenLabsAI score32

    ElevenLabs earns FedRAMP 20x Class A certification for federal voice AI

    ElevenLabs has achieved FedRAMP 20x Class A certification, covering ElevenAgents and its Text to Speech and Speech to Text APIs when run in Zero Retention Mode with US data residency. The company says the published evidence package and FedRAMP Marketplace listing should shorten security reviews for public sector and regulated buyers. The certification extends ElevenLabs for Government, its dedicated offering for federal agencies.

  5. Mustafa SuleymanAI score36

    Build your agents with the best Voice and Transcribe Streaming model in the world... now easily available to use on @vercel #1 for quality and speed... cheaper than any other hyperscaler... and a whopping 60% cheaper than Eleven Labs

    Build your agents with the best Voice and Transcribe Streaming model in the world... now easily available to use on @vercel #1 for quality and speed... cheaper than any other hyperscaler... and a whopping 60% cheaper than Eleven Labs

  6. Google · AI blogAI score58

    Google recaps September 2026 AI launches, led by Gemini 4 Argon

    Google's September 2026 roundup highlights Gemini 4 Argon, a frontier model with a 1-million-token output limit aimed at complex tasks such as cybersecurity defense. Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with developer, enterprise, and consumer access to follow after guardrail feedback. The post also covers Gemini 3.8 Flash, Connected Apps in Gemini, and WeatherNext 3.

Oct 1

Oct 1Thu
  1. LiveKitAI score40

    MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash from @MicrosoftAI are now live on LiveKit. Transcribe-2-Streaming debuts at #1 on the Artificial Analysis accuracy leaderboard → Pair it with MAI-Voice-2.1-Flash to build efficient and expressive voice agents. Try them out: https://docs.livekit.io/agents/models/tts/microsoft-ai/

    MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash from @MicrosoftAI are now live on LiveKit. Transcribe-2-Streaming debuts at #1 on the Artificial Analysis accuracy leaderboard → Pair it with MAI-Voice-2.1-Flash to build efficient and expressive voice agents. Try them out: https://docs.livekit.io/agents/models/tts/microsoft-ai/

  2. Dongxi NLP (东锡)AI score46

    Dongxi jokes about replacing remote consultants with Griffin AI agents

    The author jokes about founding a consulting firm that would use agents for work, Griffin for meetings, and Griffin for interviews to fill remote roles. They then question whether remote engineers and consultancies would still be needed if that became reality. The quoted Tavus post says Griffin passed a video Turing test with 48% of live interlocutors believing it was human.

  3. Microsoft AIAI score36

    More ways to discover, access, and build with MAI models. MAI models are now available through Vercel, giving developers another way to bring Microsoft AI models into the products and experiences they’re building. This includes our existing MAI models plus today’s newest releases: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.

    More ways to discover, access, and build with MAI models. MAI models are now available through Vercel, giving developers another way to bring Microsoft AI models into the products and experiences they’re building. This includes our existing MAI models plus today’s newest releases: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.

  4. Google · Gemini appAI score60

    Google launches Guided Vision in Gemini Live for blind and low-vision users

    Google is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.

    AIWhy it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.

  5. ElevenLabsAI score34

    Jensen Huang is coming to ElevenLabs Summit NYC. NVIDIA has partnered with ElevenLabs since our earliest days, and it's a big part of how we keep pushing what voice and audio AI can do. On November 11, @JensenHuang will join @mati for a conversation on the future of AI.

    Jensen Huang is coming to ElevenLabs Summit NYC. NVIDIA has partnered with ElevenLabs since our earliest days, and it's a big part of how we keep pushing what voice and audio AI can do. On November 11, @JensenHuang will join @mati for a conversation on the future of AI.

  6. Meta NewsroomAI score22

    Ranveer Singh Becomes Ray-Ban and Ray-Ban Meta Brand Ambassador in India

    Meta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.

Sep 30

Sep 30Wed
  1. Kling AI BlogAI score62

    Kling 4.0 enters early access with 30-second native video generation

    Kling 4.0 is entering early access, with a wider rollout planned for October, and Kling 4.0 Flash opens to Ultra Yearly subscribers on September 28. The update generates videos up to 30 seconds in a single pass, accepts up to 15 reference assets, and supports up to 10 keyframe images. Upcoming features include 10-bit HDR output at 4K and 1080p and video extension up to 2 minutes.

    AIWhy it matters: The post specifies concrete capability limits such as 30-second native generation, up to 15 references, and 10 keyframes, which help users judge fit for production workflows.

Sep 29

Sep 29Tue
  1. Hugging Face BlogAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    Hugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  2. Allie K. MillerAI score16

    By far one of the most interesting parts of Dev Day is - demo failures or not - those folks attempted to voice dictate EVERYTHING. If you’re not using voice AI, it’s clear you should. Future products, features, and experiences will be based around it. Time to get that DJI mic! 🗣️

    By far one of the most interesting parts of Dev Day is - demo failures or not - those folks attempted to voice dictate EVERYTHING. If you’re not using voice AI, it’s clear you should. Future products, features, and experiences will be based around it. Time to get that DJI mic! 🗣️

Sep 28

Sep 28Mon
  1. ModelScopeAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    Audio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

  2. Artificial IgnoranceAI score42

    OpenAI Engineer Argues Voice Agents Should Act, Not Only Talk

    An OpenAI developer experience team member argues voice agents need not always speak back, outlining speech-to-speech, speech-to-action, and event-to-speech as emerging design modes. He cites form filling, creative tools, and computer use as examples of speech-to-action, which he calls among the most underexplored areas. He says event-to-speech is still very exploratory, with hands-free recipe guidance and proactive alerts as examples.

Sep 27

Sep 27Sun
  1. DeedyAI score34

    Deedy urges explainer videos for every open source repo, citing SQLite example

    Deedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.