We got ESP32 at home
We got ESP32 at home
We got ESP32 at home
IndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.
IndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.
IndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.
Huge thank you to everyone who downloaded Nemotron 3 Diarization and helped it trend on @huggingface! And we appreciate all the comments. @sabbassi_11 answered a few of your questions:
Many thanks to The Information for sitting down with our Executive Chairman Sean Parker and CEO @premakkaraju to talk about our work building tools for music professionals.
@AssemblyAI Universal 3.6 Pro is live in LiveKit Inference! - 45% fewer wrong yes/no confirmations - ~30% less background speech - 32 languages + code-switching - Endpointing that waits out phone numbers & emails Swap to universal-3-6-pro, same $0.45/hr Try it now > https://docs.livekit.io/agents/models/stt/assemblyai/
ElevenReader is available on iOS, Android, and web. Listen to anything you want to read. https://elevenreader.io
Eleven v4, our most expressive voice model, is now in ElevenReader. Listen to any article, ebook, or PDF in a voice you choose, in 90+ languages.
we're releasing an experimental @SpaceXAI typescript sdk! npm install @xai-official/sdk get text, voice, image and video in one sdk, with the latest grok models and tools that run on our servers, like real-time X search, web search, code execution and remote mcp
A speech model can support Arabic and still struggle with local dialects. Fine-tuning NVIDIA Nemotron 3.5 ASR cut word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. Check out our new tutorial and learn how to adapt it for other dialects and languages: https://nvda.ws/4hYqGPM
ElevenLabs has achieved FedRAMP 20x Class A certification, covering ElevenAgents and its Text to Speech and Speech to Text APIs when run in Zero Retention Mode with US data residency. The company says the published evidence package and FedRAMP Marketplace listing should shorten security reviews for public sector and regulated buyers. The certification extends ElevenLabs for Government, its dedicated offering for federal agencies.
Build expressive voice agents with MAI models now in LiveKit. Learn more here: https://msft.it/6014alofr
MAI voice models are also available through OpenRouter. Learn more here: https://msft.it/6012aloIC
Expect some odd outputs while it's in beta. Your feedback helps makes it better. https://suno.com/s/jx9kt9ovaezNzfJm
Build your agents with the best Voice and Transcribe Streaming model in the world... now easily available to use on @vercel #1 for quality and speed... cheaper than any other hyperscaler... and a whopping 60% cheaper than Eleven Labs
Google's September 2026 roundup highlights Gemini 4 Argon, a frontier model with a 1-million-token output limit aimed at complex tasks such as cybersecurity defense. Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with developer, enterprise, and consumer access to follow after guardrail feedback. The post also covers Gemini 3.8 Flash, Connected Apps in Gemini, and WeatherNext 3.
🗣️ Ever felt like AI bots don't really listen? 🗣️ They just don't have a li'l Jev helping them yet!
@MicrosoftAI @MicrosoftAI STT plugin guide here: https://docs.livekit.io/agents/models/stt/microsoft-ai/
MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash from @MicrosoftAI are now live on LiveKit. Transcribe-2-Streaming debuts at #1 on the Artificial Analysis accuracy leaderboard → Pair it with MAI-Voice-2.1-Flash to build efficient and expressive voice agents. Try them out: https://docs.livekit.io/agents/models/tts/microsoft-ai/
The author jokes about founding a consulting firm that would use agents for work, Griffin for meetings, and Griffin for interviews to fill remote roles. They then question whether remote engineers and consultancies would still be needed if that became reality. The quoted Tavus post says Griffin passed a video Turing test with 48% of live interlocutors believing it was human.
More ways to discover, access, and build with MAI models. MAI models are now available through Vercel, giving developers another way to bring Microsoft AI models into the products and experiences they’re building. This includes our existing MAI models plus today’s newest releases: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.
The @MicrosoftAI team is training excellent models. Excited to bring them to @vercel, day zero.
We're launching the most accurate real time transcription model in the world... #1 !!! 55% faster and 60% cheaper than ElevenLabs. Come build agents on our platform!
Microsoft AI announced three new models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The source says the streaming transcription model is accurate and aims for natural speech with less waiting between conversational turns for building voice agents.
Google is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.
AIWhy it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.
Jensen Huang is coming to ElevenLabs Summit NYC. NVIDIA has partnered with ElevenLabs since our earliest days, and it's a big part of how we keep pushing what voice and audio AI can do. On November 11, @JensenHuang will join @mati for a conversation on the future of AI.
Meta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.
Suno has launched Speech in beta, which it describes as the first audio model that generates voice and music together as one cohesive track. Users type an idea or existing text, then describe the voice and musical style they want, and the beta is now open to everyone after a month of testing with a small group.
IndexTeam has released Index-Translate, a multilingual family covering text, speech, dubbing, and long-document translation across 150 languages. Its 9B model scores 0.8789 on FLORES, 0.8209 on instTrans, and 0.7387 on MEME, and the 2B and 9B models are released under Apache 2.0.
Kling 4.0 is entering early access, with a wider rollout planned for October, and Kling 4.0 Flash opens to Ultra Yearly subscribers on September 28. The update generates videos up to 30 seconds in a single pass, accepts up to 15 reference assets, and supports up to 10 keyframe images. Upcoming features include 10-bit HDR output at 4K and 1080p and video extension up to 2 minutes.
AIWhy it matters: The post specifies concrete capability limits such as 30-second native generation, up to 15 references, and 10 keyframes, which help users judge fit for production workflows.
Hugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.
By far one of the most interesting parts of Dev Day is - demo failures or not - those folks attempted to voice dictate EVERYTHING. If you’re not using voice AI, it’s clear you should. Future products, features, and experiences will be based around it. Time to get that DJI mic! 🗣️
Introducing OpenAI’s Dots x Higgsfield. Your always-on Higgsfield creative crew keeps working while you’re away. Check in by text, call or email, and pause the work whenever you need. Powered by GPT-6.1 Sol.
Deedy describes a video generation pipeline built around Opus 5.5 in Claude Code, routing image, video, audio, and TTS models through OpenRouter's single API key. The workflow adds reference-image consistency, animatics before full renders, a critic skill that screenshots and transcribes output for QA, and ffmpeg for most editing.
MiniCPM-o 4.5 is now in SGLang Omni v0.1.7. More flexibility for developers to run and build with the model.
Audio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.
An OpenAI developer experience team member argues voice agents need not always speak back, outlining speech-to-speech, speech-to-action, and event-to-speech as emerging design modes. He cites form filling, creative tools, and computer use as examples of speech-to-action, which he calls among the most underexplored areas. He says event-to-speech is still very exploratory, with hands-free recipe guidance and proactive alerts as examples.
so many styles and even more ways AI can help you, hands-free. even lighter, with the longest battery life to date. #MetaConnect
Deedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.