Skip to contentSkip to stories

Updated

#Product update

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Logan KilpatrickXAI score33

    Google releases Gemini 3.8 Flash TTS with lower cost than 3.1

    AIGoogle has released Gemini 3.8 Flash TTS, a text-to-speech model that costs less than its previous 3.1 Flash TTS model. The new model is accompanied by a new experience in Google AI Studio.

  2. Philipp SchmidXAI score38

    Google releases Gemini 3.8 text-to-speech models with a prompting guide

    AIGoogle has launched Gemini 3.8 text-to-speech models, with a blog post, a prompting guide for speech generation, and a migration guide for users moving from Gemini 3.1. The post itself provides few technical details, so specifics such as benchmarks or pricing are not stated here.

  3. Philipp SchmidXAI score62

    Gemini 3.8 Flash TTS adds voice replication and prompt-designed voices

    AIGoogle launched Gemini 3.8 Flash TTS and Flash-Lite TTS, letting users replicate a voice from 30 seconds of audio or design one from a text description. The post claims #1 on Hume's Voice Design Benchmark and top placement in Voice Arena across 6 languages. Voices can be directed line by line with style and inline tags, with a consent check on replication and SynthID on every clip. Availability is through the Gemini API and Google AI Studio.

    Image from @_philschmid's post
  4. Google AIOfficialAI score62

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle AI launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it describes as its most expressive audio models yet. The models support custom voices across 100+ languages, 2,000+ prebuilt voices, multi-speaker conversations, line-by-line delivery control, and cues such as <laughs> and |mhm|. Google positions Flash TTS for bespoke voices in gaming, audiobooks, and podcasts, and Flash-Lite TTS for near real-time voice agents, high-volume dubbing, and bulk audio.

    Video from @GoogleAI's post
  5. Google DeepMindOfficialAI score32

    Gemini API adds line-by-line control over AI speech delivery

    AIGoogle DeepMind says developers can fine-tune AI-generated audio line by line, adjusting pacing, emotion, and cues such as laughs or pauses. All generated audio is watermarked with SynthID so it can be reliably identified as AI-generated, and developers can start building with the Gemini API via Google AI Studio.

  6. Google DeepMindOfficialAI score44

    Google launches Gemini 3.8 Flash and Flash-Lite TTS models for custom audio

    AIGoogle DeepMind has introduced two text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS lets users design unique voices with distinct accents and characteristics, while Flash-Lite TTS is built for efficiency and scale, offering user-created styles or a production-ready voice library.

    Video from @GoogleDeepMind's post
  7. Google AI StudioOfficialAI score42

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS audio models

    AIGoogle introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as its most expressive audio generation models yet. The models are designed to help creators, developers, and enterprises build richer, more expressive audio experiences, and can be tried through the Gemini API and AI Studio.

    Video from @GoogleAIStudio's post
  8. Google DeepMindOfficialAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  9. Google DeepMind · YouTubeOfficialAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  10. Baseten BlogOfficialAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  11. Lewis Tunstall @ COLM 🌉XAI score40

    Lewis Tunstall makes voice acting debut with Reachy Mini robot

    AILewis Tunstall of Hugging Face shares his first cameo as a voice actor using the Reachy Mini robot, linked to a YouTube video about Australian football. The post itself provides no technical details, and the Reachy Mini context comes from a separate post by Andi Marafioti about NVIDIA's open-source Nemotron 3 Diarization model.

  12. Julien ChaumondXAI score36

    Hugging Face releases JS package to browse LeRobot datasets directly

    AIHugging Face has released a new JavaScript package, huggingface/lerobot, that lets developers read LeRobot datasets on the Hub directly in the browser without downloading them. The post says coding agents can use it to build custom dataset viewers quickly, with an example UI implemented in about 600 lines of JS.

    Image from @julien_c's post
  13. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Image from @Alibaba_Qwen's post
  14. QwenOfficialAI score62

    Qwen-Audio-3.1 upgrades ASR, TTS and Realtime and adds two new models

    AIAlibaba's Qwen team released Qwen-Audio-3.1, upgrading its ASR, TTS and Realtime models and adding TTS-Next and ASR-Next. The post says TTS prices fell about 70%, Realtime about 85%, and ASR up to 95%. More APIs are coming soon.

    Image from @Alibaba_Qwen's post
  15. KrASIA · Big TechNewsAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  16. Tencent HyOfficialAI score22

    Tencent Hunyuan's Hy Image3.5 preview free for two weeks on OnSolo

    AITencent Hunyuan has made its Hy Image3.5 preview available on the OnSolo platform, free for two weeks. The model targets short drama character sheets, full-motion video game assets, and keyframes, with characters kept consistent across episodes and edits that refine rather than regenerate images. OnSolo's background post says the preview supports 5 references at 2K resolution and is free for Members during the two-week window.

  17. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Tencent HyOfficialAI score43

    Tencent Hunyuan previews Hy Image3.5 in ComfyUI

    AITencent Hunyuan's Hy Image3.5 preview is now available in ComfyUI, with a claimed 30% higher win rate in human evaluation than Hy Image3.0. The model handles text-to-image and image-to-image in one model at up to 2K resolution, with multilingual text and small print rendering correctly. It also keeps identity and product features consistent across scene, outfit, and style changes.

  2. Daniel HanXAI score22

    Unsloth Desktop hotfix adds Qwen-Image-2.1 image editing and fixes

    AIUnsloth Desktop received a hotfix update adding image editing for Qwen-Image-2.1. The update also fixes diffusers update issues, GGUF loading failures for Qwen-Image, and black artifacts during diffusion on A100 and consumer GPUs. Users should receive a banner prompting them to update.

  3. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  4. Google Developers BlogOfficialAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    AIGoogle announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

    Why it matters: The source shows how to run an agent with a local Gemma 4 26B model using LiteRT, plus a hybrid cloud-planner pattern that keeps most tokens on-device.

  5. Fireworks AI BlogOfficialAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  6. Google GemmaOfficialAI score31

    Deploy DiffusionGemma-Jev on Google Cloud Run with one command

    AIGoogle Gemma says DiffusionGemma-Jev (djev) can now be deployed as a Jev API-compatible endpoint on Google Cloud Run with a single command. The post reports about 35-60 ms single-step latency and roughly 100-123 requests/sec at batch size 32, at about $3/hr that drops to $0 when idle.

    Video from @googlegemma's post
  7. Google GemmaOfficialAI score22

    Google Gemma credits DiffusionGemma-Jev deployment on Cloud Run

    AIGoogle Gemma credits @mmastrac and @dylayed for work on DiffusionGemma-Jev (djev), a Jev API-compatible endpoint. Per @dylayed, djev can be deployed to Google Cloud Run with a single gcloud command, at roughly $3/hr while active and $0 when idle.

  8. Tibor BlahoXAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

    Image from @btibor91's post
  9. Simon WillisonXAI score60

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which build on advances behind GPT-6 Astra. The company says it cut API prices 50% for Sol and Luna compared with GPT-5.6 promotional pricing, passing on caching and inference efficiency gains. Simon Willison notes GPT-6 Luna costs half of GPT-5.6 Luna and calls Luna his favorite model for building product features because of its cost and speed.

  10. Greg BrockmanXAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.

  11. Noam BrownXAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    AIOpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    Why it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

  12. Sierra BlogOfficialAI score34

    Sierra Lets Companies See, Edit, and Export Their AI Agents' Logic and Data

    AISierra says its platform makes enterprise AI agents visible and editable, with journeys, policies, and actions viewable in Agent Studio and testable through Simulations and Experiments before rollout. Customers can export agent logic in a portable structured format, access conversation logs and performance data through export APIs, and manage the agent's code in a Git repository. Sierra agents also connect to existing systems through MCP, REST, GraphQL, or custom integrations.

  13. Sherwin WuXAI score46

    GPT-6 Luna launches at $0.10 and $0.50 per million tokens

    AIOpenAI's GPT-6 Luna is priced at $0.10 per 1M input tokens and $0.50 per 1M output tokens, according to Sherwin Wu. Per OpenAI Devs, Luna and GPT-6 Sol launch today with API prices 50% lower than GPT-5.6. Wu jokes that per-billion-token pricing may soon be needed.