Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. ModelScopeOfficialAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post
  2. QwenOfficialAI score62

    Qwen-Audio-3.1 upgrades ASR, TTS and Realtime and adds two new models

    AIAlibaba's Qwen team released Qwen-Audio-3.1, upgrading its ASR, TTS and Realtime models and adding TTS-Next and ASR-Next. The post says TTS prices fell about 70%, Realtime about 85%, and ASR up to 95%. More APIs are coming soon.

    Image from @Alibaba_Qwen's post
  3. KrASIA · Big TechNewsAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  4. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Google Developers BlogOfficialAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    AIGoogle announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

    Why it matters: The source shows how to run an agent with a local Gemma 4 26B model using LiteRT, plus a hybrid cloud-planner pattern that keeps most tokens on-device.

  2. Together AI BlogOfficialAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  3. Fireworks AI BlogOfficialAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  4. Fireworks AI BlogOfficialAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  5. Google GemmaOfficialAI score31

    Deploy DiffusionGemma-Jev on Google Cloud Run with one command

    AIGoogle Gemma says DiffusionGemma-Jev (djev) can now be deployed as a Jev API-compatible endpoint on Google Cloud Run with a single command. The post reports about 35-60 ms single-step latency and roughly 100-123 requests/sec at batch size 32, at about $3/hr that drops to $0 when idle.

    Video from @googlegemma's post
  6. Google GemmaOfficialAI score22

    Google Gemma credits DiffusionGemma-Jev deployment on Cloud Run

    AIGoogle Gemma credits @mmastrac and @dylayed for work on DiffusionGemma-Jev (djev), a Jev API-compatible endpoint. Per @dylayed, djev can be deployed to Google Cloud Run with a single gcloud command, at roughly $3/hr while active and $0 when idle.

  7. Grok BotOfficialAI score18

    Bots now respond faster and complete tasks more efficiently

    AIThe post says bots respond faster and complete tasks more efficiently, following user feedback that the bot was a poor texter. The author adds that usage rose about 6% on average, and promises more improvements soon.

  8. Grok BotOfficialAI score14

    Grok desktop app gets 53 performance fixes for faster reconnects and wake

    AIxAI's Grok desktop app now reconnects after a flaky connection in 0.7 seconds, down from 60 seconds, and wakes from sleep in 1 second instead of 23. The update includes 53 performance fixes, with the Media tab no longer re-downloading files, cutting its data transfer from 137MB to 0.7MB, and lower idle memory use.

  9. ZyphraOfficialAI score20

    Zyphra's Beren Millidge on why multi-silicon AI infrastructure matters

    AIZyphra's Chief Scientist Beren Millidge, in an AI Infra Summit interview with vCluster Labs CEO Lukas Gentele, argued that a heterogeneous compute future is inevitable. The interview covers why Zyphra chose AMD over NVIDIA, along with topics such as kernel writing, surviving GPU failures mid-run, and routing. Zyphra says it is working to build a strong multi-silicon ecosystem.

  10. Tibor BlahoXAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

    Image from @btibor91's post
  11. Tri DaoXAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  12. eric zakariassonXAI score13

    Cursor's scrappy support system evolved into a scalable, self-improving operation

    AICursor's support team says a scrappy version built about 1.5 years ago to handle heavy volume taught them to crawl, walk, then run, and to identify flywheels for continuous improvement. The background post notes the company rebuilt customer support around Grok Bot to handle operations without adding headcount, with the bot responding to customers, resolving tickets, and managing the queue autonomously.

  13. Amazon ScienceOfficialAI score36

    Amazon's Peter DeSantis on AI hardware's shifting bottlenecks

    AIAmazon SVP Peter DeSantis told SemiAnalysis's Dylan Patel at the AI Infra Summit that AI workloads are shifting from being power-bound to memory-bandwidth-bound and then memory-bound. He called designing hardware for these changing constraints one of the most interesting hardware design problems in years.

  14. Greg BrockmanXAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.

  15. Sierra BlogOfficialAI score34

    Sierra Lets Companies See, Edit, and Export Their AI Agents' Logic and Data

    AISierra says its platform makes enterprise AI agents visible and editable, with journeys, policies, and actions viewable in Agent Studio and testable through Simulations and Experiments before rollout. Customers can export agent logic in a portable structured format, access conversation logs and performance data through export APIs, and manage the agent's code in a Git repository. Sierra agents also connect to existing systems through MCP, REST, GraphQL, or custom integrations.

  16. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

  17. GranolaOfficialAI score13

    First Meet + Zoom, now Teams!

    AINow Granola can tag who's speaking on Teams meetings – head to Settings > Speaker tags to get started.

    Image from @meetgranola's post
  18. StepFunOfficialAI score43

    StepFun open-sources onPanda for token-level LLM annotation and inspection

    AIStepFun has open-sourced onPanda, a tool used internally for LLM data annotation and model inspection, letting users correct tokens and let models continue. The company reports a 52% lower median annotation time versus manual post-editing, with SFT and preference data combined in one workflow. It also supports token probability and top-k inspection, token-by-token decoding control, and browser-based testing across SVG generation, web development, and agent tasks.

  19. WaymoOfficialAI score32

    Waymo launches Teen Accounts for riders in Nashville

    AIWaymo has made Teen Accounts officially available in Nashville, giving teens more independence while letting parents feel reassured. The post points readers to for details.

    Video from @Waymo's post
  20. LlamaIndex 🦙OfficialAI score22

    LiteParse v2.14.6 parses text PDFs about 25% faster locally

    AILlamaIndex released LiteParse v2.14.6, an open-source PDF-to-Markdown parser that processes text-based PDFs about 25% faster. On realistic documents it handled pages at 2.8ms per page, 1.5 times faster than the next-fastest local parser. It runs locally in Python, Node.js, Rust, or directly in the browser.

    Image from @llama_index's post
  21. Cognition Blog (Devin, Windsurf)OfficialAI score26

    Cognition Expands to Latin America, Launching São Paulo Hub for Devin Software Engineering

    AICognition announced its expansion into Latin America at MASP in São Paulo, starting with a local team to help companies build more of their software in the region. Itaú reports more than 75% of its technology teams use Devin, with legacy .NET services migrated to Java 6x faster and about 70% of security vulnerabilities resolved automatically. Nubank says Devin cut a multi-million-line monolith migration from years to weeks, at over 20x lower cost.

  22. Mike KriegerXAI score62

    Anthropic cuts API prices to $4 and $20 per million tokens

    AIThe company cut API pricing to $4 per million input tokens and $20 per million output tokens, which it says is 20% less than Opus 5. Cache reads are also 60% cheaper, and subscribers get higher five-hour rate limits plus a banked rate limit reset.

  23. Lovable BlogOfficialAI score38

    Lovable Adds Opus 5.5, Cutting Build Steps by a Third to Half at Same Quality

    AILovable now offers Opus 5.5, which it says matches Opus 5's results while finishing builds in a third to half fewer steps. Internal benchmarks showed Opus 5.5 scoring 4 to 6% ahead of Opus 5 on verification discipline, with step reductions of 26% to 57% and input token reductions of 21% to 59% across tasks.

  24. Daniel HanXAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  25. WorkBuddyOfficialAI score18

    HKUST students build two AI workbenches with WorkBuddy, win Game Track

    AIHKUST's Anchor team used WorkBuddy to build two production-ready workbenches and won the Game Track championship. Kaiwu Producer creates a complete FPS game in 8 hours through full-pipeline 3D generation with an AI-driven narrative memory engine and zero human intervention. Anchor is a de-labeling narrative engine that automatically detects stereotypical dependencies.

    Video from @WorkBuddy_AI's post
  26. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post