Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 22

Sep 22Tue
  1. Daniel HanAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  2. OpenBMBAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post
  3. TechNode · AIAI score60

    Alibaba's T-Head unveils Zhenwu V900 AI chip with full-stack system design

    AIT-Head, Alibaba's chip subsidiary, unveiled the Zhenwu V900 AI chip for training and inference at the 2026 Apsara Conference in Hangzhou. The company claims three times the performance of its predecessor, the Zhenwu M890, with 216GB of memory, 1,200GB/s inter-chip bandwidth, and mass production expected in the first quarter of 2027.

  4. Lovable BlogAI score38

    Lovable joins Blueprint Alliance to advance an open architecture for securing AI agents

    AILovable joined AWS, Google Cloud, Databricks, Salesforce, and other firms as a founding member of the Blueprint Alliance, a coalition developing an open reference architecture for securing and governing enterprise AI agents. The blueprint covers registering agents as identities with accountable owners, scoping their access to tasks, enforcing policies through gateways, and responding to incidents by revoking tokens or quarantining agents.

  5. Black Forest Labs · new models on Hugging FaceAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  6. AI SupremacyAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

  7. Tencent HyAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

  8. OpenBMBAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

  9. X.PINAI score46

    Moonshot's Kimi K3 now available on Amazon Bedrock

    AIMoonshot's Kimi K3 is now available on Amazon Bedrock, with its license requiring a paid agreement for model-hosting businesses and affiliates above $20M in annual revenue. AWS says customer data stays within its cloud, is not shared with Moonshot or used for training, and inference requests have zero data retention. Neither company disclosed financial terms.

    Image from @thexpin's post
  10. X.PINAI score42

    Alibaba targets 20GW cloud capacity by 2032, unveils Zhenwu V900 chip

    AIAlibaba CEO Eddie Wu said on September 22 that strong AI demand is driving infrastructure investment, targeting over 20GW of global cloud data-center capacity by 2032. T-Head unveiled the Zhenwu V900, claiming 3× the compute performance of the M890 and support for clusters of up to 500,000 accelerator cards. Servers using V900 chips are scheduled to launch in Q1 2027.

    Image from @thexpin's post
  11. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. Tencent HyAI score38

    Tencent Hunyuan releases Hy Image3.5 preview for image generation

    AITencent Hunyuan has launched a preview of Hy Image3.5, which it says wins 30% more often than Hy Image3.0 in human evaluation. The model supports text-to-image and image-to-image generation at up to 2K resolution with improved consistency, and is priced at $0.024 per image on the Tencent Cloud API, with reference images free. Two weeks of free access is offered through OnSolo and Miora.

    Video from @TencentHunyuan's post
  2. Tencent HyAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  3. xAI News (Grok)AI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  4. vLLM BlogAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  5. Amp NewsAI score36

    Amp Runners Add Git Worktree Creation and Secret Injection

    AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.

  6. Google Developers BlogAI score38

    Google Colab premium benefits now included in Google AI plans

    AIGoogle AI subscribers now get premium Colab benefits, including priority access to faster accelerators and more powerful machines. Google AI Ultra subscribers also get uninterrupted background execution and Premium GPU access for long training runs. The benefits roll out over the next few weeks in Colab-supported countries, and existing Colab subscriptions are unchanged.

  7. Together AI BlogAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.

  8. Latent.SpaceAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  9. Xiaomi MiMoAI score44

    MiMo-V2.6-Pro assists scientific research in materials and formal mathematics

    AIXiaomi's MiMo-V2.6-Pro, without research-specific RL training, helped Xiaomi materials researchers propose MOF materials for capturing PFAS "forever chemicals" and ran computational screening for wet-lab validation. It also helped formalize the full main theorem of Li–Yorke's "Period Three Implies Chaos" in Lean 4, producing a project of 6,000+ lines verified by Lean's kernel with no unfinished proof placeholders.

    Video from @XiaomiMiMo's post
  10. Mike KnoopAI score38

    Mike Knoop says LLM logprobs are vanishing, yet they enable useful new patterns

    AIMike Knoop notes that logprobs used to be widely exposed by LLM inference APIs and sees the market maturing so that parts of the LLM stack can be packaged in new, useful ways. He links this to Bryan Helmig's post on prompting with max_tokens: 1 plus logprobs for fast, parallel judgments, which Helmig says has a lot more depth than he expected.

  11. Amazon ScienceAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  12. ChatGPTAI score40

    ChatGPT now lets US Plus and Pro users track credit scores via Experian

    AIChatGPT users on Plus and Pro plans in the U.S. can now securely connect their Experian credit report and VantageScore 3.0 credit score in Finances. Once connected, the feature provides personalized insights into what is affecting their score and how it relates to their financial goals, with alerts when things change. It is available on web and the latest iOS and Android apps.

    Image from @ChatGPT's post