Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Tencent HyAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Liquid AI BlogAI score46

    LFM2.5-VL-DSpark speeds up vision-language model decoding on GPUs and edge devices

    AILiquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, delivering decoding throughput gains of up to 2.66× on GPUs and 3.13× on edge devices. The drafter adds about 280M parameters, an 8.9% increase in the deployed model's parameter count, and is available on Hugging Face with support in llama.cpp, SGLang, and MLX-VLM.

  3. Engineering at MetaAI score43

    Meta Brings Private Processing to AI Glasses via Confidential Cloud Computing

    AIMeta is extending its Private Processing confidential computing infrastructure to AI glasses, running AI models inside confidential virtual machines so that even Meta cannot access user data. The system relies on hardware Trusted Execution Environments, with remote attestation checked by clients before any data is sent. Meta first introduced Private Processing in 2025 for WhatsApp and the Meta AI app.

  4. vLLM BlogAI score54

    vLLM adds distortion-free Gumbel-max watermarking for text provenance

    AIvLLM now supports Gumbel-max watermarking, which embeds a keyed signal into generated text without changing the expected token distribution. Detection requires the secret key and tokenizer, and the signal accumulates over longer outputs. Benchmarks on Qwen3.5-27B with MTP-3 show throughput changes between -1.1% and +2.0% across batch sizes, with no consistent slowdown.

  5. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  6. Amp NewsAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  7. Google Developers BlogAI score62

    Google Cloud API Gateway can now expose REST APIs as MCP tools in preview

    AIGoogle Cloud API Gateway now acts as a remote MCP server in Public Preview, making REST operations in an annotated OpenAPI 3.0.x or 3.1.x spec available as agent-ready MCP tools. Existing JWT or API-key authentication, quotas, and logging apply to MCP calls, so teams do not need a separate MCP server. Current limits include no support for OpenAPI 2.0, a maximum of 1,000 tools per gateway, and no MCP and model routing in the same API config.

    Why it matters: The post shows how an existing OpenAPI spec becomes agent-callable MCP tools, with the same auth and quota policies applied, which helps teams avoid building a separate MCP server.

  8. Amp NewsAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

  9. eric zakariassonAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  10. GitHub Blog · AI & MLAI score46

    Copilot app rebuilds pull request view to render a 2,200-file diff smoothly

    AIGitHub rebuilt the pull request view in the GitHub Copilot app to keep review fast on very large diffs, testing it on an open source pull request with 2,200 files, over a million changed lines, and more than 400 inline review comments. The core difficulty is that review comment heights can only be measured at render time, which breaks the fixed-geometry virtualization used for code-only diffs. GitHub split the document height into a deterministic code domain and a separately measured domain for comment blocks.

  11. InferactAI score49

    Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

    AIInferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline. The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding. Inferact says it is open-sourcing the kernel today.

    Video from @inferact's post
  12. Greg BrockmanAI score67

    ChatGPT Voice gains plugins and ChatGPT Work access on web and mobile

    AIChatGPT Voice can now use plugins such as email, calendar, and Slack, and it can be powered by GPT-6 Astra, Sol, and Luna. It is also available in ChatGPT Work on web and mobile for creating docs, decks, sites, and spreadsheets by voice, rolling out globally in the latest app version.

    Why it matters: The quoted OpenAI post names the new voice tool access, supported models, and Work integration, showing how voice now acts across workflows.

  13. Google · AI blogAI score36

    Google Beam expands to six countries, adds Industrious flexible-workspace network

    AIGoogle is shipping Google Beam units to customers in the U.S., Canada, U.K., France, Germany, and Japan, supported by 18 channel partners, with HP Dimension with Google Beam as the flagship hardware. Starting in October, users can book Beam at select Industrious locations in Atlanta, Chicago, New York City, and Palo Alto, and an internal eight-week Google study reported 50% more connection and 21% fewer follow-up meetings.

  14. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  15. eric zakariassonAI score36

    Optimizing reading for AI agents cuts context-gathering costs

    AIEric Zakariasson argues that agents spend heavily on reading context before and after work, so optimizing that reading makes a major difference. He recommends the linked guide to builders, or handing it to an agent to implement its findings. Cursor's related post reports 7% lower token costs with no drop in agent quality, achieved through tighter prompts, selective tool loading, better caching, and compressed file reads.

    Image from @ericzakariasson's post