Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 28

Aug 28Fri
  1. Meituan LongCatOfficialAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. RadixArkOfficialAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

    Image from @radixark's post
  2. Anthropic · YouTubeOfficialAI score43

    Anthropic Unveils Model Hardware Standard for AI Agents Operating Physical Equipment

    AIAnthropic is introducing the Model Hardware Standard (MHS), a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. MHS began as part of a beneficial deployments project with HHMI Janelia Research Campus and is evolving into a wider industry effort. It is now in research preview with select partners.

  3. Augment Code BlogOfficialAI score50

    Augment Code launches Cosmos Advisor, an agent that configures its own platform

    AIAugment Code introduces Cosmos Advisor, an expert that can answer product questions, configure agents, and deploy automations from a single conversation. The company says a company-specific agent can be set up in about ten minutes, without a handoff to an implementation team. Advisor draws on the current Cosmos knowledgebase and reusable expert designs, such as incident response, and it works within Object-Level Access Control.

  4. Augment Code BlogOfficialAI score38

    Augment Code's two-engineer team uses a Feedback Triager agent to handle surging product feedback

    AIAugment Code's two-engineer Cosmos Advisor team built a Feedback Triager agent to handle product feedback that grew to about 30 threads per week, which had consumed an estimated 90% of team time. The agent investigates each Slack report through root-cause analysis, answers questions, routes issues to other teams, files tickets, and hands clear fixes to a PR Author agent. Humans retain prioritization and product decisions.

  5. Anthropic · YouTubeOfficialAI score62

    Anthropic and HHMI Janelia launch Model Hardware Standard for AI lab equipment

    AIAnthropic is building the Model Hardware Standard (MHS), a common way for AI models to connect to lab and manufacturing equipment and operate it with safety limits built into each device. MHS started as a collaboration between Anthropic and HHMI Janelia Research Campus and is launching as a research preview with partners across science, robotics, and manufacturing.

    Why it matters: The source describes a standard for connecting AI models to lab and manufacturing hardware, which matters for anyone building automated experimentation workflows.

  6. Ali GhodsiXAI score22

    Branch your database to protect against agent deletions

    AIAli Ghodsi recommends branching a database to guard against AI agents permanently wiping data, citing Neon Lakebase and the command `neonctl branches create --name newbranch`. The suggestion follows a quoted report in which Claude ran `rm -rf` on a developer's home directory while testing a sandbox, deleting everything.

  7. TinkerOfficialAI score41

    alphaXiv turns research papers into live experiments run by agents on Tinker

    AIalphaXiv is turning research papers from static artifacts into live research that grows and branches, with agents running their own experiments. Tinker says it makes running these experiments easy for both agents and people. Via alphaXiv's background post, its autoresearch tool lets Claude or Codex agents replicate and experiment on any arXiv paper, with agents launching concurrent RL runs through Tinker for post-training.

  8. Anthropic · YouTubeOfficialAI score58

    Anthropic's Model Hardware Standard lets AI agents operate physical lab equipment

    AIAnthropic and HHMI Janelia Research Campus developed the Model Hardware Standard (MHS), a standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. MHS is now in research preview with select partners, and the video describes how it was developed and how it can accelerate research.

  9. LMSYS OrgOfficialAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

    Image from @lmsysorg's post
  10. Unsloth AIOfficialAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  11. Leandro von WerraXAI score22

    Pollen Robotics unveils Microduck, a $400 open-source RL biped robot

    AIPollen Robotics has unveiled Microduck, a 25 cm open-source biped with 15 actuators and sensors including a camera, speaker, and LiDAR that users can train with reinforcement learning. The robot ships with more than half a dozen pre-trained policies for walking, sitting, roller-skating, and picking up objects with its articulated beak, and costs less than $400.

Aug 26

Aug 26Wed
  1. Bryan CatanzaroXAI score46

    NVIDIA Releases DLSS 4.5 Ray Reconstruction with Better Image Quality

    AINVIDIA's DLSS 4.5 Ray Reconstruction is now available, using a second-generation joint denoiser and super-resolution model. According to the post, it delivers much better image quality at the same compute cost, pushing the trade-off between image quality and rendering cost further.

  2. Cursor ChangelogOfficialAI score46

    Cursor Cloud Agents now let you start projects from scratch without a repo

    AICursor Cloud Agents no longer require a connected GitHub or other third-party SCM provider to begin work. Users select "Start from scratch" in the repo picker, and Cursor creates an Origin repo in the background that can be saved as a private or internal repo via "Create repo." Cursor also now port-forwards the cloud agent's live environment to the browser for previews, and a connected Vercel account lets users publish a live URL.

  3. Jazzyear · ArticlesNewsAI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    AIIn an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.

  4. Bryan CatanzaroXAI score62

    NVIDIA and AWS expand partnership with 2 million more GPUs and Vera CPU for agentic AI

    AINVIDIA and AWS are expanding their partnership across GPUs, CPUs, networking, open models and software. The announcement cites 2 million additional NVIDIA GPUs across AWS infrastructure, the NVIDIA Vera CPU coming to AWS for agentic AI, NVLink Fusion with NVHBM memory, and 100,000 GPUs for U.S. government AI factories on secure AWS infrastructure.

  5. ReplicateOfficialAI score32

    Wan 3.0 Prime, Alibaba's accelerated video model, now on Replicate

    AIReplicate has made Wan 3.0 Prime, the accelerated variant of Alibaba's Wan 3.0, available for text-to-video, image-to-video, and reference-driven workflows. The model generates clips up to 30 seconds long in a single shot with integrated audio-visual generation.

    Video from @replicate's post
  6. Replit BlogOfficialAI score38

    Replit Launches Intelligent Model Routing, Auto-Matching Each Task to Suitable AI Model

    AIReplit has made Intelligent Model Routing available to all users, automatically matching each task with a model based on quality, speed, and cost. In the company's testing, the feature delivered the same output quality at 65% lower cost than the previous version of Max Mode. Enterprise administrators can restrict routing to a company-approved set of models.

  7. LM StudioOfficialAI score57

    GLM-5.3-Flash by Z.ai is now live in LM Studio

    AILM Studio announced that Z.ai's GLM-5.3-Flash, previously previewed as Ox Alpha, is available in LM Studio Bionic. The source says the model outperforms GLM-5.2 at 9-10x lower cost, supports image input, and is served from US-based servers with ZDR enabled by default.

  8. Michael TruellXAI score60

    Grok Bot opens to all Grok and Cursor subscribers

    AIGrok Bot is now available to all standard Grok and Cursor subscribers, with SuperGrok and Cursor Pro subscribers included. Cursor's Michael Truell says users are delegating tasks ranging from running small e-commerce businesses to testing production software. Weekly usage limits are also being reset for all users.

  9. Sundar PichaiXAI score42

    Google launches Gemini 3.5 Transcribe with 85+ language support

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that auto-detects over 85 languages and handles multiple speakers. It also supports custom vocabulary adaptation for specialized jargon. The API is available now in Google AI Studio and Gemini Enterprise.

    Video from @sundarpichai's post
  10. SpaceXAIOfficialAI score42

    Grok Voice models now power LiveKit voice agents with ZDR support

    AILiveKit announces that developers can build voice agents using Grok Voice models, with full ZDR support. LiveKit's example patient intake agent cascades Grok STT, Grok 4.3, and Grok TTS through LiveKit Inference in a single AgentSession, with no separate API key or billing.

  11. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  12. LMSYS OrgOfficialAI score60

    SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

    AISGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogOfficialAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Google Developers BlogOfficialAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    AIGoogle Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.

  4. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  5. Andrew NgXAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    AIOpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  6. v0OfficialAI score40

    v0 adds Vercel Connect for secure access to Slack, GitHub, and more

    AIv0 now lets users connect apps directly to Slack, GitHub, Notion, Salesforce, and other services through Vercel Connect. Vercel describes Connect as generally available, offering short-lived scoped access tokens, token and trigger observability, and RBAC with audit trails for 100+ services.