Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Aug 28

Aug 28Fri
  1. Meituan LongCatAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. RadixArkAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

    Image from @radixark's post
  2. Anthropic · YouTubeAI score43

    Anthropic Unveils Model Hardware Standard for AI Agents Operating Physical Equipment

    AIAnthropic is introducing the Model Hardware Standard (MHS), a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. MHS began as part of a beneficial deployments project with HHMI Janelia Research Campus and is evolving into a wider industry effort. It is now in research preview with select partners.

  3. Augment Code BlogAI score50

    Augment Code launches Cosmos Advisor, an agent that configures its own platform

    AIAugment Code introduces Cosmos Advisor, an expert that can answer product questions, configure agents, and deploy automations from a single conversation. The company says a company-specific agent can be set up in about ten minutes, without a handoff to an implementation team. Advisor draws on the current Cosmos knowledgebase and reusable expert designs, such as incident response, and it works within Object-Level Access Control.

  4. Augment Code BlogAI score38

    Augment Code's two-engineer team uses a Feedback Triager agent to handle surging product feedback

    AIAugment Code's two-engineer Cosmos Advisor team built a Feedback Triager agent to handle product feedback that grew to about 30 threads per week, which had consumed an estimated 90% of team time. The agent investigates each Slack report through root-cause analysis, answers questions, routes issues to other teams, files tickets, and hands clear fixes to a PR Author agent. Humans retain prioritization and product decisions.

  5. Anthropic · YouTubeAI score62

    Anthropic and HHMI Janelia launch Model Hardware Standard for AI lab equipment

    AIAnthropic is building the Model Hardware Standard (MHS), a common way for AI models to connect to lab and manufacturing equipment and operate it with safety limits built into each device. MHS started as a collaboration between Anthropic and HHMI Janelia Research Campus and is launching as a research preview with partners across science, robotics, and manufacturing.

    Why it matters: The source describes a standard for connecting AI models to lab and manufacturing hardware, which matters for anyone building automated experimentation workflows.

  6. TinkerAI score41

    alphaXiv turns research papers into live experiments run by agents on Tinker

    AIalphaXiv is turning research papers from static artifacts into live research that grows and branches, with agents running their own experiments. Tinker says it makes running these experiments easy for both agents and people. Via alphaXiv's background post, its autoresearch tool lets Claude or Codex agents replicate and experiment on any arXiv paper, with agents launching concurrent RL runs through Tinker for post-training.

  7. Anthropic · YouTubeAI score58

    Anthropic's Model Hardware Standard lets AI agents operate physical lab equipment

    AIAnthropic and HHMI Janelia Research Campus developed the Model Hardware Standard (MHS), a standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. MHS is now in research preview with select partners, and the video describes how it was developed and how it can accelerate research.

  8. LMSYS OrgAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

    Image from @lmsysorg's post
  9. Unsloth AIAI score70

    GLM-5.3-Flash can run locally with Unsloth GGUF quantization on 128GB RAM

    AIUnsloth says GLM-5.3-Flash can run locally, with a 3-bit GGUF version needing 128GB of RAM and the 1-bit version working on 102GB of RAM or VRAM. The guide's table lists memory needs from 100GB at 1-bit to 650GB at BF16, and reports that the 1-bit quant keeps 71% of top-1% accuracy while being 85% smaller than BF16.

    Why it matters: The guide gives concrete memory requirements for each quantization level, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  10. Leandro von WerraAI score22

    Pollen Robotics unveils Microduck, a $400 open-source RL biped robot

    AIPollen Robotics has unveiled Microduck, a 25 cm open-source biped with 15 actuators and sensors including a camera, speaker, and LiDAR that users can train with reinforcement learning. The robot ships with more than half a dozen pre-trained policies for walking, sitting, roller-skating, and picking up objects with its articulated beak, and costs less than $400.

Aug 26

Aug 26Wed
  1. Cursor ChangelogAI score46

    Cursor Cloud Agents now let you start projects from scratch without a repo

    AICursor Cloud Agents no longer require a connected GitHub or other third-party SCM provider to begin work. Users select "Start from scratch" in the repo picker, and Cursor creates an Origin repo in the background that can be saved as a private or internal repo via "Create repo." Cursor also now port-forwards the cloud agent's live environment to the browser for previews, and a connected Vercel account lets users publish a live URL.

  2. Jazzyear · ArticlesAI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    AIIn an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.

  3. Google Developers BlogAI score42

    Google Developers Blog explains deep learning with Keras for astroparticle physics data analysis

    AIThe Google Developers Blog post describes how deep learning can analyze the large, image-like sensor data from astroparticle observatories such as the Pierre Auger Observatory and IceCube. The author argues these methods could improve instrument sensitivity and reveal patterns in cosmic-ray and neutrino signals that traditional analysis techniques miss.

  4. Bryan CatanzaroAI score62

    NVIDIA and AWS expand partnership with 2 million more GPUs and Vera CPU for agentic AI

    AINVIDIA and AWS are expanding their partnership across GPUs, CPUs, networking, open models and software. The announcement cites 2 million additional NVIDIA GPUs across AWS infrastructure, the NVIDIA Vera CPU coming to AWS for agentic AI, NVLink Fusion with NVHBM memory, and 100,000 GPUs for U.S. government AI factories on secure AWS infrastructure.

  5. Replit BlogAI score38

    Replit Launches Intelligent Model Routing, Auto-Matching Each Task to Suitable AI Model

    AIReplit has made Intelligent Model Routing available to all users, automatically matching each task with a model based on quality, speed, and cost. In the company's testing, the feature delivered the same output quality at 65% lower cost than the previous version of Max Mode. Enterprise administrators can restrict routing to a company-approved set of models.

  6. LMSYS OrgAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

  7. LMSYS OrgAI score60

    SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

    AISGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

  8. Microsoft AI BlogAI score19

    Microsoft Shows How AI Is Reshaping Customer Engagement Across Industries

    AIMicrosoft's Accelerating Frontier Transformation series says AI is helping organizations deliver more personalized engagement at scale and give staff time back for relationships. Examples include Lifeline Australia using AI for service insight, Brisbane Catholic Education personalizing curriculum for students with Copilot, and Uniting NSW.ACT's Buddy platform cutting some frontline tasks from 10 to 15 minutes to one to two.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Google Developers BlogAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    AIGoogle Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.