Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. AWS Machine Learning BlogAI score53

    Automate remediation after AWS DevOps Agent investigations with Lambda and Bedrock

    AIThe AWS Machine Learning Blog describes an automated remediation workflow that acts on AWS DevOps Agent investigation results. Amazon EventBridge triggers a Lambda durable function that uses Amazon Bedrock to propose fixes from an allowlist of tools, running read-only actions autonomously and pausing for human approval before infrastructure changes. The post demonstrates the flow with a Lambda function whose 3-second timeout is raised to 30 seconds after a single approval.

  2. Lucas Beyer (bl16)AI score36

    Reality Check: a public leaderboard for robot manipulation VLA models

    AILucas Beyer praises Reality Check, a new leaderboard for benchmarking VLA and related robot manipulation models. Half of its tasks are fully open, while the other half are held out to detect benchmaxxing by future model versions. The companion post from Nicolas Keller describes the launch as the first public robot manipulation benchmark, built on 14,400 real-world rollouts across four models.

  3. GitHub Copilot ChangelogAI score58

    GitHub Copilot local sandboxing now generally available across CLI, app, and VS Code

    AIGitHub has made local sandboxing for GitHub Copilot generally available in GitHub Copilot CLI, the GitHub Copilot app, and VS Code sessions using Agent Host. Sandboxes restrict the filesystem, network, and credentials that Copilot-initiated tools and commands can access, based on developer or organization policies. The feature is powered by Microsoft eXecution Container (MXC), supports Windows, macOS, and Linux, and is included at no additional cost.

  4. GitHub Copilot ChangelogAI score42

    GitHub Copilot CLI adds discovery of local Ollama models via /model

    AIGitHub Copilot CLI version 1.0.94-0 lets users run /model to discover supported models from a running local Ollama instance alongside configured and GitHub Copilot cloud models. Discovered models are not added automatically; users choose one, review its provider and endpoint, then confirm Add and use for this session or Add without switching, and models must support tool calling and streaming. Choosing a local model does not enable offline mode or disable GitHub telemetry, and COPILOT_OFFLINE=true remains a separate explicit setting.

  5. AWS Machine Learning BlogAI score32

    AWS playbook: six-week program closes AI builder gap for non-engineers

    AIAWS ran a six-week program pairing non-engineering professionals with mentors and tools like Amazon Bedrock AgentCore and the Strands Agents SDK to build working AI prototypes. Four participants with no engineering background built WealthWise, a multi-agent financial advisory tool with five agents on Amazon Nova models, which won first place. The article says participants who completed the phased program retained three times more practical skills than those in two-day intensive formats.

  6. Wired · AIAI score46

    Pentagon's Tradewinds Program Uses Five-Minute Videos to Speed AI Purchases

    AIThe Department of Defense's Chief Digital and Artificial Intelligence Office runs the Tradewinds program, which grants "post-competitive" status to AI vendors based on videos of five minutes or less. That status can let government buyers use other transaction agreements and sometimes make awards in less than a week. OpenAI, Anthropic, and Google are listed as participants, though none commented.

  7. indigoAI score34

    Grok Bot acts as a model router, using Gemini and Opus together

    AIThe poster says they already use Grok Bot as a model router, citing last weekend's personal agent livestream. In the demo, Gemini produced an infographic inside Grok Bot, and Claude Opus then checked the content. This follows Elon Musk's announcement that Grok Bot will use the best backend model for each task, including Claude Opus 5.5, MidJourney, and Suno.

    Video from @indigox's post
  8. WaymoAI score27

    Waymo releases framework for AV incident-management exercises and drills

    AIWaymo has introduced a first-of-its-kind framework for autonomous vehicle incident-management exercises, ranging from tabletop scenarios to full-scale drills. Adapted from emergency management best practices, it is designed to help AV developers, operational partners, and first responders test plans and strengthen coordination together.

    Image from @Waymo's post
  9. 🚨 AI News | TestingCatalogAI score41

    Google releases Foresight macOS app using Gemma 4 for voice notes

    AIGoogle released the Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma 2. The app can connect to Google Drive to build a knowledge graph and, when transcription is active, uses local Gemma 4 E4B or Gemma 4 12B models to transcribe voice notes into new documents. EmbeddingGemma 2 is an open-weight, Apache 2 licensed 740M-parameter multimodal embedding model with an 8K context window.

    Video from @testingcatalog's post
  10. elvisAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    AINVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

    Image from @omarsar0's post
  11. Google GemmaAI score46

    Gemma 4 E4B helps find climate-resilient crop mutations faster

    AIAI lab Living Models pairs Gemma 4 E4B with BOTANIC-1, a genomic language model, to speed up identifying DNA that makes crops climate-resilient. Gemma prepares genomic data and filters candidates, while BOTANIC-1 scores evolutionary impact to pinpoint the causal mutation. In a recent test, the system ranked a target melon yield mutation first out of 2,494 possibilities after an afternoon of computation.

    Video from @googlegemma's post
  12. Azure BlogAI score34

    Microsoft Uses AI Agents to Speed Azure Cloud Infrastructure Supply Chain Planning

    AIMicrosoft's Azure Hardware Systems and Infrastructure team is applying AI agents across its infrastructure lifecycle, starting with cloud supply chain demand planning. Following a "Lean before AI" approach, the company reports that multi-agent workflows cut planning work that took five to seven business days to hours, with approximately 50% less manual effort and cycle time down up to 75% in selected workflows. Microsoft says it is extending the approach to fulfillment, logistics, and fleet operations while keeping human judgment central.

  13. NVIDIA Technical BlogAI score25

    NVIDIA cuPhoton Speeds Up Scientific Image Analysis for High-Throughput Instruments

    AINVIDIA's cuPhoton targets the computational bottleneck in scientific image pipelines, where data from observatories, telescopes, lasers, and X-ray light sources arrives faster than CPU-bound processing can handle. The source says the bottleneck is usually the whole path from raw sensor data to decision, not one slow kernel. The available text does not give benchmark figures, pricing, or availability details.

  14. 404 MediaAI score44

    Arizona court orders resentencing after AI video of victim swayed judge

    AIAn Arizona appellate court ruled that an AI-generated video of manslaughter victim Christopher Pelkey, which his sister Stacey Wales played at sentencing, carried "undue emotional weight" and ordered the judge to reconsider the 10.5-year prison term. The conviction stands, but the sentence must be revisited. Wales said her goal was to sway the judge with the video.

  15. Google AIAI score54

    Google opens public SynthID portal for checking AI-generated images, video and audio

    AIGoogle is letting anyone check files for SynthID watermarks at synthid.com, covering content from Google and partners including OpenAI, NVIDIA and Kakao. Apple is listed as coming soon. Google says it has watermarked 180 billion images and videos and more than 240,000 years of audio, and the portal handles about 1 million verification requests daily.

    Image from @GoogleAI's post
  16. Ars Technica · AIAI score63

    Mistral releases Le Chonk, a 1 trillion-parameter open-weight model

    AIMistral has released Mistral Large 4, nicknamed Le Chonk, a 1 trillion-parameter model it says can be used and customized by anyone. It is in preview, with a final version due by the end of the month, and is optimized for coding and cyberdefense as well as manufacturing, finance, and electrical engineering tasks. Mistral claims it is the most capable open-weight model developed outside China and says it was trained from scratch rather than through distillation.

  17. GoogleAI score46

    Google's SynthID has watermarked over 180 billion images and videos

    AIGoogle says it has watermarked more than 180 billion images and videos, plus 240,000 years of audio, since launching SynthID in 2023. The verification feature is built into Search, the Gemini app, and Chrome, which together handle over 1 million verification requests daily. Google presents the SynthID Detector platform as part of its effort to give users more context about online media.

  18. Latent SpaceAI score61

    Stacklok's Mecatl harness moves coding agents from desktops to the cloud

    AIStacklok, founded by Kubernetes creators Craig McLuckie and Joe Beda, has released Mecatl, an open source cloud-native harness for coding agents on GitHub. Mecatl keeps the agent loop separate from the client, model provider, state store, and execution environment, and moves tool calling, session management, and memory into manageable systems. The article also covers ToolHive, an MCP platform, and an AI Gateway that is not yet open sourced, with a commercial enterprise control plane tying the pieces together.