Skip to contentSkip to stories

Updated

#Deployment/Engineering

Oct 7

Oct 7Wed
  1. Georgi GerganovAI score44

    llama.cpp adds ggml RPC for distributing inference across heterogeneous devices

    AIllama.cpp can distribute inference across heterogeneous devices through the ggml RPC backend, according to Georgi Gerganov. He says it is currently an advanced setting, but he expects it to become more accessible to regular users over time. A related post reports MiMo 2.6 Flash running across an RTX 6000 GPU and an M5 laptop over 10 GbE at about 40 tokens/sec.

  2. GitHub Blog · AI & MLAI score57

    GitHub argues secret protection must scale with AI-driven code growth

    AIGitHub reports that one in three pull requests now involves an AI agent, and that public secret exposures rise with the volume of pushes rather than from declining developer care. It introduces a ModernBERT-based classifier with Microsoft Applied Sciences that evaluates candidate secrets in under two milliseconds and could more than double the secrets prevented at push time. The feature is in private preview, with availability for GitHub Secret Protection customers later this month.

  3. Gergely OroszAI score31

    Samuel Newman on why LLMs aren't world models and lack causality

    AISam Newman argues the tech world misunderstands LLMs because they have no concept of causality, so "if I do A, B happens" reasoning is absent. He contends LLMs are not world models, unlike older world-model approaches that could in principle track cause and effect. He adds that people overestimate LLM capabilities because they seem smart, and that guardrails are unlikely to be the right long-term fix.

  4. AMDAI score22

    Agentic AI workloads are about 80% CPU-bound, AMD and mimik find

    AIRecent mimik tests of agentic workflows on AMD Ryzen AI Embedded X100 processors found about 80% of operations were CPU-bound, covering coordination, orchestration, scheduling and reporting. The post argues that CPUs play a major role in agentic AI rather than GPUs alone, and that heterogeneous compute matters for deploying it at the edge. A full interview with mimik founder and CEO Fayarjomandi is linked.

  5. Elvis SaraviaAI score22

    Elvis Saravia describes building personal multi-agent teams with Opus 5.5

    AIElvis Saravia reports that agent-to-agent communication with a personal agent, built on models like Opus 5.5, is already coordinating work faster and at higher quality than he can match. He describes progressing from individual Claude Code sessions to subagents, then a persistent team of eight specialized bots with his own orchestrator. He argues everyone should build a personalized agent orchestrator and says most apps like Code and Claude Desktop are behind.

  6. GoogleAI score42

    Google's Project Suncatcher tests TPUs in orbit on a satellite

    AIGoogle launched its first test satellite carrying four TPUs into orbit last week as part of Project Suncatcher, a moonshot exploring whether machine learning infrastructure could one day operate in space. The test aims to determine whether Google's AI hardware can withstand the physical stress of spaceflight and the radiation and thermal extremes of orbit.

  7. Google Cloud TechAI score34

    Antigravity agent plugins weigh eager versus lazy loading of MCP tools

    AIGoogle DevRel's James O'Reilly compares two ways Antigravity exposes local MCP tools from Agent Plugins to the model. Eager loading registers each tool as a top-level function with its full schema in every turn's system prompt, which speeds calls but consumes fixed tokens. Lazy loading, the plugin default, exposes tools through a proxy call_mcp_tool and reads schemas on demand, saving baseline context at the cost of an extra discovery step.

  8. Databricks BlogAI score41

    Databricks Apps Adds On-Behalf-of-User Authorization for Permission-Aware Apps

    AIDatabricks announced general availability of on-behalf-of-user (OBO) authorization for Databricks Apps, letting apps act with the signed-in user's identity so Unity Catalog enforces that user's row filters and column masks. Developers can request narrow API scopes such as sql:restricted-query, which allows only read-only SQL queries, while apps keep a dedicated service principal for app-owned operations.

  9. Unsloth AIAI score40

    Unsloth lets users train local decision models on 4GB VRAM

    AIUnsloth released an open-source method to fine-tune LLMs into decision models that run locally, lifting Qwen3.5 0.8B's aggregate accuracy from 20.7% to 74.3% across three decision benchmarks. The team used a Clef head with LoRA (r=64) for one epoch on just 4GB VRAM, with the approach applicable to models such as Qwen3.8 and Gemma 4. A guide and notebooks are available on the Unsloth documentation site and GitHub.

  10. SantiagoAI score22

    Model infers derived values from document data, computing yearly costs from monthly figures

    AIA new model extracts values absent from a document by computing them from figures that are present, such as deriving a yearly product cost from a monthly price. Santiago says the video shows examples of inferring complex formulas. The background post describes this as Higher-Order Extraction, which deterministically computes needed numbers from raw page values.

  11. GitHub Copilot ChangelogAI score30

    GitHub launches purpose-built AI model for leaked secret detection across developer workflows

    AIGitHub is rolling out a fine-tuned, purpose-built model for secret detection that reads surrounding code to identify likely credentials, including passwords without recognizable token formats. Existing AI-detected Password alerts have been upgraded automatically, and AI-detected secrets in push protection is in private preview. New opt-in checks in push protection and the GitHub Copilot /security-review command will consume GitHub AI Credits.

  12. LlamaIndexAI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

  13. Microsoft ResearchAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    AIMicrosoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    Why it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

  14. NVIDIA Technical BlogAI score22

    Validate AI Factory Changes with Digital Twins and AI Agents

    AINVIDIA describes using digital twins and AI agents to validate changes to AI factory infrastructure, which combines GPUs, CPUs, switches, DPUs, and SuperNICs with schedulers, orchestration services, security controls, and a fast-changing software stack. The source frames the challenge as confirming that hardware, software, and policies work together for target workloads before deployment. The available excerpt does not give further detail on specific tools or results.

  15. AWS Machine Learning BlogAI score44

    Qlik Builds Grounded Enterprise AI Answers Using Amazon Bedrock

    AIQlik built Qlik Answers, a natural-language assistant that returns sourced answers from knowledge bases, analytics apps, glossaries, and documents, using Amazon Bedrock for model access. The system routes each question through specialist agents and retrieval on Amazon OpenSearch Service, with Amazon Bedrock Guardrails applied to every request and response. Qlik serves more than 40,000 customers across regions, using Amazon SageMaker AI as an in-Region fallback when models are not yet available on Bedrock.

  16. AWS Machine Learning BlogAI score53

    Automate remediation after AWS DevOps Agent investigations with Lambda and Bedrock

    AIThe AWS Machine Learning Blog describes an automated remediation workflow that acts on AWS DevOps Agent investigation results. Amazon EventBridge triggers a Lambda durable function that uses Amazon Bedrock to propose fixes from an allowlist of tools, running read-only actions autonomously and pausing for human approval before infrastructure changes. The post demonstrates the flow with a Lambda function whose 3-second timeout is raised to 30 seconds after a single approval.

  17. GitHub Copilot ChangelogAI score58

    GitHub Copilot local sandboxing now generally available across CLI, app, and VS Code

    AIGitHub has made local sandboxing for GitHub Copilot generally available in GitHub Copilot CLI, the GitHub Copilot app, and VS Code sessions using Agent Host. Sandboxes restrict the filesystem, network, and credentials that Copilot-initiated tools and commands can access, based on developer or organization policies. The feature is powered by Microsoft eXecution Container (MXC), supports Windows, macOS, and Linux, and is included at no additional cost.

  18. GitHub Copilot ChangelogAI score42

    GitHub Copilot CLI adds discovery of local Ollama models via /model

    AIGitHub Copilot CLI version 1.0.94-0 lets users run /model to discover supported models from a running local Ollama instance alongside configured and GitHub Copilot cloud models. Discovered models are not added automatically; users choose one, review its provider and endpoint, then confirm Add and use for this session or Add without switching, and models must support tool calling and streaming. Choosing a local model does not enable offline mode or disable GitHub telemetry, and COPILOT_OFFLINE=true remains a separate explicit setting.

  19. AWS Machine Learning BlogAI score32

    AWS playbook: six-week program closes AI builder gap for non-engineers

    AIAWS ran a six-week program pairing non-engineering professionals with mentors and tools like Amazon Bedrock AgentCore and the Strands Agents SDK to build working AI prototypes. Four participants with no engineering background built WealthWise, a multi-agent financial advisory tool with five agents on Amazon Nova models, which won first place. The article says participants who completed the phased program retained three times more practical skills than those in two-day intensive formats.

  20. Elvis SaraviaAI score18

    Viktor, a Slack AI employee, reviews overnight agent eval failures

    AIElvis Saravia describes using Viktor, an AI employee in Slack, to review his nightly agent harness evaluation results. Viktor traces tasks that regressed from passing to failing back to the specific harness change that caused them and suggests reverting it, while the human makes the final decision. The post is a sponsored partnership, offering $100 in free credits with no card required.

  21. WaymoAI score27

    Waymo releases framework for AV incident-management exercises and drills

    AIWaymo has introduced a first-of-its-kind framework for autonomous vehicle incident-management exercises, ranging from tabletop scenarios to full-scale drills. Adapted from emergency management best practices, it is designed to help AV developers, operational partners, and first responders test plans and strengthen coordination together.

  22. Elvis SaraviaAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    AINVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.