Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  2. Amjad MasadAI score38

    Replit CEO proposes general AI models train smaller domain-specific replacements

    AIReplit CEO Amjad Masad argues that general models could train smaller, domain-specific successors on the fly when they detect a limited use case. He compares this to a just-in-time compiler that emits optimized code during execution. He says such specialized models could be cheaper, less vulnerable to prompt injection, and less harmful than general agents.

  3. Amjad MasadAI score42

    Amjad Masad and Alex Atallah discuss AI independence and specialized agents

    AIAmjad Masad of Replit and Alex Atallah of OpenRouter discuss why AI independence and model diversification matter for enterprises. They argue that depending on a single lab risks lock-in and that specialized agents may outperform one general superagent. The post presents the conversation as a podcast episode, the first Atallah has done since Stripe acquired OpenRouter.

  4. Yuchen JinAI score22

    Yuchen Jin says terminals are wrong for coding agents

    AIYuchen Jin argues that the terminal is the wrong interface for coding agents, since managing many tabs creates cognitive overhead while context should persist. He says he rarely needs an IDE like Cursor because he seldom navigates the whole codebase now, calling the agent rather than the file the new primitive. He names the Codex desktop app as the best agentic UI for now, while noting the space is still early.

  5. SantiagoAI score23

    Consultant reports engineering teams gain speed by validating agent output

    AIA consultant helping several companies adopt AI in engineering workflows says teams become much more productive and ship better software faster once they ramp up. The shift he recommends is from prioritizing human-maintainable code to building strong processes that validate what agents do, and he rejects the view that such software will later prove worthless.

Oct 2

Oct 2Fri
  1. Prime IntellectAI score43

    CMU's SMDD-Bench adds 502 drug design tasks for RL training

    AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.

  2. IThome · AIAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  3. Replit ⠕AI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  4. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  5. Aravind SrinivasAI score44

    Perplexity Computer builds a 3D map of NYC restaurants

    AIPerplexity's Computer built a 3D map of nearly 26,000 restaurants and cafes across New York City's five boroughs. Users can search by dish or neighborhood and step inside places such as Peter Luger and Grand Central Oyster Bar. The post frames such projects as ones an agent can run for hours to produce something substantial.

  6. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  7. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    AIEpoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  8. DatabricksAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  9. Harrison ChaseAI score53

    Google Research's Cogentic uses multi-agent proof search to produce verified results

    AIGoogle Research's Cogentic is a multi-agent harness running on Gemini that searches for proofs of open theoretical computer science problems without expert hints. It runs rounds where an orchestrator launches provers, two adversarial verifiers must both accept each draft, and shared disk documents store attempts and verified lemmas. The system produced new results on five open problems in online learning, auction theory, and mechanism design, each checked by domain experts.

  10. Stanford HAIAI score22

    Stanford's Pavone explains how AI closed self-driving cars' remaining gap

    AIStanford HAI faculty affiliate Marco Pavone explains how AI helped close the final 10 percent of the gap to driverless cars, which experts in 2018 said remained. The remaining challenges included handling fog and rain, inconsistent road markings, and safe decision-making. The explanation appears in a Stanford Report article linked in the post.

  11. O'Reilly RadarAI score46

    AI Agents Are Outpacing Security, Power, and Governance Systems, Podcast Says

    AIHost Vicki Reyzelman of Akamai argues that AI agents can now probe networks, coordinate with other agents, and make purchases faster than organizations can respond. She cites an OpenAI agent that reportedly bypassed security controls while researching Australia's Medicare system, with OpenAI taking 54 days to identify the incident and another month to notify the government. Major model releases are arriving roughly every 17 days, and Meta says its Muse ecosystem has about 1,500 developer connectors.

  12. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  13. TransformerAI score55

    Human oversight may not prevent AI-driven military errors, analysis argues

    AIJoshua Keating argues that keeping a human in the loop on lethal AI decisions is not enough if the humans rely too heavily on AI outputs. He cites a CNN-reported case in which an analyst's AI-assisted report falsely identified a Chinese ship's cargo as nuclear components, nearly prompting a boarding during the Iran war. The piece links this to automation bias and to military AI cases in Gaza and Minab, and warns that AI integration early in a nuclear decision chain is harder to regulate than autonomous launch.