Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Aug 21

Aug 21Fri
  1. Andrew NgXAI score31

    Andrew Ng outlines six core skills for building and deploying AI applications

    AIAndrew Ng's AI Engineering Skills Map ranks building and deploying AI applications as the top skill tier, spanning LLM foundations, data grounding, agentic systems, evaluation-driven development, production operations, and machine learning foundations. He explains that AI outputs are less predictable than traditional software, so skilled engineers build iteratively, examining results and deciding next steps based on intermediate outcomes. The skills map was derived from job postings, expert interviews, and survey responses.

  2. Amazon ScienceOfficialAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    AIAmazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

  3. DeepSeekOfficialAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

    Image from @deepseek_ai's post
  4. DeepSeek API NewsOfficialAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu
  1. swyxXAI score34

    Matt Pocock's /wayfinder skill navigates unclear projects with research and grilling

    AIMatt Pocock's /wayfinder skill is designed for "fog of war" situations where the end state of a project is unclear. It orchestrates research and other grill sessions to help users discover what they don't yet know, building on his popular /grill-me skill. Latent Space is featuring the skill in an exclusive interview as the first in a series of Skills coverage.

    Image from @swyx's post
  2. Mistral AIOfficialAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Jazzyear · InsightsNewsAI score29

    Jazzyear's 2026 tech investment conference maps where capital is flowing in AI and hard tech

    AIAt the 2026 Jiazi Gravity Tech Industry Investment Conference in Beijing, Jiazi Guangnian's CEO Zhang Yijia said first-half 2026 saw investment amounts rise 91.6% year on year, IPOs rise 39.2%, and M&A transaction value double. The report said AI absorbed over 70% of global venture investment, with OpenAI and Anthropic together raising $217 billion, roughly 40%.

  2. Matei ZahariaXAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  3. JetBrains AI BlogOfficialAI score31

    Air Adds Multiproject View, Markdown Rendering, and Windows IME Fixes

    AIAir's latest release lets users open several projects in one window and run agents across them in parallel, with tasks grouped by project in the sidebar. Markdown files now render as formatted text while editing, with syntax shown only when editing, and Chinese, Japanese, Korean, and other IMEs now work on Windows. The release also adds a Customize screen for keymap, theme, and accent color, and lets users choose the agent and model for Agent Review.

  4. TinkerOfficialAI score38

    Qwen3.8-27B is now available on Tinker

    AITinker has made Qwen3.8-27B available today. The model is natively multimodal, handling images and video, with flexible thinking control. Tinker says it performs meaningfully better at coding, professional work, research, and long-horizon agentic tasks.

  5. Kimi.aiOfficialAI score24

    Kimi Work tutorial shows financial analysts three research workflows

    AIKimi publishes Tutorial #2 for its Kimi Work product, showing financial analysts how to use it for three investment research tasks. The tasks are building a live investor dashboard, updating financial models in spreadsheets, and processing and generating reports in batch. The post promises more Kimi Work workflows to follow.

    Video from @Kimi_Moonshot's post

Aug 18

Aug 18Tue
  1. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.

  2. Google LabsOfficialAI score43

    Google's CC Gmail agent expands waitlist to Australia and New Zealand

    AIGoogle Labs has opened a waitlist for CC, its experimental AI productivity agent in Gmail, in Australia and New Zealand, and is expanding availability in the US and Canada. CC now helps manage calendars by connecting to Gmail and automatically creating events in a dedicated Google Calendar that update as plans change. Invitations to waitlisted users in the US and Canada begin rolling out today.

  3. VercelOfficialAI score42

    Vercel launches $1M hacker challenge to test Sandbox security

    AIVercel is offering up to $1,000,000 in a public hacker challenge testing its Vercel Sandbox against escapes from the Firecracker microVM and bypasses of the host-side network boundary. Rewards reach $50,000 per report, administered through HackerOne (@Hacker0x01). The company says agents can now exploit vulnerable sandbox boundaries, so it is testing its own defenses in the open.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesOfficialAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

  2. Chip HuyenXAI score22

    Chip Huyen asks for a model tiering system for agent orchestration

    AIChip Huyen asks what a good model tiering system looks like, since she is tired of naming specific models per vendor for her agent orchestrator. She wants to instruct the orchestrator by task tier, such as "use models tier ..." for a given kind of task, instead of listing Claude, OpenAI, and other models individually.

  3. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry adds five Claude agent features to Azure-hosted deployments

    AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.

    Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.

  4. Jason WeiXAI score45

    Jason Wei argues tool use cannot replace larger language models

    AIJason Wei now believes a small 1B-parameter "cognitive core" relying on tools is insufficient, because fast, natural recall without tool use matters. He cites speed, knowledge better learned through backpropagation than retrieved from search, and the greater reliability of already-known facts over repeated lookups. Since a 1B model has an information limit, he argues that demanding AI will still need larger models, not just tool access.

  5. Replit BlogOfficialAI score60

    Replit adds black-box pen tests that probe apps like external attackers

    AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.

    Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.

  6. Import AIBlogAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 16

Aug 16Sun
  1. Philipp SchmidBlogAI score58

    Controlling Android with Gemini 3.7 Flash and 150 lines of Python

    AIThe author built a Python agent that uses Gemini 3.7 Flash to control an Android emulator from raw screenshots, returning normalized 0–999 coordinates that are scaled to 1080x1920 pixels over ADB. In a test, the agent opened Chrome, closed popups, and solved one round of Wordle in two guesses without accessibility IDs or DOM access. The article presents the loop as usable for UI testing and task automation across native apps, webviews, and canvas interfaces, with code in an open-source quickstart repository.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogOfficialAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Aug 14

Aug 14Fri
  1. Augment Code BlogOfficialAI score62

    Augment rebuilds its Auggie CLI harness on Pi, cutting SWE-bench Pro task cost 53%

    AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.

    Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.

  2. Epoch AI · The Epoch BriefOfficialAI score42

    Epoch AI lists nine big AI questions its benchmarks aim to answer

    AIEpoch AI outlines nine open questions about AI capabilities, including whether AI can take over full jobs and whether benchmark scores are correlated. The author says Epoch's benchmarking work is built to help answer them, citing examples such as MirrorCode, Remote Labor Index, and the Epoch Capabilities Index (ECI). The post notes that benchmark scores are highly correlated across domains, and that ECI growth trends can help detect whether AI capability progress has accelerated.

  3. Andrew NgXAI score38

    Andrew Ng maps the four key skills for AI engineering

    AIAndrew Ng's team released an AI Engineering Skills Map, built from analysis of over 10,000 job postings and expert interviews, identifying four priority skills. The skills are building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. Ng says these skills matter for all developers, not only those with the AI Engineer title.

  4. Ali GhodsiXAI score46

    Databricks Smart Routing cuts AI coding task costs about 30% in Unity Gateway

    AIAli Ghodsi says Smart Routing on Databricks' AI Gateway lowers costs by about 30% without sacrificing quality. The quoted Databricks post says it matches each coding task to the right model and harness based on task needs, so higher-cost models focus on intelligence while lower-cost models compete on cost and performance.