Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Aug 19

Aug 19Wed
  1. Kimi.aiOfficialAI score24

    Kimi Work tutorial shows financial analysts three research workflows

    AIKimi publishes Tutorial #2 for its Kimi Work product, showing financial analysts how to use it for three investment research tasks. The tasks are building a live investor dashboard, updating financial models in spreadsheets, and processing and generating reports in batch. The post promises more Kimi Work workflows to follow.

    Video from @Kimi_Moonshot's post

Aug 18

Aug 18Tue
  1. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.

  2. Google LabsOfficialAI score43

    Google's CC Gmail agent expands waitlist to Australia and New Zealand

    AIGoogle Labs has opened a waitlist for CC, its experimental AI productivity agent in Gmail, in Australia and New Zealand, and is expanding availability in the US and Canada. CC now helps manage calendars by connecting to Gmail and automatically creating events in a dedicated Google Calendar that update as plans change. Invitations to waitlisted users in the US and Canada begin rolling out today.

  3. VercelOfficialAI score42

    Vercel launches $1M hacker challenge to test Sandbox security

    AIVercel is offering up to $1,000,000 in a public hacker challenge testing its Vercel Sandbox against escapes from the Firecracker microVM and bypasses of the host-side network boundary. Rewards reach $50,000 per report, administered through HackerOne (@Hacker0x01). The company says agents can now exploit vulnerable sandbox boundaries, so it is testing its own defenses in the open.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesOfficialAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

  2. Chip HuyenXAI score22

    Chip Huyen asks for a model tiering system for agent orchestration

    AIChip Huyen asks what a good model tiering system looks like, since she is tired of naming specific models per vendor for her agent orchestrator. She wants to instruct the orchestrator by task tier, such as "use models tier ..." for a given kind of task, instead of listing Claude, OpenAI, and other models individually.

  3. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry adds five Claude agent features to Azure-hosted deployments

    AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.

    Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.

  4. Jason WeiXAI score45

    Jason Wei argues tool use cannot replace larger language models

    AIJason Wei now believes a small 1B-parameter "cognitive core" relying on tools is insufficient, because fast, natural recall without tool use matters. He cites speed, knowledge better learned through backpropagation than retrieved from search, and the greater reliability of already-known facts over repeated lookups. Since a 1B model has an information limit, he argues that demanding AI will still need larger models, not just tool access.

  5. Replit BlogOfficialAI score60

    Replit adds black-box pen tests that probe apps like external attackers

    AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.

    Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.

  6. Import AIBlogAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 16

Aug 16Sun
  1. Philipp SchmidBlogAI score58

    Controlling Android with Gemini 3.7 Flash and 150 lines of Python

    AIThe author built a Python agent that uses Gemini 3.7 Flash to control an Android emulator from raw screenshots, returning normalized 0–999 coordinates that are scaled to 1080x1920 pixels over ADB. In a test, the agent opened Chrome, closed popups, and solved one round of Wordle in two guesses without accessibility IDs or DOM access. The article presents the loop as usable for UI testing and task automation across native apps, webviews, and canvas interfaces, with code in an open-source quickstart repository.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogOfficialAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Aug 14

Aug 14Fri
  1. Augment Code BlogOfficialAI score62

    Augment rebuilds its Auggie CLI harness on Pi, cutting SWE-bench Pro task cost 53%

    AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.

    Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.

  2. Epoch AI · The Epoch BriefOfficialAI score42

    Epoch AI lists nine big AI questions its benchmarks aim to answer

    AIEpoch AI outlines nine open questions about AI capabilities, including whether AI can take over full jobs and whether benchmark scores are correlated. The author says Epoch's benchmarking work is built to help answer them, citing examples such as MirrorCode, Remote Labor Index, and the Epoch Capabilities Index (ECI). The post notes that benchmark scores are highly correlated across domains, and that ECI growth trends can help detect whether AI capability progress has accelerated.

  3. Andrew NgXAI score38

    Andrew Ng maps the four key skills for AI engineering

    AIAndrew Ng's team released an AI Engineering Skills Map, built from analysis of over 10,000 job postings and expert interviews, identifying four priority skills. The skills are building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. Ng says these skills matter for all developers, not only those with the AI Engineer title.

  4. Ali GhodsiXAI score46

    Databricks Smart Routing cuts AI coding task costs about 30% in Unity Gateway

    AIAli Ghodsi says Smart Routing on Databricks' AI Gateway lowers costs by about 30% without sacrificing quality. The quoted Databricks post says it matches each coding task to the right model and harness based on task needs, so higher-cost models focus on intelligence while lower-cost models compete on cost and performance.

Aug 13

Aug 13Thu
  1. Meituan LongCatOfficialAI score46

    LongCat-2.0 Free for One Week on Nous Portal with Hermes Agent

    AILongCat-2.0, Meituan's model, is now live on the Nous Portal and free to try with Hermes Agent for one week. Nous describes it as a 1.6T-parameter MoE with a 1M context built for agentic coding, scoring 70.8 on Terminal-Bench 2.1. It can ingest an entire codebase in one pass, and the Portal is at

  2. Ali GhodsiXAI score22

    Databricks CEO says AI agents with enterprise context drive 80% growth

    AIAli Ghodsi attributes Databricks' 80% growth at $7B to enterprise AI agents becoming usable, now that Genie Ontology automates the capture of organizational context. He says over 70% of queries on the platform are now generated by Genie agents, and that this usage drives consumption and revenue.

  3. Matei ZahariaXAI score44

    Databricks adds Smart Routing to Unity AI Gateway for coding agents

    AIDatabricks has made Smart Routing available in Unity AI Gateway to improve coding agent quality and cost. It matches each coding task to a suitable model and harness based on task needs while preserving good cache hit rates. Databricks says this can match frontier quality while cutting task costs by 30% or more.

  4. Varun MohanXAI score52

    Gemini 3.7 Flash Goes Live in Google Antigravity for Coding and Agents

    AIGoogle has made Gemini 3.7 Flash available in Google Antigravity, described as its most intelligent workhorse model yet for coding and agents. Users can download or upgrade Antigravity to try the model. The author, Varun Mohan, says the model brings a big capability improvement at half the API cost.

  5. Augment Code BlogOfficialAI score44

    Augment Code Expands AI Review Loop to Automate PR-to-Merge Workflow

    AIAugment Code describes an expanded AI-native review system in which specialized agents handle review, repair, and verification from pull request to merge. Humans still make judgment calls and the final merge decision, with the company claiming a 3× increase in code output in its earlier review system.

  6. Google AI DevelopersOfficialAI score75

    Google releases Gemini 3.7 Flash for coding and agentic tasks

    AIGoogle AI Developers announced Gemini 3.7 Flash as its most intelligent workhorse model yet for coding and agents, citing higher instruction adherence, first-pass code accuracy, and high-quality agentic execution. The post shows the model building a complex 3D web game in Antigravity, covering Three.js engine logic, asset orchestration with PBR textures and Nano Banana sprite sheets, and procedural sound effects.

    Why it matters: The post shows a concrete build workflow across engine logic, assets, and audio, which helps readers judge how the model handles multi-step agentic coding.

    Video from @googleaidevs's post
  7. koray kavukcuogluXAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

    Image from @koraykv's post
  8. Augment Code BlogOfficialAI score22

    Augment Code uses Cosmos to check enterprise pilot health against usage and deal data

    AIAugment Code's Solutions Architecture lead used the Cosmos agentic orchestration platform to build a live pilot-health view that combines product usage, GitHub and PR activity, Salesforce deal data, and customer call transcripts. Each account's health and board-level one-liner was checked against the customer's own stated success criteria, such as a 30% PR merge-time reduction. The article says the view refreshed from current Salesforce data and was designed to avoid inflating usage numbers through session lineage reconciliation.

  9. Ali GhodsiXAI score38

    Databricks passes $7B revenue run-rate, growing 80% year over year

    AIDatabricks announced it crossed a $7 billion revenue run-rate, growing over 80% year over year in Q2, and raised $5 billion in its latest fundraise. Lakebase reached a $100 million-plus run-rate, and Lakehouse reached $1.5 billion-plus, growing over 100% year over year, with continued positive adjusted free cash flow. The company says it will invest the capital in Lakebase, a serverless Postgres database for AI agents; Genie, AI coworkers for business data; and Unity AI Gateway, multi-AI governance for controlling costs.

  10. DeepSeekOfficialAI score68

    DeepSeek Harness v0.1 enters Developer Preview as an open-source agent harness

    AIDeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.

    Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.

  11. DeepSeekOfficialAI score62

    DeepSeek launches V4-Pro with Agent upgrades and OpenAI Responses API support

    AIDeepSeek announced the launch of DeepSeek-V4-Pro, citing major Agent upgrades and flexible reasoning effort settings of low, high, and max for V4-Pro and V4-Flash. The model supports the native OpenAI Responses API and is optimized for Codex with one-click setup. V4-Pro is available on the app and web through Expert Mode and via API, with model names unchanged.

    Image from @deepseek_ai's post
  12. DeepSeek API NewsOfficialAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek releases DeepSeek-V4-Pro-0813 with stronger agentic benchmark results

    AIDeepSeek has released DeepSeek-V4-Pro-0813 as the official version superseding the V4-Pro preview, built on the preview structure with a DSpark speculative decoding module. The model scores higher than the preview on the listed benchmarks, including Terminal Bench 2.1 at 87.9 and DeepSWE at 62.7, and the weights are under the MIT License.

    Why it matters: The release reports agent benchmark gains over the preview and lists vLLM and SGLang setup, useful for judging deployment cost and fit.

  2. Factory NewsOfficialAI score40

    Factory Launches Agent Effectiveness to Link Droid Usage to Delivery Outcomes

    AIFactory's Agent Effectiveness, now in Private Preview within Factory Analytics, connects Droid sessions to cycle time, work intent, and shipped artifacts drawn from project, issue-tracking, and source control tools. Its Throughput, Output, and Attribution views show where delivery is speeding up, how spend splits across feature, maintenance, bug-fixing, and exploration work, and which projects and issues the output maps to. Admins enable it by connecting Jira, Linear, GitHub, or GitLab and turning on the Advanced Analytics enterprise control.

  3. Cursor ChangelogOfficialAI score42

    Cursor Cloud Agents Start 3x Faster With Builds

    AICursor's Cloud Agents now start from prebuilt copies of development environments, cutting startup time by 3x, with environments booting 10x faster internally and 3x faster time to first token. Builds are included at no additional cost, and failed builds are not activated, so agents keep using the last successful build while users debug in the background.

  4. Jason WeiXAI score22

    Jason Wei argues private knowledge and human presence remain AI-resistant moats

    AIJason Wei argues that as AI gains advantages like driving better than humans, durable human moats remain in private knowledge that language models cannot access, such as high-end real estate and venture capital. He also points to entertainment and the arts, where human creation and achievement carry value, and to human presence, since time spent on someone is meaningful because a finite life runs out.