Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Sep 10

Sep 10Thu
  1. Thomas DohmkeXAI score12

    Dohmke says agents must run on iPhone Duo to matter

    AIThomas Dohmke says he will buy the new iPhone Duo but argues it cannot become his intelligent personal hub without an agent that can use the same apps he does. He claims, like the iPad, the hardware is held back by an operating system that treats software as operable only by humans. He concludes that agents using computers is the new paradigm.

  2. DeepSeekOfficialAI score46

    DeepSeek V4.1-Flash cuts KV cache to 1/4 HBM and 1/8 SSD

    AIDeepSeek says its V4.1-Flash model needs only 1/4 the HBM and 1/8 the SSD storage for its KV cache compared with the previous generation. Because cache-hit charges often make up a large share of agent costs, the company says the compressed cache significantly reduces those costs.

    Image from @deepseek_ai's post

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  2. Cursor ChangelogOfficialAI score73

    Cursor launches Projects for long-running, multi-agent coding work

    AICursor is launching Projects, a beta feature for larger work such as a feature, migration, or full app, rolling out to all users starting today. A coordinator agent plans the work, delegates it to implementing agents that can run in parallel, and runs on a cloud computer so it continues when the laptop is closed. Each Project keeps shared context files synced across cloud and local machines, and subscriptions let the coordinator act on Slack channels, schedules, or PRs without a prompt.

    Why it matters: The source details how a coordinator agent plans, delegates, and syncs shared context across cloud and local machines, useful for judging how long-running agent work might fit a team's workflow.

  3. Fireworks AI BlogOfficialAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  4. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

  5. RadixArkOfficialAI score38

    RadixArk's Miles integrates SGLang for fast, aligned post-training rollouts

    AIRadixArk says its Miles framework natively supports SGLang for fast rollouts while keeping rollout and training aligned for reliable post-training at scale. The post thanks the community for contributions and feedback shaping Miles. A related post from @adarshxs describes Miles v0.1 running fully async agentic RL on a 744B MoE across 64 GB300 GPUs.

  6. Mistral AIOfficialAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

Sep 8

Sep 8Tue
  1. Perplexity DevelopersOfficialAI score30

    Perplexity Search API now available in Hermes Agent

    AIPerplexity says its Search API is now available in Hermes Agent, giving it access to an index of more than 400 billion URLs. The API returns real-time results with snippets ranked by relevance.

    Video from @perplexitydevs's post
  2. Google Developers BlogOfficialAI score72

    Google releases ADK for Kotlin 1.0 for building production AI agents

    AIGoogle announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.

    Why it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.

  3. Factory NewsOfficialAI score34

    Factory Now on Claude Marketplace for Enterprise Autonomous Software Development

    AIFactory is now available on the Claude Marketplace, letting enterprise customers apply their committed Anthropic spend toward its autonomous software development platform. The platform automates the software development lifecycle, covering planning, implementation, testing, and security within one system, with enterprise deployment options that keep execution close to customers' code and infrastructure.

  4. Google Developers BlogOfficialAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  5. Jazzyear · InsightsNewsAI score62

    Arm expands from mobile IP into cloud, edge, and physical AI at Shanghai event

    AIAt Arm Everywhere China on September 8, 2026, Arm launched products spanning data center CPUs, mobile compute subsystems, and robotics platforms. The article says CSS for Mobile 2 integrates CPU, GPU, and neural accelerator for agent AI on phones, and Arm's Neoverse CSS N4 and AGI CPU target agent sandboxes in data centers. It also reports that Arm's Total Design ecosystem now covers over 80 partners for physical AI.

  6. InferactOfficialAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

  7. AI at MetaOfficialAI score34

    Meta launches Muse, a personal AI agent, with a safety deep dive

    AIMeta launched Muse, a personal AI agent that learns about users over time, and published a deep dive on how safety was built into its system. The company says the agent holds substantial personal context, which is why it was designed to be secure, safe, and private. Full details are in the linked security write-up.

    Video from @AIatMeta's post
  8. AI at MetaOfficialAI score67

    Meta introduces Muse, a personal agent powered by Muse Spark 1.3

    AIMeta announced Muse, a personal AI agent designed to get things done for users across many parts of life. The product is powered by Muse Spark 1.3, and the post links to an app download and a page describing how Muse was built.

    Why it matters: The announcement names Muse Spark 1.3 as the underlying model, giving readers a concrete product and model pairing to track.

    Video from @AIatMeta's post
  9. Mckay WrigleyXAI score80

    OpenAI shares agent-produced proof of Navier-Stokes Millennium Prize problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem. The problem asks whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and it has remained unresolved for roughly 90 years. The author, Mckay Wrigley, reposted the claim with his own remark about roughly 10k agents working in a datacenter.

    Why it matters: The quoted OpenAI post makes a major mathematical claim about the Navier-Stokes problem, so readers should weigh it against the proof's verification status.

  10. Werner VogelsXAI score50

    Werner Vogels highlights Kiro Crew's memory system drawing on brain evolution

    AIWerner Vogels says that after spending time with Kiro Crew since its launch, its memory system stands out for deciding what to keep, compress, and let go. He notes that Amazon engineers, starting from engineering constraints, arrived at an approach resembling the brain's evolved architecture. Per the referenced post, Kiro Crew is a persistent workspace that retains project context across sessions and runs scheduled jobs.

  11. Xiaomi MiMoOfficialAI score52

    Xiaomi MiMo Desktop enters invite-only beta as a desktop agent

    AIXiaomi MiMo has launched MiMo Desktop in invite-only beta, a desktop agent that turns Office files, images, video, audio, and zips into finished, editable output. Invitees also get limited access to next-gen MiMo models, and the post lists features including live previews, region-based editing with versioned rollback, automatic model routing, and browser and computer use with record and replay.

  12. NVIDIA · new models on Hugging FaceOfficialAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 7

Sep 7Mon
  1. Baidu Inc.OfficialAI score22

    Baidu launches AI, Evolving podcast on AI in scientific discovery

    AIBaidu has launched AI, Evolving, a new podcast series, with its first episode examining AI's growing role in scientific discovery through Famou's work on pine wilt disease. The post frames this as part of a broader trend in which AI takes on more of the research process itself. It asks whether research agents could become part of the infrastructure of discovery.

    Video from @Baidu_Inc's post
  2. Ian Johnson 🔬🤖XAI score38

    Ian Johnson: knowing what to ask AI for matters most for value

    AIOrbital is building an operating system that lets non-CS domains like science and mechanical engineering use its team's computer science expertise to build complex apps and research tools. The author argues that clearly specifying what you want from AI is the key lever for getting value, and that robust results are possible without a CS degree if the right pieces are in place. A quoted post on an ETH Zurich study of 100 developers suggests computer science background predicts vibe coding success more strongly than writing skill.

Sep 6

Sep 6Sun
  1. Satya NadellaXAI score22

    Opal powers new Copilot Autopilot experiences, now in Frontier rings

    AISatya Nadella highlights Opal, the technology behind one of Microsoft's new Autopilot experiences, now available in Frontier rings and coming to Copilot soon. The post links to a Microsoft Tech Community blog introducing Project Opal as a new way to complete task-based work.

  2. Satya NadellaXAI score34

    Copilot Autopilots now complete long-running multi-step work tasks

    AISatya Nadella says Microsoft is bringing new models into Copilot to handle increasingly complex work, from quick questions to delegated tasks and complete long-running jobs via Autopilots. As an example, an Opal-powered Autopilot on a secure Windows 365 Cloud PC sorts a month of trail cam footage, extracts species sightings, and builds a highlight reel, spreadsheet, PowerPoint, and Teams share.

    Video from @satyanadella's post

Sep 5

Sep 5Sat
  1. AI at MetaOfficialAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaOfficialAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

    Image from @AIatMeta's post
  3. AI at MetaOfficialAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. Matei ZahariaXAI score46

    Qwen3.8-Flash-Next runs at 68.3 tok/s on a single RTX 5090

    AIA Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.

  2. Andrew NgXAI score42

    Andrew Ng maps key skills for using AI coding agents effectively

    AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.

  3. Lewis Tunstall @ COLM 🌉XAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  4. Lewis Tunstall @ COLM 🌉XAI score22

    Research Preference Models Rank AI Research Ideas to Save Compute

    AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.

  5. Lewis Tunstall @ COLM 🌉XAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

    Image from @_lewtun's post