Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 20

Sep 20Sun

Sep 19

Sep 19Sat
  1. StepFunAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

  2. StepFunAI score38

    StepFun previews Step 5 for large-scale research and analytical deliverables

    AIStepFun has previewed Step 5, an agent built for professional knowledge work spanning large-scale research, structured analysis, and interactive reporting. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years. In another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.

  3. StepFunAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

Sep 18

Sep 18Fri
  1. TinkerAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. Mike KnoopAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  3. Google AIAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  4. One Useful Thing (Ethan Mollick)AI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  5. Noam BrownAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  6. GitHub Blog · AI & MLAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

  7. Google · AI blogAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

Sep 17

Sep 17Thu
  1. KrASIA · Big TechAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  2. Together AI BlogAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  3. AnthropicAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  4. Josh WoodwardAI score40

    Google Labs launches CC, an AI agent for family logistics

    AIGoogle Labs has announced CC, an AI agent built for families that can be connected to up to 5 members. It syncs schedules and to-dos through shared Google Calendar and Tasks, sends a shared "Your Day Ahead" brief each morning, and handles tasks such as meal plans, shopping lists, and paperwork under user direction. It is available by waitlist or upgrade in the US for users 18 and older.

  5. Google LabsAI score44

    Google Labs' CC agent expands to families, sharing one daily brief and calendar across up to six members

    AIGoogle Labs has turned its experimental CC agent into a family and household assistant that supports up to six members, each with a shared view of the day ahead. CC has its own Google account, sees only what members choose to share, and connects to Calendar and Tasks. It is available as an early experiment on web and mobile for U.S. users 18 and older with a personal Google account.

  6. Google AI StudioAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  7. Dwarkesh PatelAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

  8. OpenBMBAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

  9. Baidu Inc.AI score22

    Apollo Go plans to seek commercial autonomous driving approval in Hong Kong

    AIBaidu's Apollo Go plans to apply for commercial operation of its autonomous driving service in Hong Kong, citing the HKSAR Government's support in its first Five-Year Plan and the 2026 Policy Address. The company says it builds on fully driverless trials already conducted in the city. Baidu hopes Hong Kong can become a global benchmark for commercial autonomous driving in right-hand-drive markets.

  10. KrASIA · Big TechAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

  11. Gemini API ChangelogAI score38

    Antigravity Agent 09-2026 replaces 05-2026 with new built-in file and search tools

    AIGoogle released the antigravity-preview-09-2026 agent, which replaces and deprecates antigravity-preview-05-2026. Remote sandbox users reading only output_text or model_output steps need only update the agent string, while local-environment users or those parsing function_call steps must adapt to renamed tools, PascalCase parameters, and line-range file edits. The 05-2026 preview shuts down on October 5, 2026.

Sep 16

Sep 16Wed
  1. hardmaruAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Google Developers BlogAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  3. Greg BrockmanAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.