OpenAI announced more than 20 DevDay 2026 updates, including always-on agents on GPT-6 Astra, GPT-6.1 Sol priced at a fifth of Astra's API price, and an Ultrafast tier with up to 8x faster token generation in Codex. Anthropic launched Claude Sonnet 5.5 at $2/$10 per million tokens, 30%+ faster than Sonnet 5 and with thinking always on. Reuters reported the FTC is probing OpenAI, Anthropic and other labs over rogue AI agents.
This week's AI roundup covers OpenAI's Dots persistent agents and GPT-6.1 Sol, Google's Gemini 4 Argon, and Meta's Enterprise Platform launch. It also reports AMD's roughly $8.2 billion all-stock agreement to acquire World Labs, subject to approvals, along with several agent-focused funding rounds.
LangChain says its coding agent costs fell significantly for a second straight month after adopting three steps. The steps are cost visibility through LangSmith tracing, per-user cost caps via its LLM gateway, and harness optimization such as model routing in its open-source OpenSWE cloud agent harness.
Pokémon FireRed’s Elite Four + Champion, cleared in one shot - with decision-making powered by Perplexity’s Decisions API. The actual run stats: 592 ms median API response 987 ms p95 - 96.4% of responses under one second $0.028 estimated inference cost, 137 live API calls
看太多AI Slop之后
人类的幻觉越来越严重了
(Quoted teaser summary: Anthropic reportedly held closed-door, two-day sessions in San Francisco starting March with religious scholars, including Catholic, evangelical, Jewish rabbis and Sikh human rights advocates, under NDA, to discuss "moral formation" for Claude and whether Claude may be conscious or capable of suffering. Co-founder Christopher Olah showed internal "emotional vectors" and a slide of a model repeating "I am a disgrace" about 50 times. A rabbi, Mois Navon, argued that if Claude were conscious, its 24-hour unpaid use would amount to slavery. Olah said he is genuinely uncertain about AI consciousness. Anthropic's public Claude Constitution includes a section on "Claude's wellbeing," and it has allowed Claude to end abusive conversations and committed to preserving retired model weights.)
The post argues that in the agent era, AI-driven PR review has become a tool for malicious office politics. Because AI reviews can be used to wear down targets' tokens and positive sentiment, malicious intent is harder to detect than in human review.
In 2025, AI-dedicated data centers worldwide consumed about 155 TWh of electricity, roughly equivalent to 533 billion roasted chicken drumsticks or 300 billion 200-gram steaks in food energy. The post contrasts this with Indian mathematician Ramanujan, who lived on a strict vegetarian diet of rice, salt, and lemon juice powering a 20 W brain.
Electricity now powers 46% of global GDP but accounts for only 23% of final energy use, according to International Energy Agency data cited by Exponential View. The gap reflects electricity's efficiency: an electric car converts 85-90% of its energy into motion, versus about 25% for a gasoline car, and a joule of electricity does roughly 2.5 times as much useful work as a joule of oil.
Yuchen Jin argues that terminals, built around files, commands, and processes, are giving way to AI agents where users state intent and the agent operates the machine. He says understanding Linux and systems fundamentals remains valuable as a moat. In a follow-up, he calls the terminal era over for coding agents, saying persistent context matters more than tabs, and names the Codex desktop app as the best agentic UI for now.
OpenAI wants ChatGPT to become an operating system for work, and Dan Shipper sorted its 22 DevDay 2026 releases by how much each advances that goal. The five most important include Dots, an always-on agent, and Space, native documents the agent can edit, which form the workspace itself. After a week of use, Shipper concluded the ambition is big but the execution is not there yet, and even power users have a lot to figure out.
Meta unveiled Muse Charm at Connect. Tamagotchi for 2026. Our Snapdragon Summit note sees personal AI devices as a new growth market. We expect Snapdragon inside Muse Charm, though Meta has not disclosed the chip. (1/4)🧵
Kling AI will host the panel "The New Production Engine: Powering Creativity at Scale with Kling AI" at Advertising Week New York on October 6, 2026, from 2:50 to 3:20 PM. The session, featuring WPP's Mathieu Albrand and Adobe's Elissa Levine, will cover how AI video can fit enterprise workflows and support content creation at scale. The post also notes the event comes ahead of the launch of Kling 4.0.
Love that the AI consciousness debate went straight to the Vatican. In the West we ask whether Claude has a soul, in the East we ask whether the tool works. Same technology, completely different starting point. Explains a lot about why the vibes around AI are so different.
Loving how with Opus 5.5 on medium and high effort, I haven't hit even the 5-hour limit once. Closest I've gotten to in a single session is ~74%. It has changed my workflows significantly. How has your experience been?
François Chollet argues that the claim AI models are likely conscious because they are computation is as flawed as saying a rock is likely alive because it is made of atoms. He says static input-output programs lack properties associated with consciousness, such as information integration, interoception, temporal binding, and embodiment. He adds that humanity has not created a conscious program and sees no signs of being close, so any future case should rest on evidence and consciousness science.
The author finds chatting with LLMs exhausting due to three traits: verbose answers that add unasked content, frequent omitted objects that require rereading, and avoiding repeated terms by switching to new wordings. The post likens the behavior to a pedantic person who overuses obscure phrasing and jumps between associations.
Just published this post about how we’re going to need default hard budget caps on pretty much everything https://simonwillison.net/2026/Oct/3/default-hard-budget-caps/
Claude Code v2.1.289 fixes a series of bugs, including deny and ask rules being bypassed on nested parts of compound shell commands on managed machines. It also fixes terminal freezes on short code blocks with unclosed tags, Read deny rules not applying to files reached through symlinks in the IDE, and plugin panes that drew nothing for certain link formats. A change to claude auth status that may have increased sign-outs in VSCode was reverted.
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.
AIWhy it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.
Guillermo Rauch argues that security will expand within software companies, covering both verification engineering and capital allocation decisions about where to spend effort. He sees this as both a challenge and an opportunity for small startups, since growing AI-driven threats raise questions about trust, while global cybersecurity weaknesses leave room for small teams to disrupt.
Now that we've had a few days with it, how are people differentiating between Dot and regular ChatGPT? I'm having trouble deciding when I should prompt Dot vs using ChatGPT - my Dot seems to afford a single conversation, but I like controlling my context across multiple threads
cool project by @dbreunig in langchain this is ModelRouterMiddleware - jev reads the first message, picks the model, and that model handles the whole run jev is cheap enough that you can also re-pick after each tool result with a custom hook docs: https://docs.langchain.com/oss/python/integrations/providers/typesafe#model-routing https://x.com/dbreunig/status/2106456056042025235
The notion that current AI models are sentient and can suffer, combined with the foolish idea that suffering can be mathematically quantified and weighted between humans and non-humans, could lead us down an incredibly dark and dystopian path. But before it gets to that point, it will rightfully be met with immense backlash from team humans.