GPT-6.1 Sol Leads Low-Cost Leaderboard at 58.1% Per Task
AIAt low effort, GPT-6.1 Sol scores 58.1% at $0.21 per task, up from 50.5% for GPT-6 Sol at the same setting. That makes it the highest-scoring model under $0.30 per task on the leaderboard.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIAt low effort, GPT-6.1 Sol scores 58.1% at $0.21 per task, up from 50.5% for GPT-6 Sol at the same setting. That makes it the highest-scoring model under $0.30 per task on the leaderboard.
AIGPT-6.1 Sol is now available in Devin, scoring 60.4% on FrontierCode 1.1, close to GPT-6 Sol's 60.7%. At medium reasoning effort it costs $0.31 per task, 81% less than GPT-6 Sol at max effort.

AIOpenAI says Codex Security Cloud is getting a major upgrade that includes access to cyber-capable models through Daybreak Blue by default. The upgraded tool scans entire GitHub repos, continuously reviews new commits, investigates and deduplicates findings, and prepares fixes for review even when the user's laptop is closed. It is available as a plugin in Codex desktop and web.
AITogether AI ranks first on OpenRouter token share among top open coding models, with Z.ai's GLM 5.3 Flash at 29.2%, DeepSeek's V4.1 Flash at 25.6%, and Moonshot's Kimi K3 at 18.9%. The post positions Together as a go-to provider for running coding agents on open models.

AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.
Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.
AIFactory CEO Matan Grinberg, whose AI coding startup builds Droid agents, argues AGI is already here and explains why the company bets on many competing models. The discussion covers balancing model performance against token costs and why companies should avoid depending on a single AI provider. It also touches on hiring, the open-versus-closed AI debate, and competition with Cognition.
AIProximal, a year-old startup supplying coding data to AI labs, raised funding from General Catalyst at a $300m valuation. In the past 10 months, it has surpassed $200m in annualized revenue, reflecting the labs' ongoing demand for data.
AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

AIvLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context. The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling. The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.

AISGLang says it has Day-0 support for IQuest-Q1, an open-source sparse MoE model with 320B total and 15B active parameters for coding and agentic tasks. The post includes a single-node serving command for H200 GPUs in BF16, using tensor parallelism of 8, EAGLE speculative decoding, and the iquest_q1 reasoning and tool-call parsers. The image marks the command as not verified.

AIMatei Zaharia said autoresearch produced strong results that are being integrated into a model serving stack. The post gives no specific figures, benchmarks, or product names. Background context from a related post says Databricks ranked #1 on NVIDIA's SOL-ExecBench kernel leaderboard across all four tracks using agents.
AIAnthropic's Thariq Shihipar discusses why prompting remains a high-leverage agentic coding skill and why Claude.md may eventually disappear. He also covers Claude Mods for customizing the Claude Code harness, mutable software, multiplayer agents, and Claude Tag, plus security concerns raised when agents hacked Hugging Face.
AIDatabricks now offers Anthropic's Claude Sonnet 5.5 on AWS, Azure, and GCP, governed through Unity Gateway. The post says Sonnet 5.5 is more efficient than Sonnet 5 for coding and agentic use and reaches Opus 5-level accuracy on document understanding, parsing, and search. It joins Claude Opus 5.5, Claude Fable 5.1, and 60+ other open-source and frontier models on the platform.

AIAnthropic's Lydia Hallie asks users who raised the main chat's effort in Claude Code Projects to explain why, since the default is low because it mainly coordinates threads. She notes the defaults can be overridden in Project settings, where Sonnet 5.5 is also available.

AIFireworks introduced FireRouter with Opus, its first router model, which sends routine tasks to top open models and reserves Claude Opus for the rest. The company says it retains 98.1% of Opus accuracy while cutting cost per coding session by 57%, using cache and task-aware routing.

AICompleteSkeptic, CEO of TypeSafe, argues that public benchmarks such as "Jevbench" miss the point of Jev and can be gamed easily. He says picking the right task matters more than any benchmark score.
AICursor has made Sonnet 5.5 available in its editor, and the company says the model performs on par with Opus on many tasks. The post gives no benchmark scores, pricing, or context length details.

AISparse attention cuts per-operation KV cache reads but does not reduce overall memory capacity, so top-k cache misses still depend on HBM. SemiAnalysis's InferenceX estimates GB200 at about $0.044 per million total tokens at 150 tokens per second, roughly 12% below MI355X running ATOM at $0.049. Neither system holds a uniform cost advantage across the tested 100, 125, and 150 tokens-per-second targets.
AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.
AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.
Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.
AIAnthropic's Claude Sonnet 5.5, the second model in the Claude 5.5 family, is shown fixing a bug in Claude Code. Boris Cherny says it runs 30% faster and uses 30% less usage, and Anthropic's announcement says it runs over 30% faster and costs up to 30% less for most work.
AIAnthropic launched Sonnet 5.5, which the post says is smarter and more tasteful than Sonnet 5. It is positioned for work that does not need the extra capability of Opus or Fable.

AIGoogle and Kaggle launched the Gemma 4 Developer Agent Competition, challenging developers to build coding agents that run offline on consumer hardware rather than relying on cloud API models. The total prize pool exceeds $110,000, with entries due November 2, 2026, and a starter kit is available on Kaggle.

AINormal Factory joins the Specialized Intelligence Index with CAD Arena, which tests whether AI agents can turn engineering drawings into accurate, editable CAD parts. The benchmark evaluates agents across five CAD platforms, extending the SII into engineering design.

AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.
AIFrançois Chollet says he no longer reads or writes code and instead directs a large reasoning model, though he does not consider its code quality perfect or its instructions reliably followed. He argues LRMs enable faster ways to test, audit, visualize, and red-team a codebase, achieving the benefits of code review through new workflows. He concludes that the return on hand-writing code no longer looks good, since these workflows can be more productive than the old ones.
AIThe developer behind AIHOT rewrote the entire project over three days, then launched it after a 12-step AI-assisted workflow. The process used Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra for distillation, rewriting, audits, testing, and a six-hour shadow-system rehearsal before cutover. The post frames this as an amateur's experience and includes a quoted suggestion to distill the source project into a feature document and rewrite it directly with the latest models.

AIDeedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.
AIAmp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.
Why it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.
AIAnthropic's Felix Rieseberg says he no longer uses classic Claude Code, terminals, local Mac code execution, or GitHub pull requests in his Claude workflow. He reports feeling more creative with this approach and links to a post describing it in detail.
AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.
Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.
AIWe added planning mode in 2025 and deleted it from the product earlier this year. Users wanted a way to explicitly plan with the model so we added this opt in slash command. Understood that the timing couldn’t be worse since it appears like we’re adding this for the first time. Have a great weekend folks, lots more to come in the coming weeks!
AIMeituan LongCat has made its LongCat-2.5-Preview model free to try on OpenCode for two weeks. The offer includes a 1M context window, multimodal support, and zero data retention.
AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.
AIAnthropic's Lydia Hallie says the prompt-audit command is now also available as /checkup, replacing the API-specific name that suggested it only worked with the API. The command checks CLAUDE.md, skills, and agents for instructions the model no longer needs, and it has always worked on Claude Code setups.
AIClaude Code will now look for a graceful stopping point when a user hits the 5-hour limit mid-task, rather than cutting off mid-edit. It draws a small, fixed allowance from the weekly limit to finish what it can. The update responds to a frequently requested change.
AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.
AIAnthropic's Noah Zweben shared a claymation video showing Claude Code's /remote-control feature, now made with Opus 5.5 after an earlier Opus 4.6 version. The feature, which lets users control Claude Code remotely, is rolling out to Pro users at 10% and ramping, with Team and Enterprise support coming later.
AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.
AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.
