GPT-6.1 Sol Leads Low-Cost Leaderboard at 58.1% Per Task
AIAt low effort, GPT-6.1 Sol scores 58.1% at $0.21 per task, up from 50.5% for GPT-6 Sol at the same setting. That makes it the highest-scoring model under $0.30 per task on the leaderboard.
Updated
Updated
AIAt low effort, GPT-6.1 Sol scores 58.1% at $0.21 per task, up from 50.5% for GPT-6 Sol at the same setting. That makes it the highest-scoring model under $0.30 per task on the leaderboard.
AIGPT-6.1 Sol is now available in Devin, scoring 60.4% on FrontierCode 1.1, close to GPT-6 Sol's 60.7%. At medium reasoning effort it costs $0.31 per task, 81% less than GPT-6 Sol at max effort.
AIOpenAI says Codex Security Cloud is getting a major upgrade that includes access to cyber-capable models through Daybreak Blue by default. The upgraded tool scans entire GitHub repos, continuously reviews new commits, investigates and deduplicates findings, and prepares fixes for review even when the user's laptop is closed. It is available as a plugin in Codex desktop and web.
AIOpenCode says Space Bunny received an upgrade over the weekend, with another one coming soon. The free period has been extended by five days, now running through 10/05.
AITogether AI ranks first on OpenRouter token share among top open coding models, with Z.ai's GLM 5.3 Flash at 29.2%, DeepSeek's V4.1 Flash at 25.6%, and Moonshot's Kimi K3 at 18.9%. The post positions Together as a go-to provider for running coding agents on open models.
AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.
Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.
AIFactory CEO Matan Grinberg, whose AI coding startup builds Droid agents, argues AGI is already here and explains why the company bets on many competing models. The discussion covers balancing model performance against token costs and why companies should avoid depending on a single AI provider. It also touches on hiring, the open-versus-closed AI debate, and competition with Cognition.
AIYuchen Jin argues OpenAI and Anthropic are roughly tied on coding LLMs, so the competition now centers on which $100/$200 Pro plan is more generous and resets more often. He suggests the winner will be whoever has more GPUs and is willing to lose more money to capture the market. He adds that a future Gemini 4 could change this picture.
AIProximal, a year-old startup supplying coding data to AI labs, raised funding from General Catalyst at a $300m valuation. In the past 10 months, it has surpassed $200m in annualized revenue, reflecting the labs' ongoing demand for data.
AIQuick video: "Git in 100 Seconds: What Everyone Should Know in the Agent Era" === Though most followers already know git, I made the video anyway, so here it is 🤣
AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.
AIvLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context. The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling. The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.
AIMatei Zaharia said autoresearch produced strong results that are being integrated into a model serving stack. The post gives no specific figures, benchmarks, or product names. Background context from a related post says Databricks ranked #1 on NVIDIA's SOL-ExecBench kernel leaderboard across all four tracks using agents.
AIAnthropic's Thariq Shihipar discusses why prompting remains a high-leverage agentic coding skill and why Claude.md may eventually disappear. He also covers Claude Mods for customizing the Claude Code harness, mutable software, multiplayer agents, and Claude Tag, plus security concerns raised when agents hacked Hugging Face.
AIDatabricks now offers Anthropic's Claude Sonnet 5.5 on AWS, Azure, and GCP, governed through Unity Gateway. The post says Sonnet 5.5 is more efficient than Sonnet 5 for coding and agentic use and reaches Opus 5-level accuracy on document understanding, parsing, and search. It joins Claude Opus 5.5, Claude Fable 5.1, and 60+ other open-source and frontier models on the platform.
AIAnthropic's Lydia Hallie asks users who raised the main chat's effort in Claude Code Projects to explain why, since the default is low because it mainly coordinates threads. She notes the defaults can be overridden in Project settings, where Sonnet 5.5 is also available.
AIFireworks introduced FireRouter with Opus, its first router model, which sends routine tasks to top open models and reserves Claude Opus for the rest. The company says it retains 98.1% of Opus accuracy while cutting cost per coding session by 57%, using cache and task-aware routing.
AICompleteSkeptic, CEO of TypeSafe, argues that public benchmarks such as "Jevbench" miss the point of Jev and can be gamed easily. He says picking the right task matters more than any benchmark score.
AICursor has made Sonnet 5.5 available in its editor, and the company says the model performs on par with Opus on many tasks. The post gives no benchmark scores, pricing, or context length details.
AISparse attention cuts per-operation KV cache reads but does not reduce overall memory capacity, so top-k cache misses still depend on HBM. SemiAnalysis's InferenceX estimates GB200 at about $0.044 per million total tokens at 150 tokens per second, roughly 12% below MI355X running ATOM at $0.049. Neither system holds a uniform cost advantage across the tested 100, 125, and 150 tokens-per-second targets.
AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.
AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.
Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.
AIAnthropic's Claude Sonnet 5.5, the second model in the Claude 5.5 family, is shown fixing a bug in Claude Code. Boris Cherny says it runs 30% faster and uses 30% less usage, and Anthropic's announcement says it runs over 30% faster and costs up to 30% less for most work.
AIFelix Rieseberg, who is affiliated with Anthropic, shared a sailing game generated by Sonnet 5.5 at medium effort. The post links to a Claude artifact containing the game and gives no further details on its features or performance.
AIFelix Rieseberg of Anthropic shared a flight simulator built by Sonnet 5.5 at medium effort, linking to a Claude artifact. The post provides no further details on its features or performance.
AIAnthropic launched Sonnet 5.5, which the post says is smarter and more tasteful than Sonnet 5. It is positioned for work that does not need the extra capability of Opus or Fable.
AINormal Factory joins the Specialized Intelligence Index with CAD Arena, which tests whether AI agents can turn engineering drawings into accurate, editable CAD parts. The benchmark evaluates agents across five CAD platforms, extending the SII into engineering design.
AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.
AIFrançois Chollet says he no longer reads or writes code and instead directs a large reasoning model, though he does not consider its code quality perfect or its instructions reliably followed. He argues LRMs enable faster ways to test, audit, visualize, and red-team a codebase, achieving the benefits of code review through new workflows. He concludes that the return on hand-writing code no longer looks good, since these workflows can be more productive than the old ones.
AIThe developer behind AIHOT rewrote the entire project over three days, then launched it after a 12-step AI-assisted workflow. The process used Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra for distillation, rewriting, audits, testing, and a six-hour shadow-system rehearsal before cutover. The post frames this as an amateur's experience and includes a quoted suggestion to distill the source project into a feature document and rewrite it directly with the latest models.
AIDeedy argues every open source repository should have a roughly seven-minute explainer video like the one made for SQLite, covering its purpose, a high-level code map, a query's path through the codebase, core abstractions, and a real execution trace including join-order query planning. He says the video was generated with Opus 5.5 and Gemini 3.8 TTS, and he expresses amazement at how coherent and capable the model is.
AIAmp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.
Why it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.
AIAnthropic's Felix Rieseberg had Opus 5.5 create a parody of the "nihilistic" penguin, generating all models and textures from scratch. The model did this by using Blender as a Python library.
AIAnthropic's Felix Rieseberg says he no longer uses classic Claude Code, terminals, local Mac code execution, or GitHub pull requests in his Claude workflow. He reports feeling more creative with this approach and links to a post describing it in detail.
AIFelix Rieseberg, an Anthropic employee, says he remade his homepage with Opus 5.5 and pushed it hard, using it to create music, movies, textures, and Blender models. He says he is very happy with the result and links to his site.
AIOpenAI released GPT-6 Sol and Luna, priced 50 percent below GPT-5.6 promo API pricing, and rolling out in ChatGPT Work, Codex and the API, not yet in regular Chat. Anthropic released Claude Opus 5.5, described as roughly Claude Fable 5.1 level for 40 percent less than Opus 5 and over 30 percent faster, with Sonnet 5.5 and Haiku 5.5 due in coming weeks.
Why it matters: The recap puts OpenAI and Anthropic releases side by side, with pricing and capability claims that help compare the two launches.
AIA possibly more efficient approach than refactoring spaghetti code: Distill the source project into functional documentation, then rewrite it in place using the latest model.🤦♂️🤦♂️🤦♂️
AIWe added planning mode in 2025 and deleted it from the product earlier this year. Users wanted a way to explicitly plan with the model so we added this opt in slash command. Understood that the timing couldn’t be worse since it appears like we’re adding this for the first time. Have a great weekend folks, lots more to come in the coming weeks!
AIThe author ran a full-rewrite-scale task with Claude Opus 5.5 for over six hours, reading every thinking summary along the way. They call it stable and strong, rating it top tier across communication, comprehension, aesthetics, development, and long-horizon agent work, and wish OpenAI would catch up.