Skip to contentSkip to stories

Updated

Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. StepFunOfficialAI score37

    Step 5 Preview free in Cline for one week, StepFun says

    AIStepFun's flagship model Step 5 Preview is free to use in Cline for one week. The quoted background says it scores ahead of Kimi K3 and GLM-5.3 on DeepSWE, positioning it among the strongest open-weights coding models.

  2. Zhihao JiaXAI score62

    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post
  3. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  4. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  5. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  6. Unsloth AIOfficialAI score44

    Unsloth adds Windows OS-level sandboxing via Microsoft's mxc

    AIUnsloth now supports OS-level sandboxing on Windows by integrating Microsoft's open-source mxc repository for sandboxed code execution. The integration adds under 100 ms of overhead, according to the post. A setup guide is available in Unsloth's documentation.

    Image from @UnslothAI's post
  7. OpenAI NewsOfficialAI score26

    How Oracle turns days of work into minutes with ChatGPT and Codex

    AIOracle is using ChatGPT Work and Codex to turn specialist knowledge into fast, repeatable workflows across recruiting, engineering, and operations. The source does not provide figures, timelines, or specific results beyond the headline's claim that days of work can take minutes.

  8. Augment Code BlogOfficialAI score62

    Augment Code sells Cosmos, Auggie CLI, and Context Engine assets to Harness

    AIAugment Code is selling select assets, including Cosmos, Auggie CLI, and the Code Context Engine, to Harness, and the product team is moving to Harness. The company says Harness's integrated platform delivers these capabilities to customers more effectively than building them independently. Harness describes itself as building the Autonomous SDLC Platform for shipping AI-written code across enterprises.

    Why it matters: The announcement shows how a coding AI company is folding its products into a larger software delivery platform, a shift that shapes how enterprise teams will buy these tools.

  9. Tessl BlogOfficialAI score44

    Tessl Code Review Uses Repo-Owned Lenses to Make AI Review Context-Driven

    AITessl's Code Review defines review standards as skills in the repository, called lenses, routed to files by a repo-owned profile file. Because these team-visible standards are portable, lessons from review can feed back into code generation and maintenance, not only the next review.

  10. Charlie Barmore, CPA, CFE, CVAXAI score40

    Accountant AI lessons: 30 takeaways from 80 calls with firm owners

    AIOver six months, a solo founder held more than 80 calls with accountants about AI and wrote down 30 lessons from the recurring problems. The post's opening lessons say to treat AI setup like onboarding a new employee and to start by listing the tasks people hate doing. The author argues that many "model problems" are actually setup problems, and that a firm's software stack limits what AI can do.

  11. OpenRouterOfficialAI score44

    StepFun's Step 5 Preview launches on OpenRouter as agentic flagship

    AIStepFun's Step 5 Preview is now live on OpenRouter as the company's new flagship for agentic work. It uses a sparse MoE design with 27B active and 600B total parameters, a 1M context window, and accepts text, image, and video input. The post highlights strength in coding and professional knowledge work, especially finance.

    Image from @OpenRouter's post
  12. Gergely OroszXAI score26

    Developers working more with AI tools, citing more context switching

    AISoftware developer Gergely Orosz questions why he is working more despite AI tools, quoting Sam Newman's view that AI was meant to free developers from drudgery. Newman says most developers are doing more work, with more context switching and a loss of the big picture. The quoted post adds that AI assistants are not human partners and that pairing with them fragments the shared mental model of a program.

  13. OpenRouter · New modelsBlogAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  14. SiliconANGLE · AINewsAI score62

    Google Cloud launches Gemini agent for enterprise work across devices and apps

    AIGoogle Cloud introduced Gemini agent, a unified AI assistant that acts autonomously, generates code, and completes work across web, mobile, desktop, and third-party apps. It runs jobs on models matched to each task, including Gemini Flash and a flagship frontier model, with Anthropic Claude models also available. Hard spend limits per project let companies enforce budgets and charge AI costs to departments.

  15. JetBrains AI BlogOfficialAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  16. vLLMOfficialAI score62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    AIvLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

    Image from @vllm_project's post
  17. The DecoderNewsAI score72

    AI hacking tools let a likely single attacker breach multiple South Korean banks

    AIA suspected Chinese-speaking attacker breached several South Korean financial institutions between late September and early October 2026, reportedly stealing over 25,000 records from Shinhan Bank alone. The attacker used ARTEX, a Chinese open-source tool that uses AI language models to automate finding security flaws, and models named in the report include DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6.

  18. Ant LingOfficialAI score22

    Ant Ling's Ling-3.1-flash now live on AI/ML API

    AIAnt Ling announced a day-zero collaboration with AI/ML API, making Ling-3.1-flash available there for agentic and cowork scenarios. AI/ML API describes it as a 560B-parameter MoE model with about 25B active per token and up to 1M context, built for agents, coding, and long documents. The model is free to try on AI/ML API until October 13.

  19. Ant LingOfficialAI score34

    Ling-3.1-flash is free on OpenCode for a limited time

    AIAnt Ling's Ling-3.1-flash is now free on the OpenCode coding harness for a limited week-long period. The model has 560B total parameters, 25B active parameters, and a 262K context window, according to the OpenCode background post. Ant Ling says it shows strong coding performance and thanks OpenCode for day-zero support.

  20. Latent SpaceBlogAI score73

    Claude Haiku 5.5 launches at GPT-6 Luna pricing with 1M context

    AIAnthropic released Claude Haiku 5.5, priced the same as OpenAI's GPT-6 Luna, with a 1M-token context window. Artificial Analysis scored it 43 on its Intelligence Index, slightly ahead of GPT-6 Luna at 38, but it uses about 3x more output tokens at max effort.

  21. howie.seriousXAI score46

    Agent bottleneck is human understanding, not model capability

    AIThe author argues that in agent workflows, the real bottleneck is whether users can precisely express requirements, not the model or agent capability. When people work outside their expertise, they lack the precision needed for prompts and plans, forcing many imprecise iterations that waste time and tokens. The suggested fix is to have the model first teach the unfamiliar domain knowledge before acting.

  22. Claude Code · GitHub ReleasesOfficialAI score22

    Claude Code v2.1.294 fixes prompt and agent hook judgment

    AIClaude Code v2.1.294 fixes prompt and agent hooks written as instructions, which had allowed actions they should block. It also improves how prompt hooks on Stop and SubagentStop are judged, making Claude less likely to stop early.

  23. Claude BlogOfficialAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

Oct 7

Oct 7Wed
  1. Orange AIXAI score34

    Next Token episode 5 covers Personal Agents, open-source software, and hardware projects

    AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.

  2. HorizzonXAI score22

    Solo founder shares Devin Max experience and favorite features

    AIA solo founder writes that after a rocky first trial, they gave Devin Max a second chance and now favor Devin Cloud for shipping real projects. The post praises Cognition and Devin Max's model access and pricing, and covers SWE-1.7 Lightning's speed and its tendency to consume limits quickly. It also compares GPT 6 Astra and Fable 5.1, finding Fable more efficient for bug fixes and features.

  3. Gizmodo · AINewsAI score51

    Vibe-Coded Artcraft Suite Offers Free Photoshop Alternative on GitHub

    AIDeveloper Brandon Thomas used Claude Opus 5.5 and Rust to build Artcraft, a free open-source suite with Photocraft, Vectorcraft, and other apps that mimic Adobe products. The author tested Photocraft and found its basic editing commands worked where expected, but Free Transform behaved unpredictably. Thomas describes the software as early alpha and invites developers to contribute.

  4. Epoch AIOfficialAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. IThome · AINewsAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.