Replit's Muse can now build apps, per Amjad Masad
AIReplit's Muse can now make apps, according to Replit CEO Amjad Masad's post on X. The post gives no further details on features, availability, or pricing.

Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIReplit's Muse can now make apps, according to Replit CEO Amjad Masad's post on X. The post gives no further details on features, availability, or pricing.

AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.
AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.
AIClaude builds a model of a program's most complex parts, such as state machines or race-prone code, and searches that model for counterexamples that signal suspected bugs. It then reproduces those bugs and fixes them in the code. The post clarifies that the whole codebase is not formally verified, only the riskiest sections are modeled and checked.
AIClaude Code cloud sessions are now generally available, letting work continue while a laptop is closed. Existing Pro subscribers get a one-time $100 credit and Max subscribers get $250 to try them.
AICursor has made Rollouts and Security Reviewer available today on Teams and Enterprise plans. The company is including Rollouts usage credits for the next 10 days so customers can try it on real changes. More details are available in Cursor's blog post.
AICursor's Security Reviewer now completes reviews in 3.8 minutes on average, down from 4.8 minutes, a 21% speedup. The update was released alongside Rollouts.
AICursor has launched Rollouts, which write a monitoring plan and watch changes as they deploy. The post says deployments are verified so regressions can be caught before users see them.

AIAnthropic published a write-up on how it made the Claude frontend faster using agentic optimization techniques. The post coincides with Max Woolf's separate blog post showing that prompting agents can make code faster than current state-of-the-art libraries, with prompts and benchmark results included.
AICursor announced improvements to the token efficiency of its agent harness, linking to a blog post with the details. The post itself offers no figures or specifics beyond that headline claim.
AIHugging Face has released a new JavaScript package, huggingface/lerobot, that lets developers read LeRobot datasets on the Hub directly in the browser without downloading them. The post says coding agents can use it to build custom dataset viewers quickly, with an example UI implemented in about 600 lines of JS.

AIMinecraft creator Notch says on X that he is enjoying vibe coding and admits he may have been slightly wrong. Months earlier he had publicly rejected AI-written code, but he later began having AI build internal tools such as a map editor and node graph tools. The main post adds that he has accumulated a set of small tools for himself before much game development has happened.
AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.
AIAnthropic's Cat Wu points users to ways to try Opus 5.5 and Claude Tag in Slack. The main post carries no further details on access, features, or availability.
AICursor is launching two bots, Rollouts and Security Review, for Teams and Enterprise plans. Rollouts monitors each pull request as it deploys and reports change health per environment, while Security Review reports exploitable bugs on every pull request.
AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.
AISwitching effort levels mid-session on Opus 5.5 does not break the prompt cache, according to Lydia Hallie of Anthropic. This holds when running Claude Code v2.1.280 or later.
AIOn FrontierCode, Opus 5.5 at xHigh reasoning performed worse than at lower effort, a problem also seen with Opus 5. The author attributes this to scope creep, since FrontierCode penalizes unnecessary changes and models consistently score lower at higher reasoning efforts.

AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.
AIOpenAI announced GPT-6 Sol and GPT-6 Luna, rolling out today in ChatGPT Work and Codex. The rollout covers Plus, Pro, Business, Enterprise, and Edu users.
AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.
Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.
AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.
AIClaude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. Anthropic is defaulting to effort medium across products, which it says is comparable to Fable 5.1 on intelligence but faster. Rate limits will go 25% further on Opus 5.5 compared to Opus 5.
AILovable now offers Opus 5.5, which it says matches Opus 5's results while finishing builds in a third to half fewer steps. Internal benchmarks showed Opus 5.5 scoring 4 to 6% ahead of Opus 5 on verification discipline, with step reductions of 26% to 57% and input token reductions of 21% to 59% across tasks.
AIStepFun released an install script for Step Code that runs on macOS, Linux, or WSL with a single curl command. WSL is the recommended route on Windows, while PowerShell support is currently in beta. The project accepts issues and pull requests on GitHub.
AIStepFun's Step Code passed 72 of 89 tasks (80.9%) on Terminal-Bench 2.1, tying for the highest pass rate among evaluated harnesses while using fewer tokens than the other tied leaders. On Multi-Frame, it passed 110 of 150 tasks (73.3%) and averaged 5.09M tokens per task, the highest pass rate and lowest token use among six harnesses evaluated.

AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.
AIMoonshot AI's Kimi K3 is now available on Amazon Bedrock for coding, document analysis, and extended agent workflows. Bedrock provides access, encryption, and auditing controls, and explicit prompt caching is supported.

AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.
AIFrançois Chollet argues that engineers can delegate coding to AI but must never delegate understanding. He says that if code was previously the source of truth for understanding a system, teams need a new source-of-truth artifact and new workflows to replace code-centric practices.
AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.
AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.
AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.
Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

AIGoogle AI Developers showed a coding tutor built with Gemini 3.8 Live Extended Thinking that views the user's screen and calls functions to reference the p5.js library. The tutor points out the exact bug on screen and talks the user through the logic.
AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.
AIMax Woolf reports that prompting agents to iteratively optimize code can make it faster than current state-of-the-art libraries. He says the blog post includes his prompts and benchmark results.
AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.
AIThe post links to an official xAI news page for more details.