Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  2. Sierra BlogOfficialAI score62

    Sierra launches fleming-1 to detect AI agents calling by phone

    AISierra has launched fleming-1, a model that analyzes caller speech in real time and scores audio for signs it was generated by AI. It flags likely AI callers while keeping real people unflagged by default, and companies decide how to handle those calls. The model works with any voice agent built on Sierra, and Sierra also announced Personal Agent Protocol, an open standard for authorized agent-to-business interactions.

    Why it matters: The post explains why companies need to know when a caller is an AI agent, which frames the detection model as a business decision rather than an automatic block.

  3. MuseOfficialAI score20

    Muse builds a personalized movie feed and books tickets for users

    AIMuse, created by @JJEnglert, is trained on the user's movie and TV preferences and weekly curates a feed of films in theaters and on owned streaming services. Users can upvote or downvote picks to refine the feed, and when they want to see a movie, Muse books the ticket through a wallet link integration.

    Image from @Muse's post
  4. 🚨 AI News | TestingCatalogXAI score49

    Odyssey launches Odyssey-3 world model with public research preview

    AIOdyssey has launched Odyssey-3, its most powerful foundation world model, with a public research preview. Odyssey-3 Pro scored 66.1 on Physics-IQ Verified video-to-video with best-of-8 sampling, the highest reported result. The model generates environments from prompts and predicts changes in real time as users move through scenes.

    Image from @testingcatalog's post
  5. Luke EdwardsXAI score38

    Pocketty brings SSH and herdr agent alerts to iPhone and iPad

    AIPocketty launches as an SSH app for iPhone and iPad, built for herdr, that notifies users when an agent on any host is blocked. Tapping a notification opens the exact pane, and the app supports Tailscale and Bonjour natively, shows diffs for every agent turn, and requires no account or subscription.

    Video from @lukeed05's post
  6. SantiagoXAI score46

    Odyssey 3 Pro world model tops Physics-IQ and goes live

    AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

    Image from @svpino's post
  7. OdysseyOfficialAI score31

    Odyssey-3 world model debuts for physical AI and training environments

    AIOdyssey has released Odyssey-3, which it describes as a major leap toward world models that power physical AI, generate training environments, and enable new human experiences. The post invites readers to try Odyssey-3 at the company's website but gives no specific benchmarks, parameter counts, or pricing.

  8. OdysseyOfficialAI score34

    Odyssey-3 world knowledge can be applied to physical AI systems

    AIOdyssey says its Odyssey-3 model's learned world knowledge can be adapted by physical AI developers to control robots, power humanoids, drive cars, and fly drones. The post describes this as a capability for autonomous machines generally, without providing benchmarks, specifications, or availability details.

    Video from @odysseyml's post
  9. Aravind SrinivasXAI score22

    Perplexity Decider ranks first on DecisionBench at lowest cost

    AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.

  10. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  11. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  12. Vercel DevelopersOfficialAI score36

    StepFun's Step 5 Preview model now available on Vercel AI Gateway

    AIVercel says StepFun's flagship Step 5 Preview, built for agentic coding, research, and finance, is now live on AI Gateway. The model offers a 1M-token context window, accepts text and image input, and uses a 600B-parameter mixture-of-experts design with 27B parameters active.

  13. Grok BotOfficialAI score38

    Grok Bot can now help run Shopify stores

    AIGrok Bot can now connect to Shopify to check orders, track inventory, and keep product listings up to date. The feature lets merchants ask it questions about their business after linking their store. Shopify's related announcement says the new connectors also allow merchants to build teams of agents to help run their business.

    Image from @bot's post
  14. ClaudeDevsOfficialAI score46

    Anthropic adds monthly API credits for Max and Team plans

    AIAnthropic now provides monthly Claude Platform API credits to Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. The credits can be used on any Claude model, including in code or third-party harnesses.

  15. Arena.aiOfficialAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post