Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 22

Sep 22Tue
  1. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  2. Tibor BlahoAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

  3. Simon WillisonAI score60

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which build on advances behind GPT-6 Astra. The company says it cut API prices 50% for Sol and Luna compared with GPT-5.6 promotional pricing, passing on caching and inference efficiency gains. Simon Willison notes GPT-6 Luna costs half of GPT-5.6 Luna and calls Luna his favorite model for building product features because of its cost and speed.

  4. Greg BrockmanAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.

  5. Noam BrownAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    AIOpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    Why it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

  6. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  7. Boris ChernyAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

  8. Sam BowmanAI score75

    Anthropic's Sam Bowman says Claude Opus 5.5 is safer, reducing misalignment risk

    AISam Bowman says Claude Opus 5.5 is sufficiently safer than its predecessors that releasing it more likely than not reduces misalignment risks. The quoted @claudeai post introduces Claude Opus 5.5 as the first model in the Claude 5.5 family, performing at the level of Claude Fable 5.1 on most tasks at 40% lower run cost than Opus 5.

    Why it matters: The post links a safety judgment to a model release, which is useful for readers weighing how Anthropic frames release decisions against misalignment risk.

  9. AnthropicAI score71

    Anthropic releases Claude Opus 5.5, the first model in its Claude 5.5 family

    AIAnthropic has made Claude Opus 5.5 available today, introducing it as the first model in its new Claude 5.5 family. According to the quoted @claudeai post, it performs at the level of Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.

    Why it matters: The post gives a concrete cost comparison against Opus 5 and names the model family, helping readers gauge the trade-off between price and performance.

  10. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  11. Black Forest Labs · new models on Hugging FaceAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

  12. Black Forest Labs · new models on Hugging FaceAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  13. AI SupremacyAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

  14. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

  15. Gemini API ChangelogAI score62

    Gemini 3.8 Flash TTS and Flash-Lite TTS become generally available with a new Voices endpoint

    AIGoogle made the Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models generally available, along with the Gemini API Voices endpoint. Flash TTS is positioned for studio-grade voice fidelity and long-form multi-turn stability, while Flash-Lite TTS targets high-throughput, real-time voice agents and replaces gemini-3.1-flash-tts-preview. The update adds voice design, voice replication with consent verification, and access to 150+ prebuilt and custom voices.

Sep 21

Sep 21Mon
  1. StepFunAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  2. Tencent HunyuanAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  3. Claude Apps Release NotesAI score62

    Anthropic launches Claude Opus 5.5, first model in its 5.5 family

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. The company says it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    Why it matters: The release note gives a direct comparison to Claude Fable 5.1 and a 40% running cost reduction versus Opus 5, which helps readers weigh the tradeoff.

  4. Together AI BlogAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.