Skip to content

#Reasoning

Oct 6

Oct 6Tue
  1. ARC PrizeAI score29

    @SpaceXAI On ARC-AGI-2, Grok 4.7's best score is 61.4% at $2.01/task, compared to Grok 4.6's 67.1% at $0.76/task, a 166% increase in cost per task despite identical input and output token pricing. Full results: https://arcprize.org/results/xai-grok-4-7

    @SpaceXAI On ARC-AGI-2, Grok 4.7's best score is 61.4% at $2.01/task, compared to Grok 4.6's 67.1% at $0.76/task, a 166% increase in cost per task despite identical input and output token pricing. Full results: https://arcprize.org/results/xai-grok-4-7

  2. ARC PrizeAI score38

    Grok 4.7 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-3: 1.8%, $2.7k (standard harness), 10.0%, $4.8k (provider adapter harness) - ARC-AGI-2: 61.4%, $2.01/task - ARC-AGI-1: 90.2%, $0.64/task Grok 4.7 scores higher than Grok 4.6 on ARC-AGI-1, but lower on ARC-AGI-2 and 3.

    Grok 4.7 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-3: 1.8%, $2.7k (standard harness), 10.0%, $4.8k (provider adapter harness) - ARC-AGI-2: 61.4%, $2.01/task - ARC-AGI-1: 90.2%, $0.64/task Grok 4.7 scores higher than Grok 4.6 on ARC-AGI-1, but lower on ARC-AGI-2 and 3.

  3. Latent SpaceAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    Reflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. Sophia YangAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    Sophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

  2. Clément DelangueAI score72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    Reflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    AIWhy it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.

Oct 3

Oct 3Sat
  1. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Homura-2B-FP4 released as NVFP4 quantized translation model

    IndexTeam released Index-Homura-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Homura-2B multilingual translation model, which supports 150 languages. The quantized checkpoint shows a 5.73% perplexity increase over the BF16 original (3.5011 to 3.7017) on a fixed corpus, and its zh-en and en-zh outputs are semantically equivalent under greedy decoding. Full FP4 acceleration requires an NVIDIA Blackwell GPU, while the source recommends the FP8 build for Hopper and Ampere hardware.

Oct 1

Oct 1Thu

Sep 30

Sep 30Wed
  1. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    Google AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    AIWhy it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

  2. Google DeepMindAI score36

    With a 1M token output limit, Argon adds a deeper level of reasoning to tackle long, multi-step problems in one go. Feedback from early testers will help us strengthen our systems before we roll out more broadly to developers, enterprises, and consumers soon. Find out more → https://goo.gle/4rZBiSd

    With a 1M token output limit, Argon adds a deeper level of reasoning to tackle long, multi-step problems in one go. Feedback from early testers will help us strengthen our systems before we roll out more broadly to developers, enterprises, and consumers soon. Find out more → https://goo.gle/4rZBiSd

  3. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    Google DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    AIWhy it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  4. Artificial Analysis ArticlesAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    Korean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  5. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    Artificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    AIWhy it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

Sep 29

Sep 29Tue
  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    BAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    AIWhy it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score40

    InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs

    InternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.

Sep 22

Sep 22Tue
  1. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    Fireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    AIWhy it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    Xiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    AIWhy it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

  2. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    Xiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    AIWhy it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  3. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    Xiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    AIWhy it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

Sep 17

Sep 17Thu
  1. OpenBMBAI score36

    Love this. 🙌 MiniCPM5-2B was built for exactly this: on-device, offline, download from Hugging Face and run. 2.5B, native 128K, hybrid Think / No-Think in one checkpoint. First calling the puzzle impossible, then correcting itself and writing a working verifier is the kind of reasoning we want to see at this size. Thanks for shipping it. @RunAnywhereAI

    Love this. 🙌 MiniCPM5-2B was built for exactly this: on-device, offline, download from Hugging Face and run. 2.5B, native 128K, hybrid Think / No-Think in one checkpoint. First calling the puzzle impossible, then correcting itself and writing a working verifier is the kind of reasoning we want to see at this size. Thanks for shipping it. @RunAnywhereAI

Sep 15

Sep 15Tue

Sep 14

Sep 14Mon
  1. Intern Large ModelsAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    Intern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

Sep 13

Sep 13Sun
  1. Fireworks AI BlogAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    Fireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

Sep 12

Sep 12Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score58

    Shanghai AI Lab releases Intern-S2-397B, a 397B multimodal scientific model

    Shanghai AI Lab's InternLM team released Intern-S2-397B, a multimodal foundation model for scientific intelligence and long-horizon agents. The model uses visual pre-training on raw scientific literature pages, multi-task reinforcement learning across more than 20 scientific domains, and agentic reinforcement learning in sandboxed environments.

Sep 11

Sep 11Fri
  1. BAAIAI score46

    Introducing AREX. A research agent from @BAAIBeijing . It does not treat a hard question as one long search. It drafts a candidate, checks every constraint, and goes back for what is still open. 122B MoE. 10B active. On hard search benches it sits next to GPT-5.4. Here is how it works👇

    Introducing AREX. A research agent from @BAAIBeijing . It does not treat a hard question as one long search. It drafts a candidate, checks every constraint, and goes back for what is still open. 122B MoE. 10B active. On hard search benches it sits next to GPT-5.4. Here is how it works👇

Sep 10

Sep 10Thu
  1. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    Cognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    AIWhy it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

Sep 7

Sep 7Mon
  1. Tencent HunyuanAI score44

    ✔️Hy4 preview just shipped an upgrade. You flagged it: long thinking + over-verification on complex tasks. We optimized it. Now live for everyone. Same task quality. Fewer turns. Lower in/out tokens. Bench + human eval both confirm. We’ll keep iterating fast. Try it and tell us what still breaks. Try on WorkBuddy: https://WorkBuddy.ai

    ✔️Hy4 preview just shipped an upgrade. You flagged it: long thinking + over-verification on complex tasks. We optimized it. Now live for everyone. Same task quality. Fewer turns. Lower in/out tokens. Bench + human eval both confirm. We’ll keep iterating fast. Try it and tell us what still breaks. Try on WorkBuddy: https://WorkBuddy.ai

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    OpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.