Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  2. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  3. Mike KnoopAI score57

    Tufa Labs reaches 83.06% on ARC-AGI-2, 2% short of the grand prize

    AIMike Knoop says the top ARC Prize 2026 ARC-AGI-2 score of 83.06% by Tufa Labs is only 2% short of the 85% grand prize threshold. The challenge runs under strict Kaggle compute limits with no internet access, and the winning solution is set to be open sourced. The image shows the leaderboard with RabbitHole at 76.94%, nvbanana at 74.17%, Yi-Chia Chen at 55.14%, and Kha Vo at 37.50%.

  4. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS adds voice replication and prompt-designed voices

    AIGoogle launched Gemini 3.8 Flash TTS and Flash-Lite TTS, letting users replicate a voice from 30 seconds of audio or design one from a text description. The post claims #1 on Hume's Voice Design Benchmark and top placement in Voice Arena across 6 languages. Voices can be directed line by line with style and inline tags, with a consent check on replication and SynthID on every clip. Availability is through the Gemini API and Google AI Studio.

    Image from @_philschmid's post
  5. QwenAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Image from @Alibaba_Qwen's post
  6. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  7. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Image from @ModelScope2022's post
  8. KrASIA · Big TechAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  9. Mike KnoopAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. ModelScopeAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Image from @ModelScope2022's post
  2. TinkerAI score25

    Tinker fine-tunes Qwen3.6 for Jev-style probability prompts in 10 minutes

    AITinker says an open LLM can serve a Jev-like interface that takes discrete options and returns fast probabilities, since next-token prediction is already a probabilistic classifier. A post by @ekzhang1 reports that a $5, 10-minute supervised fine-tuning run on Tinker improved Qwen3.6-35B-A3B's handling of Jev-style prompts, with +8% on GPQA Diamond and +12% on MMLU-Pro.

  3. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  4. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

  5. Tri DaoAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  6. StepFunAI score27

    StepFun's Step Code tops Terminal-Bench 2.1 and Multi-Frame with fewer tokens

    AIStepFun's Step Code passed 72 of 89 tasks (80.9%) on Terminal-Bench 2.1, tying for the highest pass rate among evaluated harnesses while using fewer tokens than the other tied leaders. On Multi-Frame, it passed 110 of 150 tasks (73.3%) and averaged 5.09M tokens per task, the highest pass rate and lowest token use among six harnesses evaluated.

    Image from @StepFun_ai's post
  7. StepFunAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  8. Sebastian RaschkaAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  9. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  10. AI SupremacyAI score45

    TypeSafe AI's Jev Is a Non-LLM Probabilistic Classifier for Fast Software Decisions

    AITypeSafe AI released Jev, a transformer-based System-1 model that outputs calibrated probabilistic decisions instead of generating tokens, returning answers in 70–500 ms at $0.042 per million input tokens. The model is built for typed Choice, Score, and yes/no questions inside software pipelines, and it is available to everyone without a waitlist, with $5 in starting credits. Vercel, Cloudflare, LangChain, and Langfuse have added Jev to their platforms.

  11. Tencent HyAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

  12. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Sep 21

Sep 21Mon
  1. StepFunAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  2. Xiaomi MiMoAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  3. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  4. Xiaomi MiMo · new models on Hugging FaceAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.