Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 24

Jun 24Wed
  1. PaddlePaddleAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    AIPaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

    Image from @PaddlePaddle's post

Jun 23

Jun 23Tue
  1. Lil'Log (Lilian Weng)AI score40

    Scaling Laws, Carefully: Early Empirical Power-Law Studies of Loss, Data and Model Size

    AILil'Log examines early empirical work showing that deep learning generalization error follows power-law curves as training data and model size grow. Hestness et al. (2017) found the exponent reflects the problem domain rather than the architecture, while Rosenfeld et al. (2020) modeled loss jointly as a function of model size N and data size D, fitting parametric forms on small configurations to extrapolate to larger ones.

  2. PaddlePaddleAI score38

    PP-OCRv6 lightweight OCR model challenges large VLMs with 34.5M params

    AIPaddlePaddle introduced PP-OCRv6, a lightweight OCR architecture built on the LCNetV4 backbone, in the first episode of its tech deep dive series. The post says PP-OCRv6_medium reaches 86.2% detection Hmean and 83.2% recognition accuracy, surpassing PP-OCRv5_server while running faster. Three model specs—Tiny, Small, and Medium—target edge CPU devices, balanced deployment, and industrial high-accuracy pipelines.

    Image from @PaddlePaddle's post

Jun 19

Jun 19Fri

Jun 18

Jun 18Thu
  1. Cohere · new models on Hugging FaceAI score43

    Cohere Releases Open-Source 2B Arabic Speech Recognition Model Transcribe Arabic

    AICohere and Cohere Labs released Cohere Transcribe Arabic, an open-source 2B-parameter Arabic automatic speech recognition model under Apache 2.0. It is optimized for Arabic, Arabic dialects, English, and Arabic-English code-switched speech, using a Conformer encoder-decoder architecture supported natively in Transformers. The model's average WER of 25.87 and CER of 11.80 on the Open Universal Arabic ASR Leaderboard, as of 07.07.2026, is reported in the source.

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    AIOpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    Why it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

  2. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. ByteDance · new models on Hugging FaceAI score24

    Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

    AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 11

Jun 11Thu
  1. OpenRouter BlogAI score74

    OpenRouter Fusion panels beat individual models on the DRACO deep research benchmark

    AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.

    Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.

  2. Moonshot AI (Kimi) · new models on Hugging FaceAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score34

    EvoQuality: ByteDance's self-evolving VLM for image quality assessment without human labels

    AIEvoQuality is a ByteDance vision-language model for no-reference image quality assessment that generates pseudo-ranking labels through pairwise majority voting and refines them with GRPO, requiring no human-annotated quality scores. On the paper's setting, it raised weighted-average PLCC from 0.615 to 0.770 and SRCC from 0.570 to 0.726 over its Qwen2.5-VL-7B backbone. The model is recommended for research and pre-production assessment, not as the sole criterion for high-stakes decisions.

Jun 9

Jun 9Tue
  1. ByteDance · new models on Hugging FaceAI score28

    ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation

    AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.

  2. Andrej KarpathyAI score65

    Karpathy Calls Claude Fable 5 a Major Step Forward for Long Tasks

    AIAndrej Karpathy says Claude Fable 5 is the same underlying model as Mythos with added safeguards, and that it leads on nearly all benchmarks. He describes it as a step change, especially for long, difficult problem-solving sessions where it handles more ambitious tasks without close supervision. He notes that its safeguards are set a bit too aggressively at launch and may be tuned over time.

Jun 8

Jun 8Mon
  1. Cognition Blog (Devin, Windsurf)AI score70

    Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

    AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

    Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

  2. Xiaomi MiMoAI score65

    Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model

    AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.

    Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

Jun 3

Jun 3Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition launches $10M AI Productivity Guarantee for enterprise Devin customers

    AICognition introduced the AI Productivity Guarantee, under which it will issue credits up to $10M if Devin delivers less engineering value than enterprise customers pay for. The company uses an AI estimator to measure hours of productive output, validated against engineers' own estimates of how long the same work would have taken by hand. Value is converted to dollars at a standard global rate and compared against each customer's consumption near the end of the annual contract.

    Why it matters: The post explains how Cognition estimates Devin's output in hours and backs the estimate with a $10M credit commitment, a concrete model for measuring AI vendor value.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition Estimates Engineering Hours Saved by Its Devin Coding Agent

    AICognition built an automated agent that classifies Devin sessions as productive and estimates the human engineering hours each one would have taken. On 233 held-out sessions the estimator reached an rlog of 0.74, with individual errors often 2 to 3 times in either direction but roughly unbiased in aggregate. The system is calibrated to underestimate and is currently running with Devin customers.

    Why it matters: The post shows how the measurement design, from hours-based metrics to conservative calibration, determines whether agent productivity estimates can be trusted in aggregate.

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 29

May 29Fri
  1. Fei-Fei LiAI score38

    Fei-Fei Li Highlights GPIC, a Permissive Image Corpus for Visual Generation

    AIFei-Fei Li praised GPIC, a new benchmark dataset for visual generation built for modern large-scale generative models. The corpus includes 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, totaling about 28 trillion pixels. It is centrally hosted and fully permissive for research and commercial use.

May 28

May 28Thu
  1. PaddlePaddleAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

    Image from @PaddlePaddle's post

May 26

May 26Tue

May 22

May 22Fri
  1. AI Snake OilAI score60

    Google's $916 agent-built operating system claim lacks key methodology details

    AIGoogle claimed a team of agents built an operating system from a single prompt for about $916 in API fees, using Gemini 3.5 Flash and Antigravity 2.0. The authors argue the prompt was many thousands of lines, the scaffold and human intervention are undefined, and no code, logs, or similarity analysis were released to verify the claim. They still see value in such open-world evaluations, which need stronger methodological norms and independent scrutiny.

May 21

May 21Thu
  1. Mark ChenAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

May 19

May 19Tue
  1. koray kavukcuogluAI score72

    Google's Gemini 3.5 Flash beats Gemini 3.1 Pro on coding and agentic benchmarks

    AIGoogle's Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%). The post also claims it is 4x faster than other frontier models, or 12x in Antigravity, and reports 83.6% on MMMU-Pro for multimodal performance.

    Why it matters: The post gives specific benchmark scores against Gemini 3.1 Pro, letting readers compare coding, agentic, and multimodal results directly.

    Image from @koraykv's post

May 15

May 15Fri
  1. Eugene YanAI score54

    Eugene Yan reviews Claude Mythos Preview exploit case study transcripts

    AIEugene Yan reviewed the Claude Mythos Preview transcripts to verify their legitimacy and check for reward-hacking behavior. He reports the model reasoned through a bug, tested hypotheses, debugged issues, and found ways to bypass the V8 sandbox, which he judged consistent with a competent browser and JavaScript engine security researcher. The case study cites CVE-2024-051912, an exploited bug with no public report or working PoC, which had resisted reproduction by researchers for a year.

May 13

May 13Wed
  1. Eugene YanAI score72

    Mythos completes 32-step network attack in six of ten UK AISI trials

    AIEugene Yan relays two evaluations of Mythos: UK AISI reports it completed a 32-step network attack, estimated at about 20 expert hours, in 6 of 10 tries and was the first model to solve its end-to-end cyber ranges. XBOW's evaluation describes its performance as token-for-token and unprecedented in precision. The post links both AISI and XBOW blog posts for details.

May 9

May 9Sat
  1. PaddlePaddleAI score60

    Baidu releases ERNIE 5.1 with reduced pretraining cost and parameter scale

    AIBaidu's PaddlePaddle account announced ERNIE 5.1, which it says cuts total parameters to about one-third and activated parameters to about one-half, using roughly 6% of the pretraining cost of models at similar scale. The post reports benchmark results including 99.6 on AIME26 with tools, surpassing DeepSeek-V4-Pro on τ3-bench and SpreadsheetBench-Verified, and ranking #4 globally on Arena Search. ERNIE 5.1 is available through the ERNIE website and Baidu AI Studio Model Playground.

May 8

May 8Fri
  1. Berkeley AI ResearchAI score46

    Adaptive Parallel Reasoning Lets Models Decide When to Parallelize Inference

    AIBerkeley AI Research describes adaptive parallel reasoning, in which a reasoning model decides when to split independent subtasks, how many concurrent threads to spawn, and how to coordinate them. The approach targets the latency, context-rot, and cost problems of long sequential reasoning, which can require millions of tokens and tens of minutes for complex tasks. Existing methods such as self-consistency, Tree of Thoughts, ParaThinker, and Hogwild! Inference fix the parallel structure outside the model, which wastes compute on simple problems.

May 7

May 7Thu
  1. Sam BowmanAI score38

    Anthropic donates open-source alignment testing tool Petri to Meridian Labs

    AIAnthropic is donating Petri, its open-source interactive behavioral-evals tool for alignment testing, to Meridian Labs so development can continue independently. Working with Meridian, Anthropic has also released a major update improving the adaptability, realism, and depth of Petri's tests. Developers are invited to try the tool and contribute.

May 4

May 4Mon
  1. HyperdimensionalAI score63

    Dean W. Ball argues against overreacting to Anthropic's Mythos cyber capabilities

    AIDean W. Ball argues that Anthropic's Mythos, which finds software vulnerabilities by chaining bugs into exploits, shifts the cost of vulnerability discovery and should not prompt an overreaction. He contends governments hold a uniquely mixed incentive over vulnerabilities, so heavy state control risks making software less secure. He proposes a narrow, testable government role focused on cyber-discovery risk thresholds, with private verification bodies supporting it.

Apr 30

Apr 30Thu
  1. Mark ChenAI score62

    OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range

    AIMark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture. He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations. The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.

  2. ARC PrizeAI score44

    GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models

    AIOpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs. The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic. ARC Prize is open-sourcing its analysis package.