Skip to contentSkip to stories

Updated

#Model release

Jun 9

Jun 9Tue
  1. Andrej KarpathyAI score65

    Karpathy Calls Claude Fable 5 a Major Step Forward for Long Tasks

    AIAndrej Karpathy says Claude Fable 5 is the same underlying model as Mythos with added safeguards, and that it leads on nearly all benchmarks. He describes it as a step change, especially for long, difficult problem-solving sessions where it handles more ambitious tasks without close supervision. He notes that its safeguards are set a bit too aggressively at launch and may be tuned over time.

  2. Z.ai (GLM) · new models on Hugging FaceAI score52

    Z.ai releases SCAIL-2, an open-source end-to-end character animation model

    AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.

Jun 8

Jun 8Mon
  1. Xiaomi MiMoAI score62

    Xiaomi MiMo-V2.5-Pro UltraSpeed claims 1,000+ tokens/s on a 1T model

    AIXiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, which the post says reaches output speeds above 1,000 tokens/s on a 1 trillion parameter MoE model. The post says this runs on a single standard 8-GPGPU node rather than wafer-scale or pure on-chip SRAM hardware. UltraSpeed access is application-based from Jun 8 to Jun 23 (PDT), and the UltraSpeed API costs 3x the standard price.

  2. ByteDance · new models on Hugging FaceAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

  3. Xiaomi MiMo · new models on Hugging FaceAI score41

    Xiaomi releases MiMo-V2.5-Pro-FP4-DFlash, an FP4 model with block-diffusion decoding

    AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.

  4. Xiaomi MiMoAI score65

    Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model

    AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.

    Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

  3. ByteDance · new models on Hugging FaceAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat

May 28

May 28Thu
  1. PaddlePaddleAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

May 26

May 26Tue

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens

    AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.

    Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 22

May 22Fri

May 21

May 21Thu
  1. Mark ChenAI score92

    OpenAI model disproves Erdős's unit distance conjecture in planar geometry

    AIAn OpenAI model disproved Erdős's longstanding planar unit distance conjecture, which Paul Erdős posed in 1946, by discovering a new family of constructions that performs better than the square grids mathematicians had long assumed. Mark Chen says the proof draws on algebraic number theory and describes it as the first time AI has autonomously solved a prominent open problem central to a field of mathematics.

    Why it matters: The post names the specific open problem and the approach used, giving readers a concrete case of AI producing a research proof in mathematics.

May 20

May 20Wed
  1. Stability AIAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

May 19

May 19Tue
  1. koray kavukcuogluAI score72

    Google's Gemini 3.5 Flash beats Gemini 3.1 Pro on coding and agentic benchmarks

    AIGoogle's Gemini 3.5 Flash outperforms Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%). The post also claims it is 4x faster than other frontier models, or 12x in Antigravity, and reports 83.6% on MMMU-Pro for multimodal performance.

    Why it matters: The post gives specific benchmark scores against Gemini 3.1 Pro, letting readers compare coding, agentic, and multimodal results directly.

  2. koray kavukcuogluAI score72

    Google rolls out Gemini 3.5 Flash globally across consumer, developer, and enterprise platforms

    AIGoogle is rolling out Gemini 3.5 Flash globally for consumers in the Gemini app and Search AI Mode. It is also available to developers through the Gemini API, Google Antigravity, and Google AI Studio, and to businesses on the Gemini Enterprise Agent Platform.

    Why it matters: The post shows where each Gemini 3.5 Flash access path goes, from consumer apps to developer and enterprise platforms, which helps readers pick the right entry point.

  3. koray kavukcuogluAI score62

    Google introduces Gemini 3.5 Flash, used with agents to rebuild AlphaZero

    AIAt Google I/O, Google introduced Gemini 3.5 Flash, which the author says has become part of the daily research cycle. The author says a team of agents in Antigravity 2.0 recreated the original AlphaZero paper and built a playable web version from two prompts, coding the reinforcement learning pipeline in JAX/Flax and training a ResNet model via self-play on multi-TPU pods.

May 18

May 18Mon
  1. Michael TruellAI score40

    Cursor's Composer 2.5 is a significant upgrade over Composer 2

    AIMichael Truell of Cursor says Composer 2.5 is a significant step up from Composer 2. He adds that this is only the start of work with SpaceXAI, with more improvements expected soon. Cursor's announcement describes the model as more intelligent, better at long-running tasks, and more reliable at complex instructions, with doubled included usage for the next week.

May 15

May 15Fri
  1. Intern Large ModelsAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

May 11

May 11Mon

May 9

May 9Sat
  1. PaddlePaddleAI score60

    Baidu releases ERNIE 5.1 with reduced pretraining cost and parameter scale

    AIBaidu's PaddlePaddle account announced ERNIE 5.1, which it says cuts total parameters to about one-third and activated parameters to about one-half, using roughly 6% of the pretraining cost of models at similar scale. The post reports benchmark results including 99.6 on AIME26 with tools, surpassing DeepSeek-V4-Pro on τ3-bench and SpreadsheetBench-Verified, and ranking #4 globally on Arena Search. ERNIE 5.1 is available through the ERNIE website and Baidu AI Studio Model Playground.

May 6

May 6Wed

Apr 28

Apr 28Tue

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  2. Mistral AI · new models on Hugging FaceAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.