Skip to content

#Model release

Aug 26

Aug 26Wed
  1. Ai2 · new models on Hugging Face38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    Ai2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  2. Ai2 · new models on Hugging Face37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    Ai2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  3. Ai2 · new models on Hugging Face38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    Ai2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  4. Ai2 · new models on Hugging Face38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    Ai2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Fireworks AI Blog40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    DeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Z.ai Release Notes62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    Z.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  3. Daniel Han34

    Yes, you can fine-tune Qwen3.8-27B completely for free by just having a Google account! 🦥 Kaggle, like Colab, provides 30 hours of free GPU with 2× Tesla T4s. Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.

    Yes, you can fine-tune Qwen3.8-27B completely for free by just having a Google account! 🦥 Kaggle, like Colab, provides 30 hours of free GPU with 2× Tesla T4s. Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.

  4. Z.ai (GLM) · new models on Hugging Face72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    Z.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  5. Z.ai (GLM) · new models on Hugging Face72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    Z.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging Face40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    Google Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Microsoft Research34

    Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5

    Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5

  3. Generalist38

    We've reduced the time it takes to go from physical prompt → robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.

    We've reduced the time it takes to go from physical prompt → robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.

  4. Qwen · new models on Hugging Face75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    Qwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.

Aug 21

Aug 21Fri
  1. Thinking Machines36

    We want to improve Inkling’s agentic performance. To help us understand its real-world behavior, we are making it available for free on OpenRouter (only with agentic harnesses) for the next few weeks, starting now. We’ll use the data, disassociated from accounts, to better it.

    We want to improve Inkling’s agentic performance. To help us understand its real-world behavior, we are making it available for free on OpenRouter (only with agentic harnesses) for the next few weeks, starting now. We’ll use the data, disassociated from accounts, to better it.

  2. Z.ai28

    The community has already built so many interesting projects with ZCode + GLM-5.3. To thank everyone for all the support, we’re turning Build Week into an ongoing series. From now to Aug 23 at 6 PM PT, we’re giving 50,000 new ZCode users 100M free GLM-5.3 tokens each.

    The community has already built so many interesting projects with ZCode + GLM-5.3. To thank everyone for all the support, we’re turning Build Week into an ongoing series. From now to Aug 23 at 6 PM PT, we’re giving 50,000 new ZCode users 100M free GLM-5.3 tokens each.

  3. DeepSeek62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    DeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

  4. DeepSeek API News60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    DeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu
  1. swyx38

    proud that @vibhuuuus and i did the most recent pod with @eisokant on why NVIDIA just paid him $6B to buy the incredible model factory that is pumping out Thinky-beating models (actually not exaggeration, look at the numbers) tune in / subscribe / whatever https://www.latent.space/p/poolside only on @latentspacepod

    proud that @vibhuuuus and i did the most recent pod with @eisokant on why NVIDIA just paid him $6B to buy the incredible model factory that is pumping out Thinky-beating models (actually not exaggeration, look at the numbers) tune in / subscribe / whatever https://www.latent.space/p/poolside only on @latentspacepod

Aug 19

Aug 19Wed
  1. Liquid AI Blog60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    Liquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

  2. Generalist42

    To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog: https://generalistai.com/blog/gen-1.5

    To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog: https://generalistai.com/blog/gen-1.5

  3. Daniel Han40

    We also released 1-bit Qwen3.8-27B quants which run on 8GB RAM. They retain ~77% of accuracy compared to BF16. We originally didn't want to release them but were shocked that they worked extremely well on our internal testing!

    We also released 1-bit Qwen3.8-27B quants which run on 8GB RAM. They retain ~77% of accuracy compared to BF16. We originally didn't want to release them but were shocked that they worked extremely well on our internal testing!

  4. Google · new models on Hugging Face22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    Google has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  5. Google · new models on Hugging Face26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    Google has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

  6. Daniel Han40

    New Qwen3.8-27B GGUFs with Unsloth Dynamic v3! Extra 10% top-1% accuracy at the same size! Divergence-300 is a new metric which extends top-1% greedy acc to 32 tokens & more, and we used 300 unseen examples from Terminal Bench, DeepSWE and more We also made 6-8GB 1-bit quants!

    New Qwen3.8-27B GGUFs with Unsloth Dynamic v3! Extra 10% top-1% accuracy at the same size! Divergence-300 is a new metric which extends top-1% greedy acc to 32 tokens & more, and we used 300 unseen examples from Terminal Bench, DeepSWE and more We also made 6-8GB 1-bit quants!