Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
Ai2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.
Ai2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.
Ai2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.
Ai2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.
Ai2 has released Bolmo-1B-Stage1, a 1B-parameter byte-level language model retrofitted from OLMo 2 1B to process bytes rather than tokens. This checkpoint includes Stage 1 training only, with inner model parameters unchanged, and is available on Hugging Face under an Apache 2.0 license for research and educational use.
Ai2 has released Bolmo-7B-Stage1, a 7B byte-level language model retrofitted from Olmo 3 7B through a short additional training procedure. This checkpoint includes only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.
DeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).
Harvey and Fireworks post-trained Tenet from the Kimi K3 base using asynchronous reinforcement learning on the Fireworks Training API for long-horizon legal work. On the Legal Agent Benchmark, Tenet reached 19.7% all-pass versus 10.8% for base Kimi K3, and its cost per task was $5.92 versus $5.62.
Z.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.
Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.
Yes, you can fine-tune Qwen3.8-27B completely for free by just having a Google account! 🦥 Kaggle, like Colab, provides 30 hours of free GPU with 2× Tesla T4s. Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.
Qwen announced Qwen3.8-Flash-Next, an open-weight multimodal MoE model built on the new Qwen4 architecture. The model will be released tomorrow, and Unsloth is preparing day-zero support.
Z.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.
Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.
Z.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.
Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.
real question: why does gpt 5.6 over engineer so much?
Google Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.
Try here: https://replicate.com/alibaba/wan-3
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5
Read more about GEN-1.5, our latest foundation model for the physical world: https://generalistai.com/blog/gen-1.5
We've reduced the time it takes to go from physical prompt → robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.
LongCat-2.0 just landed in Go on @opencode 🐱 Give it a try and let us know what you build!
Qwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.
Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.
Sundar Pichai says Gemini 3.7 Flash set new Gemini growth records in its first week, making it the fastest-growing Gemini model so far. The model is now running in Search and the Gemini app. A quoted ARC-AGI post reports 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12 per task.
We want to improve Inkling’s agentic performance. To help us understand its real-world behavior, we are making it available for free on OpenRouter (only with agentic harnesses) for the next few weeks, starting now. We’ll use the data, disassociated from accounts, to better it.
Try out Inkling and Inkling-Small on OpenRouter here: https://openrouter.ai/provider/thinkingmachines
The community has already built so many interesting projects with ZCode + GLM-5.3. To thank everyone for all the support, we’re turning Build Week into an ongoing series. From now to Aug 23 at 6 PM PT, we’re giving 50,000 new ZCode users 100M free GLM-5.3 tokens each.
Try it today on MAI Playground: https://msft.it/6017aMcTl
Bring your mood board to life with MAI-Image-2.6. Our latest image model can modify color, style, and visual details while maintaining consistency across iterations.
Multimodality unlocks more agent use cases. 👀 V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows. 2/n
DeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.
DeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.
Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.
proud that @vibhuuuus and i did the most recent pod with @eisokant on why NVIDIA just paid him $6B to buy the incredible model factory that is pumping out Thinky-beating models (actually not exaggeration, look at the numbers) tune in / subscribe / whatever https://www.latent.space/p/poolside only on @latentspacepod
Liquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.
Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.
GPT 5.6 Sol & Fast mode are 50% off on v0. Available until September 18, 2026. Try it https://v0.app/?sol
To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog: https://generalistai.com/blog/gen-1.5
We also released 1-bit Qwen3.8-27B quants which run on 8GB RAM. They retain ~77% of accuracy compared to BF16. We originally didn't want to release them but were shocked that they worked extremely well on our internal testing!
Google has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.
Google has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.
New Qwen3.8-27B GGUFs with Unsloth Dynamic v3! Extra 10% top-1% accuracy at the same size! Divergence-300 is a new metric which extends top-1% greedy acc to 32 tokens & more, and we used 300 unseen examples from Terminal Bench, DeepSWE and more We also made 6-8GB 1-bit quants!
Google released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.