Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. Ai2OfficialAI score36

    Ai2 adapts existing models to bytes with a short training run

    AIAi2 describes a recipe that adapts existing models to process raw bytes through a brief additional training run. The approach keeps the model's core intact and adds components that group bytes into variable-length patches for processing, then expand outputs back to byte-level predictions of the next byte.

  2. Ars Technica · AINewsAI score63

    Mistral releases Le Chonk, a 1 trillion-parameter open-weight model

    AIMistral has released Mistral Large 4, nicknamed Le Chonk, a 1 trillion-parameter model it says can be used and customized by anyone. It is in preview, with a final version due by the end of the month, and is optimized for coding and cyberdefense as well as manufacturing, finance, and electrical engineering tasks. Mistral claims it is the most capable open-weight model developed outside China and says it was trained from scratch rather than through distillation.

  3. Hugging Face BlogOfficialAI score53

    TII releases Falcon-ASR, a 1.6B speech recognition model focused on Emirati Arabic

    AIThe Technology Innovation Institute introduces Falcon-ASR, a 1.6 billion parameter speech recognition model for Arabic with a focus on the Emirati dialect. On six Arabic test sets it reports an average word error rate of 20.92%, versus 23.17% for the best published leaderboard result it compared against. The model also transcribes English, French, Spanish and Portuguese with the same weights, and a demo Space is available while API access and native apps are planned.

  4. 🚨 AI News | TestingCatalogXAI score47

    Daily AI brief covers Mistral Large 4, Google, OpenAI, and Anthropic updates

    AIMistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks. Google rolled out Nano Banana 2.1 across Gemini, AI Studio, and the Gemini API, and released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0. OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.

  5. Ai2 (Allen Institute for AI)OfficialAI score57

    Ai2's Bolmo byte-level language models are published in Nature

    AIAi2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

  6. Latent SpaceBlogAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

  7. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  8. Artificial Analysis ArticlesOfficialAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    Why it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.

Oct 6

Oct 6Tue
  1. Josh WoodwardOfficialAI score34

    Nano Banana 2.1 adds mask-based editing and improved visual quality

    AIGoogle's Nano Banana 2.1 is an upgraded image model that outperforms prior versions in visual design, mask-based editing, subject consistency, and natural-looking imagery. Josh Woodward calls mask-based editing his favorite feature from the launch and says more is coming soon.

  2. meng shaoXAI score62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.

    Image from @shao__meng's post
  3. SpaceXAIOfficialAI score38

    Grok 4.7 is now live on Microsoft Foundry

    AIGrok 4.7 is now available on Microsoft Foundry. The post announces the model's availability on the platform without additional details on features, pricing, or benchmarks.

    Video from @SpaceXAI's post
  4. Liquid AI BlogOfficialAI score62

    Liquid AI releases open d1-3B and d1-omni-600M decision models for edge devices

    AILiquid AI released two open-weight d1 decision models, d1-3B and d1-omni-600M, on Hugging Face. d1-3B scores 48.57 on the Decision Index v0.2.1 public split and answers a single question in 8 ms on an NVIDIA GeForce RTX 4090 and 50 ms on a Jetson Orin Nano. d1-omni-600M is an experimental checkpoint that handles text with images or audio and scores 15.95 on the same index.

    Why it matters: The release pairs open-weight decision models with measured latency across Apple, NVIDIA, and Jetson hardware, showing how edge deployment changes what is practical.

  5. Claude Apps Release NotesOfficialAI score60

    Claude Haiku 5.5 launches as a fast, low-cost small model, and Max and Team plans gain monthly API credits

    AIAnthropic launched Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released, aimed at high-volume, cost-sensitive tasks. Max and Team plans now include monthly API credits for running their own apps and agents on the Claude Platform, rolling out over a few days. Users claim the credits by linking a Claude Console organization in Settings > Billing for Max or Organization settings > Billing for Team.

    Why it matters: The notes name a new small model and a credit change for Max and Team plans, with the claim path, which matters for teams budgeting API use.

  6. Comfy BlogOfficialAI score43

    Gemini Nano Banana 2.1 Is Now Available via ComfyUI Partner Nodes

    AIGoogle's Gemini Nano Banana 2.1 image generation and editing model is now available through ComfyUI Partner Nodes, succeeding Nano Banana 2 with a balance of price and performance. It accepts up to 14 reference images, outputs images up to 4K, and offers Minimal, Medium, and High thinking levels plus search grounding and 9:21 aspect ratio support.

  7. ollamaOfficialAI score55

    Google DeepMind's EmbeddingGemma 2 is now available on Ollama

    AIOllama announced that Google DeepMind's EmbeddingGemma 2 is now available on Ollama. The author describes it as made for consumer devices and multimodal, and gives the command ollama pull embeddinggemma-2 to download it. The quoted DeepMind post says the model is a natively multimodal open model for on-device embeddings that unifies code, images, audio, and video in a shared space.

  8. OpenAI DevelopersOfficialAI score22

    OpenAI's GPT-6 Luna adds predicate, choice, and score outputs

    AIOpenAI's GPT-6 Luna accepts text and image inputs and supports three output types: predicates that estimate the probability a statement is true, choices that select from predefined options with confidence scores, and scores that evaluate an input against a numeric range.

    Video from @OpenAIDevs's post
  9. Aravind SrinivasXAI score40

    Perplexity halves Decision API input pricing to $0.02 per million tokens

    AIPerplexity cut its Decision API input pricing by half, to $0.02 per million input tokens. The reduction follows the release of pplx-decider-v1.1-27b, an open-weights multimodal decision model that scores highest on Hugging Face's Decision Index 0.3 benchmark. The model costs half as much as v1.

  10. Google DeepMindOfficialAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    Why it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  11. Philipp SchmidXAI score70

    EmbeddingGemma 2 releases native multimodal embeddings built on Gemma 4

    AIGoogle releases EmbeddingGemma 2, its first native multimodal embedding model, built on Gemma 4 under Apache 2.0. It embeds over 100 languages, code, images, audio, and video into one vector, with an 8,192-token context and four sizes from 270M to 740M parameters. Matryoshka output dimensions of 768, 512, 256, or 128 are supported, and the model is available in Sentence Transformers and LiteRT-LM, with a reported 14% gain on MTEB Code.

    Why it matters: The release extends an embedding model to text, code, images, audio, and video in one vector, a useful option for retrieval systems that mix media types.

  12. vLLMOfficialAI score60

    vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings

    AIvLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.

    Image from @vllm_project's post
  13. ZDNet · AINewsAI score60

    Google limits free Gemini users to Flash-Lite model from October 9

    AIStarting October 9, free Google Gemini users will only have access to the Flash-Lite model, and AI Plus subscribers will lose Pro access. Google is also moving to compute-based limits that refresh every five hours, with higher limits for paid plans. The AI Pro plan at $20 per month will gain access to the Deep Think reasoning mode.

  14. ReplicateOfficialAI score29

    Nano Banana 2.1 image model now live on Replicate

    AIReplicate has made Nano Banana 2.1, Google DeepMind's latest image model, available, optimized for a balance of price and performance. The model offers improved visual design, mask-based editing, and subject consistency for more natural-looking images.

    Image from @replicate's post
  15. Nano Banana 2.1OfficialAI score40

    Nano Banana 2.1 released with gains in design, editing, and consistency

    AIGoogle announces Nano Banana 2.1, an upgraded image model that outperforms its previous versions across the board. The company says it brings notable improvements in visual design, mask-based editing, subject consistency, and more natural-looking imagery. It is available now in the Gemini app and Google AI Studio.

    Image from @NanoBanana's post
  16. Yuchen JinXAI score34

    Reflection's Beam and Mistral Large 4 near GLM-5.2 level

    AIYuchen Jin says Reflection's Beam and Mistral Large 4 both reached roughly GLM-5.2 level within the past two days. He suggests the Western versus Chinese open-source model gap may come down to Chinese labs being able to distill Anthropic and OpenAI models, which Western labs cannot.