Skip to content

#Open-source ecosystem

Oct 7

Oct 7Wed
  1. Lucas BeyerAI score36

    Missed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.

    Missed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.

  2. GitHub Copilot ChangelogAI score42

    GitHub Copilot CLI adds discovery of local Ollama models via /model

    GitHub Copilot CLI version 1.0.94-0 lets users run /model to discover supported models from a running local Ollama instance alongside configured and GitHub Copilot cloud models. Discovered models are not added automatically; users choose one, review its provider and endpoint, then confirm Add and use for this session or Add without switching, and models must support tool calling and streaming. Choosing a local model does not enable offline mode or disable GitHub telemetry, and COPILOT_OFFLINE=true remains a separate explicit setting.

  3. Andrew CurranAI score27

    Nous Research has raised $90 million at a $1.5 billion valuation, I know a lot of people from the team, and they are a really great group. Teknium helped me when I was starting out here. Hermes has been a huge success for them.

    Nous Research has raised $90 million at a $1.5 billion valuation, I know a lot of people from the team, and they are a really great group. Teknium helped me when I was starting out here. Hermes has been a huge success for them.

  4. Elvis SaraviaAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    NVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

  5. Mistral AIAI score22

    Gmail. Google Sheets. Slack. Salesforce. i.e. the four horsemen of Monday. On AutomationBench, Mistral Large 4 took on 657 business workflows across simulated workplace apps, all four included, finishing ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.

    Gmail. Google Sheets. Slack. Salesforce. i.e. the four horsemen of Monday. On AutomationBench, Mistral Large 4 took on 657 business workflows across simulated workplace apps, all four included, finishing ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro.

  6. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    NVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    AIWhy it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  7. AI SupremacyAI score44

    Reflection AI's Beam and Mistral Large 4 advance Western open-source models

    Reflection AI announced Beam, a model trained end-to-end from scratch that appears to advance the Western open frontier on coding and agentic tasks. Mistral then released Mistral Large 4, a 1 trillion-parameter natively multimodal model with 49 billion active parameters, though the piece says neither model yet matches leading Chinese open-weight models.

  8. O'Reilly RadarAI score42

    Build Your Own Post-Training Pipeline: SFT, Reward Model, and PPO

    The final post in O'Reilly Radar's four-part post-training series walks readers through implementing the classic ChatGPT pipeline on Qwen2.5-1.5B, covering SFT, reward model training, and PPO. The walkthrough uses torchtune for SFT and verl, a Ray-based RL framework from ByteDance's team, for reinforcement learning. The author says the goal is hands-on understanding rather than reproducing InstructGPT, which took a large team and thousands of GPU-hours.

  9. The Register · AIAI score38

    COSMIC bans AI-generated contributions as GNOME debates accepting AI bug reports

    System76's COSMIC desktop now requires contributors to declare no LLM-generated content in pull requests, including code, comments, and descriptions. GNOME Calendar and GNOME Extensions also restrict AI-generated contributions, while GNOME developer Michael Catanzaro argues the project should accept AI-generated bug reports. Catanzaro's case rests on memory-unsafe languages such as C, C++, and Vala, and he has shortened GNOME Security's disclosure deadline from 90 days to 30, effective August 1.

  10. Ai2 (Allen Institute for AI)AI score57

    Ai2's Bolmo byte-level language models are published in Nature

    Ai2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

  11. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    OpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

Oct 6

Oct 6Tue
  1. Lewis TunstallAI score25

    This is the most important plot from the Beam release IMO. The Chinese models are great, but horribly token inefficient (try running an eval with max reasoning to feel the pain). I'm looking forward to a future where open models start competing on this axis!

    This is the most important plot from the Beam release IMO. The Chinese models are great, but horribly token inefficient (try running an eval with max reasoning to feel the pain). I'm looking forward to a future where open models start competing on this axis!

  2. Epoch AIAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    Epoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  3. vLLM BlogAI score62

    vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations

    Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks. Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.

    AIWhy it matters: The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.

  4. Alex HeathAI score42

    I really liked Reflection CEO @MishaLaskin's framing on the state of closed vs. open AI models: "Closed models are the equivalent in real estate to renting an apartment... As AI adoption has increased, an ownership market is basically coming in.. The only way to own intelligence, by definition, is if it's open"

    I really liked Reflection CEO @MishaLaskin's framing on the state of closed vs. open AI models: "Closed models are the equivalent in real estate to renting an apartment... As AI adoption has increased, an ownership market is basically coming in.. The only way to own intelligence, by definition, is if it's open"

  5. OpenAIAI score62

    OpenAI releases new mathematical results from an internal frontier model

    OpenAI is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and drew on its advice and public recommendations for how the results are released. The results are available at https://github.com/openai/math.

  6. TekniumAI score33

    We just launched Hermes Index! This combines the scores of our new HermesBench and 3 other leading and relevant benchmarks for agents to give every Hermes Agent user a way to find both the best model at a given time, as well as the best model at a given price point! Check it out

    We just launched Hermes Index! This combines the scores of our new HermesBench and 3 other leading and relevant benchmarks for agents to give every Hermes Agent user a way to find both the best model at a given time, as well as the best model at a given price point! Check it out

  7. OllamaAI score55

    Google DeepMind's EmbeddingGemma 2 is now available on Ollama

    Ollama announced that Google DeepMind's EmbeddingGemma 2 is now available on Ollama. The author describes it as made for consumer devices and multimodal, and gives the command ollama pull embeddinggemma-2 to download it. The quoted DeepMind post says the model is a natively multimodal open model for on-device embeddings that unifies code, images, audio, and video in a shared space.

  8. AMDAI score18

    Open models need an open stack. @ZyphraAI VP of AI Engineering Quentin Anthony shares how access to open software libraries and direct collaboration with AMD are helping the team train bigger models, keep growing compute resources working efficiently and pave new ground in AI. Learn how Zyphra trained its ZAYA1-8B reasoning model from scratch on a full-stack AMD platform: https://bit.ly/4ixfdaB

    Open models need an open stack. @ZyphraAI VP of AI Engineering Quentin Anthony shares how access to open software libraries and direct collaboration with AMD are helping the team train bigger models, keep growing compute resources working efficiently and pave new ground in AI. Learn how Zyphra trained its ZAYA1-8B reasoning model from scratch on a full-stack AMD platform: https://bit.ly/4ixfdaB

  9. NVIDIA Technical BlogAI score37

    Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

    NVIDIA's technical blog describes bitwise determinism for large-scale pretraining with Megatron Core, which makes training runs easier to debug, validate, and resume reproducibly. The source says these benefits matter most for models with trillions of parameters trained across thousands of GPUs, where multiple parallelism dimensions, low-precision computation, and distributed checkpointing complicate failure reproduction and fix validation.

  10. MiniMaxAI score12

    MiniMax hosts AI events at SF Tech Week with partner companies

    MiniMax is taking part in SF Tech Week with a series of events on October 6, 7, and 8, featuring partners including Friendli.AI, Anaconda, Kilo Code, Novita AI, Artificial Analysis, Nous Research, RadixArk, Vercel, Fireworks AI, DigitalOcean, Modular, and Evermind. The programming covers frontier models, high-speed inference, agents, and open-source AI stacks, plus a Magnific-hosted talk on growing creative AI products.

  11. vLLMAI score60

    vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings

    vLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.

  12. Clément DelangueAI score62

    Mistral Large 4 announced with API access today and open weights due end of October

    Mistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active parameters. It is available via API today, with open weights planned for the end of October. Clément Delangue, Hugging Face's CEO, reacted by noting that the model cannot be the best open-weight model until its weights are actually released.

  13. Yuchen JinAI score34

    Exciting to see Reflection’s Beam and Mistral Large 4 both reach roughly GLM-5.2 level in the past two days. Makes me wonder if the real Western vs. Chinese OSS models gap is simply this: Chinese labs can distill Anthropic and OpenAI models. Western labs can’t.

    Exciting to see Reflection’s Beam and Mistral Large 4 both reach roughly GLM-5.2 level in the past two days. Makes me wonder if the real Western vs. Chinese OSS models gap is simply this: Chinese labs can distill Anthropic and OpenAI models. Western labs can’t.

  14. Merve NoyanAI score25

    releasing my local AI slide deck covers from prefill vs decode, MoE vs dense, VRAM vs unified memory, quantization to speculative decoding, everything with llama.cpp feel free to reuse with attribution

    releasing my local AI slide deck covers from prefill vs decode, MoE vs dense, VRAM vs unified memory, quantization to speculative decoding, everything with llama.cpp feel free to reuse with attribution

  15. Google for DevelopersAI score36

    — Weights are live on @HuggingFace and @Kaggle — Get out-of-the-box support for LiteRT, MediaPipe, @LangChain, and @llama_index — Start Mapping: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

    — Weights are live on @HuggingFace and @Kaggle — Get out-of-the-box support for LiteRT, MediaPipe, @LangChain, and @llama_index — Start Mapping: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

  16. Google DeepMindAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    Google DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

  17. Sundar PichaiAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.