Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Sep 15

Sep 15Tue
  1. NVIDIA · new models on Hugging FaceOfficialAI score34

    NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time RGB Hand Localization

    AINVIDIA's RT-DETR Hand Detection v1.0 detects and localizes left and right hands in RGB images, outputting 2D bounding boxes with per-hand confidence scores in a single pass. The model, built on RT-DETRv2-S with HGNetv2-S backbone and about 20M parameters, is intended as a region-of-interest stage for downstream 3D hand pose estimation and is exported to ONNX. The source describes it as for demonstration purposes rather than production use, runs on NVIDIA Lovelace GPUs under Linux, and is licensed under the NVIDIA Software and Model Evaluation License.

  2. RadixArkOfficialAI score42

    Periodic Labs builds Neon on SGLang and Miles for 2.5x faster inference

    AIPeriodic Labs chose SGLang and Miles to build Neon, an open-source model it says surpasses GPT-6 Astra on its analysis benchmark after mid-training and RL on 1,300 H200s. RadixArk says Periodic extended both frameworks for scientific RL at trillion-parameter scale, delivering more efficient training, lower memory use, and 2.5x faster inference. The work has been contributed back to both projects.

  3. OdysseyOfficialAI score22

    Odyssey-3 aims to enable physical agents that interact with the world

    AIOdyssey says its new Odyssey-3 model could enable physical agents, a new kind of agent that interfaces natively with physical and virtual systems. The company expressed excitement about the promise its models are showing, linking to an introduction page for Odyssey-3.

  4. OdysseyOfficialAI score38

    Odyssey unveils Odyssey-3, a foundation world model for robotics and more

    AIOdyssey announced Odyssey-3, a foundation world model it describes as a major step forward. The post claims it can control robots, power humanoids, drive cars, train AIs, pilot drones, and play video games, though it gives no benchmarks or technical specifications.

    Video from @odysseyml's post
  5. Baseten BlogOfficialAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  6. Gemini API ChangelogOfficialAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 14

Sep 14Mon
  1. Intern Large ModelsOfficialAI score23

    Intern-S2-397B, a scientific multimodal model, gets SGLang Day-0 support

    AISGLang announces Day-0 support for Intern-S2-397B from Intern Large Models, a 397B multimodal foundation model built for scientific intelligence and long-horizon agents. The model is pre-trained directly on raw scientific literature pages without parsing and uses reinforcement learning across more than 20 scientific domains, from biomolecule design to material generation. It also applies black-box agentic reinforcement learning in large-scale sandboxed environments.

  2. NVIDIA · new models on Hugging FaceOfficialAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  3. NVIDIA · new models on Hugging FaceOfficialAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

  4. Google · new models on Hugging FaceOfficialAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model

    AIGoogle DeepMind released EmbeddingGemma 2, an open model under Apache 2.0 that maps text, images, video, and audio into one shared 768-dimensional vector space. The model has 740M total parameters and supports 8,192-token context, with Matryoshka truncation to 128d, 256d, and 512d. The source reports 14% better code-task performance than EmbeddingGemma 1 and says it is designed for consumer hardware such as phones and laptops.

    Why it matters: The release combines text, image, video, and audio retrieval in one 768-dimensional space at 740M parameters, a useful reference for on-device multimodal search design.

  5. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

    Why it matters: The post names the benchmark scores and training scope behind Intern-S2-397B, letting readers compare its scientific and agentic claims against the table.

  6. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Why it matters: The post pairs a new open multimodal model with benchmark tables against named Qwen, DeepSeek, Kimi, GLM, GPT, Gemini, and Claude models, letting readers compare scientific and agentic results directly.

    Image from @intern_lm's post
  7. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

  8. Sherwin WuXAI score40

    GPT Image 2.5 now live in ChatGPT with better editing consistency

    AIGPT Image 2.5 is now available in ChatGPT, and the post says it keeps consistency well while editing images. It also claims state-of-the-art results on all image leaderboards, and suggests users who saw faces shift in GPT Image 2 edits try again.

    Video from @sherwinwu's post

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceOfficialAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    Why it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score38

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.8-27B

    AIinclusionAI has released Qwen3.8-27B-singprobe, a 10.1M-parameter intrinsic streaming guardrail that reuses Qwen3.8-27B hidden states to score query intent, response unsafety, and hallucination risk at every token. The probe adds less than 0.5% decode-time overhead and reports a 0.03% benign-response false-positive rate averaged across five datasets. It is supported through SGLang and vLLM integration branches, with training code available at inclusionAI/SingProbe.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  5. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    SingProbe: inclusionAI releases streaming safety probe for gpt-oss-120b

    AIinclusionAI released SingProbe, a 5.8M-parameter intrinsic guardrail built on openai/gpt-oss-120b that scores query intent, response unsafety, and hallucination risk at every token. It reuses the base model's hidden states, adding less than 0.5% decode-time overhead, and reports a 0.06% benign-response false-positive rate. The probe is available on Hugging Face and supported through SGLang and vLLM integrations.

  6. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

Sep 12

Sep 12Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score58

    Shanghai AI Lab releases Intern-S2-397B, a 397B multimodal scientific model

    AIShanghai AI Lab's InternLM team released Intern-S2-397B, a multimodal foundation model for scientific intelligence and long-horizon agents. The model uses visual pre-training on raw scientific literature pages, multi-task reinforcement learning across more than 20 scientific domains, and agentic reinforcement learning in sandboxed environments.

Sep 11

Sep 11Fri
  1. Baseten BlogOfficialAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  2. ChatGPTOfficialAI score4

    GPT-VI ASTRA: OpenAI's new model name announced

    AIThe official ChatGPT account posted the single line "GPT-VI ASTRA," naming a new model without giving any details. The post offers no information on capabilities, pricing, or availability.

    Video from @ChatGPT's post
  3. Dwarkesh PatelXAI score18

    Dwarkesh Patel on why Sonnet 5 and Opus 5 trail GLM 5.3

    AIDwarkesh Patel said a discussion with John, Beren, and Charlie questioned why Sonnet 5 and Opus 5 feel weaker than GLM 5.3, even though Anthropic could use raw logit distillation from Fable and train on Fable's environments. The discussion raised questions about the value of distillation, what makes it effective, and which model behaviors are hard to extract through it.

    Video from @dwarkesh_sp's post
  4. Cognition Blog (Devin, Windsurf)OfficialAI score51

    Cognition introduces Fusion in Devin Desktop and CLI for lower-cost coding

    AICognition is making Fusion available in Devin Desktop and CLI, a harness where a frontier lead model plans and reviews while a cheaper sidekick executes. Across listed coding benchmarks, Cognition reports Fusion cuts cost per task by about 11% to 46% versus the lead model alone, while the sidekick does the implementation work. The post recommends pairing Fable 5.1 with SWE-2, and argues price per task matters more than price per token.

  5. BAAIOfficialAI score46

    BAAI unveils AREX, a 122B MoE research agent for hard search

    AIBAAI introduced AREX, a research agent built on a 122B-parameter mixture-of-experts model with 10B active parameters. It drafts candidate answers, checks each constraint, and revisits unresolved points rather than running one long search. The post says AREX performs on hard search benchmarks comparable to GPT-5.4.

    Video from @BAAIBeijing's post
  6. LM StudioOfficialAI score40

    LM Studio Bionic now live with DeepSeek-V4.1-Flash support

    AILM Studio has launched its Bionic version, now live for users. The announcement is tied to DeepSeek-V4.1-Flash, the smallest model in DeepSeek's new architecture family, which adds native visual understanding and targets faster inference and higher throughput.

  7. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

  8. ollamaOfficialAI score60

    DeepSeek-V4.1-Flash becomes available on Ollama's cloud

    AIOllama says DeepSeek-V4.1-Flash is now fully rolled out on its cloud, hosted in the US and Europe. Prompts and responses are not logged or trained on, and per-token pricing matches the DeepSeek API, including off-peak pricing. The post repeats DeepSeek's claim that the model is more capable, faster, and more cost effective than prior DeepSeek models, including DeepSeek-V4-Pro.

    Why it matters: The post names the hosting regions, zero data retention policy, and pricing parity with the DeepSeek API, which matter for teams weighing cloud access to this model.

Sep 10

Sep 10Thu
  1. hardmaruXAI score52

    Sakana Fugu releases Fugu Max and Fugu Ultra v2 multi-agent orchestration models

    AISakana AI released Fugu Max and Fugu Ultra v2, multi-agent orchestration systems that route tasks across a pool of open-weights and specialized models. The source says Fugu Max delivers performance within striking distance of elite models at two to six times lower cost, while Fugu Ultra v2 outperforms Opus 5 and Fable 5 on Chartography and outperforms models costing three to five times more per token on DeepSWE.

    Image from @hardmaru's post
  2. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  3. Understanding AI (Timothy B. Lee)BlogAI score78

    OpenAI's AI-driven Navier-Stokes result draws anger from mathematicians

    AIOpenAI announced that a swarm of 10,000 agents produced a solution to the Navier-Stokes Millennium Problem, a result that angered mathematicians. NYU mathematician Tristan Buckmaster and Anthropic-employed collaborator Levent Alpöge had been working on related problems and released three draft papers of about 245 pages. Buckmaster said OpenAI's offer to merge efforts required acknowledging an OpenAI model and excluded Alpöge as co-author.

    Why it matters: The piece separates the mathematical result from the collaboration dispute, showing how AI labs' compute spending is straining academic norms around credit and openness.