Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 24

Sep 24Thu
  1. ModelScopeAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

    Video from @ModelScope2022's post
  2. Google · Gemini appAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  3. Google Cloud · AI & Machine LearningAI score55

    Gemini 3.8 Live with Live Avatar becomes generally available in Gemini Enterprise

    AIGoogle says Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, with US and EU endpoints, provisioned throughput, and enterprise compliance. Its video avatars use synchronized lip-syncing, custom avatars are limited to an allowlist, and generated audio and video carry SynthID watermarks. The model also understands and speaks 97 languages and can run tool calls in the background while the conversation continues.

  4. ModelScopeAI score38

    Qwen-Image-2.1-Fun-Controlnet-Union adds eight controls and inpainting

    AIModelScope released Qwen-Image-2.1-Fun-Controlnet-Union, a single checkpoint adding eight structural controls, including Canny, Depth, Pose, and Scribble, plus inpainting to Qwen-Image 2.1. Control and inpainting share one branch with 16 injection points across every second Transformer block, keeping the base model frozen and requiring no checkpoint switching. It runs at guidance scale 1.0 with CFG-distilled sampling and prefix KV caching, and is available under the Qwen Research License with base Qwen-Image 2.1 weights required.

    Image from @ModelScope2022's post
  5. KrASIA · Big TechAI score55

    Mind Lab launches Mint Recursive, a post-training platform for companies

    AIMind Lab unveiled Mint Recursive, a post-training and inference platform for industry use, alongside Macaron-V1.1, a model post-trained entirely on it. Macaron-V1.1 is a 752-billion-parameter model built from GLM-5.3 with four two-billion-parameter LoRA expert modules for chat, agents, coding, and generation. The platform is serverless and bills by token usage, and it collects feedback from models in use to support continued training.

  6. inclusionAI (Ant Ling) · new models on Hugging FaceAI score22

    inclusionAI Publishes Training-Content Summaries for Ling and Ring Models

    AIinclusionAI has published public training-content summaries on Hugging Face for its Ling and Ring model versions, including Ling-2.0, Ling-2.5, Ling-2.6-1T, Ling-3.0, Ring-2.0, Ring-2.5-1T, and Ring-2.6-1T. The documents, organized under the template associated with Article 53(1)(d) of Regulation (EU) 2024/1689, contain documentation only, not model weights or training datasets. Each summary covers only the model versions it names.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  2. Black Forest LabsAI score40

    FLUX 3 Action uses a smaller architecture to predict actions and frames

    AIBlack Forest Labs says FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3 but uses a smaller architecture. The company attributes this smaller size to more efficient representations learned through its Self-Flow research. During midtraining, the model was trained to predict actions and future frames together.

    Video from @bfl_ai's post
  3. Black Forest LabsAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

    Image from @bfl_ai's post
  4. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  5. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  6. Google for DevelopersAI score52

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle announced two new Gemini 3.8 text-to-speech models, positioned as its most expressive yet. Gemini 3.8 Flash TTS targets creative work, letting developers use natural language to define vocal personas, cues, pacing, and dialects, while Gemini 3.8 Flash-Lite TTS is built for high-volume pipelines such as bulk audiobook production and audio dubbing. Both are available now through the Gemini API in Google AI Studio.

  7. Google AIAI score62

    Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle AI launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it describes as its most expressive audio models yet. The models support custom voices across 100+ languages, 2,000+ prebuilt voices, multi-speaker conversations, line-by-line delivery control, and cues such as <laughs> and |mhm|. Google positions Flash TTS for bespoke voices in gaming, audiobooks, and podcasts, and Flash-Lite TTS for near real-time voice agents, high-volume dubbing, and bulk audio.

    Video from @GoogleAI's post
  8. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  9. Google DeepMind · YouTubeAI score46

    Gemini 3.8 text-to-speech lets developers design and clone custom voices

    AIGoogle DeepMind's latest Gemini Audio models let developers design new vocal personas from natural language prompts, directing pacing, back channeling, and dialect shifts line by line. Developers can also recreate consistent adult voice profiles from a 30-second audio sample, with built-in consent verification, SynthID watermarking, and C2PA credentials.

  10. Baseten BlogAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  11. ModelScopeAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post
  12. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  13. howie.seriousAI score22

    Opus 5.5 praised for language, visual taste, code quality, and token efficiency

    AIThe X user howie.serious says Claude Opus 5.5 delivers high language quality, good visual taste, strong code quality, and notably low token usage. He also compares Anthropic's reported roughly $2 trillion IPO valuation with OpenAI's roughly $1.2 trillion fundraising valuation, arguing OpenAI is worth about 0.6 Anthropics and the gap may widen.

  14. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Image from @ModelScope2022's post
  15. KrASIA · Big TechAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  16. ModelScopeAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Image from @ModelScope2022's post