Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Cloudflare Blog · AIAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  2. Varun MohanAI score40

    Google announces Gemini 4 Argon, a new frontier model for software tasks

    AIGoogle announced Gemini 4 Argon, a new frontier model that delivers frontier performance across complex software tasks, according to Varun Mohan. Thousands of Googlers have been using it internally in Antigravity, and it is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

  3. Google AIAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  4. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  5. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  6. MiniMax (official)AI score38

    Creatify's Boreal-H3 ad video model built on MiniMax H3

    AICreatify Labs released Boreal-H3, a video model built on MiniMax H3 and post-trained specifically for advertising. Reported results include 85.3% reference fidelity, brief success rising from 28% to 50%, and identity match improving from 83% to 94%. Visible defects per clip dropped 70%, while generation time and estimated cost fell 20%.

  7. IdeogramAI score23

    Ideogram 4.5 Performs Strongly Across General Image Editing Tasks

    AIIdeogram 4.5 was built for precise, targeted editing but also performs very well across general editing tasks. Design Arena ranks it 15th in Image Editing with an Elo of 1250, placing it in the same performance band as MAI-Image-2.6 and Gemini 3 Pro Image Preview. It is especially strong at typography edits, such as modifying text in infographics.

  8. Ant LingAI score46

    Ant Ling releases Ling-3.1-flash with 1M-token context, plans open-source

    AIAnt Ling introduced Ling-3.1-flash, a model with about 560B total parameters, about 25B active per token, and up to a 1M-token context window. The company plans to open-source the model soon. It reports 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional across work, coding, and healthcare tasks.

    Image from @AntLingAGI's post
  9. IdeogramAI score46

    Ideogram 4.5 launches with four quality modes at native 2K resolution

    AIIdeogram 4.5 comes in four quality modes, ranging from 0.8¢ to 22¢ per image, all at native 2K resolution. Available now on our launch partners: @superscale_ai @Picsart @cfabricacom @luminal_ai @arena @runware @florafaunaai @krea_ai @trymoda @LeonardoAi @runwayml @pika_labs @fal @ComfyUI @magnific @GammaApp @LumaLabsAI @DesignArena

    Image from @ideogram_ai's post
  10. IdeogramAI score38

    Ideogram 4.5 launches as a precise image edit model

    AIIdeogram released Ideogram 4.5, which it calls the most precise edit model, claiming it avoids the artifacts, pixel shifts, and color changes that leading models add with each edit. The company says this eliminates artifact buildup and makes multi-turn editing possible. It is live in Ideogram, via the API, and with launch partners, with open weights promised soon.

    Video from @ideogram_ai's post
  11. ModelScopeAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post
  12. Kling AI BlogAI score49

    Kling 4.0 Extends Native Video to 30 Seconds With Up to 10 Keyframes

    AIKling 4.0 extends native single-pass video generation from 15 to 30 seconds and adds Multiple Keyframes supporting up to 10 keyframe images, versus Start & End Frames in Kling 3.0. It also expands reference inputs to up to 15 combined assets, including up to 5 videos totaling 30 seconds, and adds 10-bit HDR at 1080p and 4K. The all-new Kling 4.0 will officially launch in October, and Kling 4.0 Flash became available to a limited group of early-access users on September 28.

  13. Artificial Analysis ArticlesAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    AIKorean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  14. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    AIArtificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    Why it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

Sep 29

Sep 29Tue
  1. Liquid AIAI score32

    Liquid AI launches d1, first decision model, beating Jev on HF index

    AILiquid AI announced d1, its first decision model, which it says is the first to outperform Jev on Hugging Face's Decision Index. The company claims d1 wins on multilingual evals, resists prompt injection better, handles longer inputs more effectively, and is built for fast, structured decision-making in software environments. It is available via the Liquid API at console.liquid.ai, with OpenRouter availability coming soon.

    Image from @liquidai's post
  2. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.