Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 6

Oct 6Tue
  1. ARC PrizeOfficialAI score28

    Grok 4.7 scores 1.8% on ARC-AGI-3 standard harness

    AIGrok 4.7 scored 1.8% on ARC-AGI-3 in the standard harness, which lets models carry notes between turns, slightly below the 2.1% reported for Grok 4.7 in that setting. In a new provider adapter harness that preserves opaque reasoning and enables auto compaction, the score rose to 10.0%.

  2. ARC PrizeOfficialAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

    Image from @arcprize's post
  3. ARC PrizeOfficialAI score38

    Grok 4.7 scores 1.8% on ARC-AGI-3, trails Grok 4.6 on ARC-AGI-2

    AISpaceXAI's Grok 4.7 scored 90.2% on ARC-AGI-1 at $0.64 per task, higher than Grok 4.6, according to ARC Prize. It reached 61.4% on ARC-AGI-2 at $2.01 per task and 1.8% on ARC-AGI-3 under the standard harness ($2.7k), or 10.0% with the provider adapter harness ($4.8k), both lower than Grok 4.6 on those two benchmarks.

    Image from @arcprize's post
  4. 👩‍💻 Paige BaileyXAI score54

    EmbeddingGemma 2 launches as an Apache 2.0 multimodal embeddings model

    AIGoogle's EmbeddingGemma 2 is an open embeddings model for on-device use that covers code, image, video, audio, and text. It comes in modular sizes from 270M text/code to 740M full multimodal, supports Matryoshka truncation down to 128 dimensions, and reports a 14% gain on MTEB Code over v1 under an Apache 2.0 license. The author's post highlights the release and a Hugging Face demo, while the benchmark table compares it with several models.

    Video from @DynamicWebPaige's post
  5. KreaOfficialAI score36

    Nano Banana 2.1 launches on Krea with 4K generation support

    AIKrea has made Nano Banana 2.1 available on its platform, promising better prompt understanding, sharper edits, and stronger subject consistency. The model now supports generations up to 4K resolution.

    Video from @krea_ai's post
  6. KhazixXAI score32

    Khazix builds an enterprise platform replacing Feishu's workspace in two days

    AIThe author spent two days building an internal enterprise platform on all Feishu data and a self-built MCP, replacing Feishu's native workbench to handle Vibe Coding app deployment, security, permissions, and app and skill circulation. A custom configuration interface is planned so employees can use their own Agents to modify their homepages and data pages.

    Image from @Khazix0918's post
  7. SantiagoXAI score34

    Ampersand packages Salesforce integration work for enterprise AI agents

    AIAmpersand lets developers connect an AI agent to a customer's Salesforce by configuring object and field mappings. The platform then handles API calls, authentication, token refreshes, and retries, which the post presents as a major advantage given the difficulty of managing multiple Salesforce accounts.

  8. NewcomerBlogAI score23

    Fintech Founders and Investors Debate AI Agents at Machine Earning Summit

    AIAt the Machine Earning AI Summit in San Francisco, Town co-founder Jean-Denis Greze said personal AI agent purchases will initially require human approval, with mistakes budgeted in like credit card fraud. Lead Bank CEO Jackie Reses raised liability questions, asking "If a model hallucinates, whose responsibility is that?" Chime co-founder Ryan King said AI agents will make switching banks easier, though regular people are not yet ready to let AI manage their money.

  9. Joshua AchiamXAI score26

    Joshua Achiam argues success lies in human inner lives, not cosmic control

    AIJoshua Achiam argues that many in Silicon Valley wrongly define success as controlling the largest share of matter and energy in the universe, a goal beyond human limits that can drive them toward successionism. He contends that success instead comes from inner lives, relationships, creativity, cooperation, and striving to overcome human limitations, which could make them less pessimistic.

  10. Dongxi NLPXAI score34

    Mistral AI releases Mistral Large 4, dubbed "Le Chonk"

    AIMistral AI has released Mistral Large 4, a model nicknamed "Le Chonk," according to a post by Dongxi NLP. The post also highlights "sovereign AI" as a keyword, tying the release to the theme of national or independent AI capability. No specifications, benchmarks, or pricing are given in the post itself.

  11. Google FlowOfficialAI score20

    Google Flow partners with Divine on "Tequila Dance" music video

    AIGoogle Flow says it partnered with Divine to create the music video for "Tequila Dance," framing the project as a showcase of how technology can amplify human expression. The post provides no further details about the tools or production process used.

    Video from @FlowbyGoogle's post
  12. 👩‍💻 Paige BaileyXAI score60

    Google releases Nano Banana 2.1 image model at $0.034 per image

    AIGoogle's Nano Banana 2.1, model gemini-nano-banana-2.1, is now available and is said to outperform the previous Pro model at about a quarter of the price, $0.034 per image versus $0.134. The quoted post lists improved instruction following, better in-image text rendering, grounding with Google Image Search, and up to 5 characters of consistency plus 14 reference images. It is available in Google AI Studio, the Gemini API, Google Cloud, the Gemini app, and Flow. The author's own post is a playful reaction praising its design ability and shows a generated vegan basketball food truck poster.

    Video from @DynamicWebPaige's post
  13. Microsoft ResearchOfficialAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  14. SGLangOfficialAI score46

    SGLang adds Day-0 support for Google's EmbeddingGemma 2

    AISGLang now supports EmbeddingGemma 2 from Google DeepMind on day zero. The multimodal embedding model maps text, code, images, video, and audio into one shared 768d space, with 8K context and 100+ languages. Its modular encoders range from a 270M text-only footprint to 740M for all modalities, with Matryoshka embeddings.

    Image from @sgl_project's post
  15. GoogleOfficialAI score43

    EmbeddingGemma 2 Delivers Best-in-Class Performance at 740M Parameters

    AIGoogle's EmbeddingGemma 2 is a 740M-parameter embedding model that outperforms some models more than twice its size while using about 191MB to 567MB of active RAM. It offers an 8K context window, 4x larger than the first generation, and can process up to 5.5 minutes of audio, 29 images, or 58 video frames in one pass.

    Image from @Google's post
  16. GoogleOfficialAI score44

    EmbeddingGemma 2 pairs with Gemma 4 for on-device RAG

    AIGoogle says EmbeddingGemma 2, paired with Gemma 4, enables efficient on-device retrieval-augmented generation with a lower memory footprint. In this setup, EmbeddingGemma 2 retrieves local files and Gemma 4 reasons over them to produce grounded answers while keeping data private.

    Video from @Google's post
  17. TiboXAI score23

    Codex adds "Approve for me" auto-review permission mode

    AICodex now offers an "Approve for me" permission mode, which automatically reviews actions instead of requiring manual approval. To enable it, open the permissions menu below the composer and select "Approve for me."

  18. ElevenLabsOfficialAI score20

    ElevenLabs' ElevenAgents Architect analyzes transcripts and implements agent improvements

    AIElevenAgents Architect is a tool that answers questions about your agents, such as why customers asked for a human agent on refund calls and how to improve resolution rates on account queries. It analyzes your transcripts, suggests improvements, implements them, and can build a test set to keep your agent on brand.

    Image from @ElevenLabs's post
  19. ElevenLabsOfficialAI score40

    ElevenLabs launches ElevenAgents Architect to help teams build AI agents

    AIElevenLabs introduced ElevenAgents Architect, an expert built into ElevenAgents that helps teams launch and improve AI agents through voice or text. The post describes it as a conversational way to create and refine agents without further technical detail provided.

    Video from @ElevenLabs's post
  20. 👩‍💻 Paige BaileyXAI score38

    Google's Nano Banana 2.1 image model now available in Google AI Studio

    AIGoogle has released Nano Banana 2.1, its latest image generation and editing model, which it says outperforms previous versions with gains in visual design, mask-based editing, and subject consistency. Paige Bailey shared a pug smile example made with the model and invited users to try it in Google AI Studio.

    Image from @DynamicWebPaige's post
  21. Philipp SchmidXAI score62

    Google releases Nano Banana 2.1 image model at $0.034 per image

    AIGoogle's Nano Banana 2.1 (gemini-nano-banana-2.1) is now available and outperforms the previous Pro model at $0.034 per image, versus $0.134 before. It adds improved instruction following, better in-image text rendering, grounding with Google Image Search, and consistency for up to 5 characters with 14 reference images. It is available in Google AI Studio, the Gemini API, Google Cloud, the Gemini app, and Flow by Google.

    Image from @_philschmid's post
  22. falOfficialAI score38

    Nano Banana 2.1 image model now available on fal

    AIGoogle's Gemini Nano Banana 2.1 is now available on fal, offering significantly faster generation than Nano Banana 2. The model adds major gains in visual design, mask-based editing, and subject consistency.

    Video from @fal's post
  23. Microsoft CopilotOfficialAI score22

    Microsoft unveils new Copilot identity connecting apps, tools, and context

    AIMicrosoft's Copilot has a new identity designed around how work actually happens, bringing apps, tools, and context into one connected experience. The post offers a behind-the-scenes look at the redesign but gives no further specifics.

    Video from @MSFTCopilot's post
  24. Google GemmaOfficialAI score62

    Google Gemma introduces EmbeddingGemma 2, a multimodal on-device embedding model

    AIGoogle Gemma announces EmbeddingGemma 2, a lightweight embedding model that maps text, code, images, video, and audio into a single unified embedding space. The model has a 740M parameter form factor with modular encoders, Matryoshka Representation Learning dimensions from 768 down to 128, and an 8K context window that is 4x larger than the text-only EmbeddingGemma. It is released under the commercially permissive Apache 2.0 license.

    Video from @googlegemma's post
  25. Google for DevelopersOfficialAI score48

    EmbeddingGemma 2 arrives as compact multimodal on-device embedding model

    AIGoogle has released EmbeddingGemma 2, a compact 740M-parameter model built for on-device and edge applications. It natively embeds text, images, audio, and video into one shared space, enabling cross-format search without manual organization, translation, or labeling.

    Video from @googledevs's post
  26. Google for DevelopersOfficialAI score40

    Google's multimodal embedding toolkit runs fully offline on device

    AIGoogle's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.

    Video from @googledevs's post
  27. Google DeepMindOfficialAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    AIGoogle DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

    Image from @GoogleDeepMind's post
  28. Sundar PichaiXAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

    Video from @sundarpichai's post