Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Aug 24

Aug 24Mon

Aug 23

Aug 23Sun

Aug 21

Aug 21Fri
  1. Jim FanAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

    Video from @DrJimFan's post
  2. Amazon ScienceAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    AIAmazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

Aug 20

Aug 20Thu
  1. Mistral AIAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Ali GhodsiAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    AIDatabricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

    Image from @alighodsi's post
  2. Matei ZahariaAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  3. GeneralistAI score46

    Generalist model learns physical tasks from one or few demonstrations

    AIGeneralist's model reached 59% average success on 10 diverse physical tasks with one-shot prompting straight from pretraining. With few-shot learning, using 10 gradient steps on 5 minutes of data per task, performance rose to 83%. The post calls it the first model it knows of that learns a wide range of dexterous closed-loop physical tasks from one or few demonstrations.

    Video from @GeneralistAI's post
  4. Google · new models on Hugging FaceAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  5. Google · new models on Hugging FaceAI score26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    AIGoogle has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

Aug 18

Aug 18Tue
  1. Liquid AI BlogAI score65

    Liquid AI releases QAD 4-bit LFM2.5 checkpoints for edge deployment

    AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.

    Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.

  2. Stability AIAI score40

    Stability AI launches Stable Audio plugin and enhanced web app for Stable Audio 3.0

    AIStability AI released a Stable Audio plugin that runs Stable Audio 3.0 generation inside DAWs as an instrument, available as a macOS AU and VST3 with Apple Silicon and Intel support. The enhanced StableAudio.com web app adds iterative prompting, audio-to-audio variations, per-track mixing controls, and export, and both tools are in beta and powered by commercially-safe models that users can distribute freely.

Aug 16

Aug 16Sun
  1. Ian Johnson 🔬🤖AI score34

    Ian Johnson maps Prelinger film dataset with UMAP and Marlin-2B vision latents

    AIIan Johnson used UMAP to visualize a video dataset, adding vision latents extracted from Marlin-2B for each clip alongside the included embeddings. He built the interactive map to render smoothly in the browser, with a writeup linked in the post. The quoted post by Daniel van Strien describes indexing 370 hours of Prelinger Archives films into 23,148 timestamped searchable moments.

    Video from @enjalot's post
  2. Replit BlogAI score50

    Replit launches audit logs, Admin API, and workspace settings for enterprises

    AIReplit announced enterprise governance updates including more than 50 audit log events that can stream to SIEM tools like Datadog and Splunk. It also launched a beta Admin API for pulling usage, workspace, member, and project data, with workspace settings for company-wide policies and team-level exceptions. Some features are available now, while workspace settings roll out at the end of the week and the Compliance API at the end of August.

Aug 13

Aug 13Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score38

    MathForm-8B Translates Natural-Language Math Statements into Lean 4 Formal Proofs

    AIMathForm-8B is an open-source autoformalization model from OpenBMB that translates natural-language mathematical statements into Lean 4. It was trained on FormalVerse through supervised fine-tuning, then reinforcement learning using Lean compilation and semantic-consistency feedback. The model is available on Hugging Face under Apache License 2.0 and can be served with Transformers, vLLM, or SGLang, using a recommended max_new_tokens of 16384.

  2. Air Street PressAI score52

    Air Street Press argues logged research decisions could teach AI scientific taste

    AIThe article argues that scientific papers omit the failed experiments and rejected branches that could train AI systems to develop scientific judgment. It describes Alasdair Russell's Cambridge group logging discovery paths as graphs of ideas, and proposes recording six fields per decision, including candidates and outcomes, to test whether this taste transfers to unfamiliar projects.

Aug 12

Aug 12Wed
  1. MiniMax BlogAI score62

    MiniMax releases Music 3.0, an open-weights model for full-length songs

    AIMiniMax introduces Music 3.0, a music generation model that composes, arranges, performs, and produces a complete song from a creative concept and optional lyrics. The post describes an eight-layer RVQ tokenizer, a Hybrid-LM pairing an 8B Global LLM with a 0.6B Local LLM, and a flow-matching and Flow-VAE audio renderer. It says songs can run up to five minutes and that the model focuses on creative intent, arrangement, and vocal naturalness.

    Why it matters: The post explains how the model's pipeline targets structure, acoustic detail, and vocal realism, which helps readers judge where open-weights music generation stands.

Aug 11

Aug 11Tue

Aug 10

Aug 10Mon
  1. Andy JassyAI score38

    Novo Nordisk selects AWS as preferred cloud and strategic AI partner

    AINovo Nordisk has chosen AWS as its preferred cloud provider and strategic AI partner to accelerate drug discovery. The collaboration will combine Novo Nordisk's scientific expertise with AWS AI tools, including Amazon Bio Discovery and Bedrock AgentCore, and establish a co-innovation hub in London. The partnership already spans AWS, Amazon Pharmacy, and One Medical.

    Image from @ajassy's post

Aug 7

Aug 7Fri
  1. Matei ZahariaAI score44

    Matei Zaharia says AI Gateways let teams cut token costs centrally

    AIMatei Zaharia argues AI tokens are now a resource to optimize in software engineering, with companies routing all AI usage through an AI Gateway. The approach enables centralized analysis, which found settings on Claude Code and Codex that can substantially lower cost, plus smart routing and per-task budgets for engineers.

Aug 6

Aug 6Thu
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed

Aug 4

Aug 4Tue
  1. Fireworks AI BlogAI score26

    Voyage AI's embedding and reranking models now run natively on Fireworks AI

    AIVoyage AI by MongoDB's full lineup, including the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5, now runs natively on the Fireworks inference platform. The partnership lets teams run embedding, retrieval, reranking, and generation on one platform and one API. Fireworks says Voyage 4 Large outperforms Voyage 4, Voyage 4 Lite, Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large on average retrieval quality.

Aug 3

Aug 3Mon