Skip to contentSkip to stories

Updated

#Data/Training

Aug 19

Aug 19Wed
  1. Ali GhodsiAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    AIDatabricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

  2. Matei ZahariaAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  3. GeneralistAI score46

    Generalist model learns physical tasks from one or few demonstrations

    AIGeneralist's model reached 59% average success on 10 diverse physical tasks with one-shot prompting straight from pretraining. With few-shot learning, using 10 gradient steps on 5 minutes of data per task, performance rose to 83%. The post calls it the first model it knows of that learns a wide range of dexterous closed-loop physical tasks from one or few demonstrations.

  4. Google · new models on Hugging FaceAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  5. Google · new models on Hugging FaceAI score26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    AIGoogle has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

Aug 18

Aug 18Tue
  1. Liquid AI BlogAI score65

    Liquid AI releases QAD 4-bit LFM2.5 checkpoints for edge deployment

    AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.

    Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.

  2. Stability AIAI score40

    Stability AI launches Stable Audio plugin and enhanced web app for Stable Audio 3.0

    AIStability AI released a Stable Audio plugin that runs Stable Audio 3.0 generation inside DAWs as an instrument, available as a macOS AU and VST3 with Apple Silicon and Intel support. The enhanced StableAudio.com web app adds iterative prompting, audio-to-audio variations, per-track mixing controls, and export, and both tools are in beta and powered by commercially-safe models that users can distribute freely.

Aug 16

Aug 16Sun
  1. Ian JohnsonAI score34

    Ian Johnson maps Prelinger film dataset with UMAP and Marlin-2B vision latents

    AIIan Johnson used UMAP to visualize a video dataset, adding vision latents extracted from Marlin-2B for each clip alongside the included embeddings. He built the interactive map to render smoothly in the browser, with a writeup linked in the post. The quoted post by Daniel van Strien describes indexing 370 hours of Prelinger Archives films into 23,148 timestamped searchable moments.

  2. Replit BlogAI score50

    Replit launches audit logs, Admin API, and workspace settings for enterprises

    AIReplit announced enterprise governance updates including more than 50 audit log events that can stream to SIEM tools like Datadog and Splunk. It also launched a beta Admin API for pulling usage, workspace, member, and project data, with workspace settings for company-wide policies and team-level exceptions. Some features are available now, while workspace settings roll out at the end of the week and the Compliance API at the end of August.

Aug 13

Aug 13Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score38

    MathForm-8B Translates Natural-Language Math Statements into Lean 4 Formal Proofs

    AIMathForm-8B is an open-source autoformalization model from OpenBMB that translates natural-language mathematical statements into Lean 4. It was trained on FormalVerse through supervised fine-tuning, then reinforcement learning using Lean compilation and semantic-consistency feedback. The model is available on Hugging Face under Apache License 2.0 and can be served with Transformers, vLLM, or SGLang, using a recommended max_new_tokens of 16384.

  2. Air Street PressAI score52

    Air Street Press argues logged research decisions could teach AI scientific taste

    AIThe article argues that scientific papers omit the failed experiments and rejected branches that could train AI systems to develop scientific judgment. It describes Alasdair Russell's Cambridge group logging discovery paths as graphs of ideas, and proposes recording six fields per decision, including candidates and outcomes, to test whether this taste transfers to unfamiliar projects.

Aug 12

Aug 12Wed
  1. MiniMax BlogAI score62

    MiniMax releases Music 3.0, an open-weights model for full-length songs

    AIMiniMax introduces Music 3.0, a music generation model that composes, arranges, performs, and produces a complete song from a creative concept and optional lyrics. The post describes an eight-layer RVQ tokenizer, a Hybrid-LM pairing an 8B Global LLM with a 0.6B Local LLM, and a flow-matching and Flow-VAE audio renderer. It says songs can run up to five minutes and that the model focuses on creative intent, arrangement, and vocal naturalness.

    Why it matters: The post explains how the model's pipeline targets structure, acoustic detail, and vocal realism, which helps readers judge where open-weights music generation stands.

Aug 11

Aug 11Tue

Aug 10

Aug 10Mon
  1. Andy JassyAI score38

    Novo Nordisk selects AWS as preferred cloud and strategic AI partner

    AINovo Nordisk has chosen AWS as its preferred cloud provider and strategic AI partner to accelerate drug discovery. The collaboration will combine Novo Nordisk's scientific expertise with AWS AI tools, including Amazon Bio Discovery and Bedrock AgentCore, and establish a co-innovation hub in London. The partnership already spans AWS, Amazon Pharmacy, and One Medical.

Aug 7

Aug 7Fri

Aug 6

Aug 6Thu
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed

Aug 4

Aug 4Tue
  1. Fireworks AI BlogAI score26

    Voyage AI's embedding and reranking models now run natively on Fireworks AI

    AIVoyage AI by MongoDB's full lineup, including the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5, now runs natively on the Fireworks inference platform. The partnership lets teams run embedding, retrieval, reranking, and generation on one platform and one API. Fireworks says Voyage 4 Large outperforms Voyage 4, Voyage 4 Lite, Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large on average retrieval quality.

Aug 3

Aug 3Mon

Jul 30

Jul 30Thu

Jul 29

Jul 29Wed
  1. Fireworks AI BlogAI score54

    Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank

    AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.

  2. Ahmad Al-DahleAI score52

    Ahmad Al-Dahle argues AI capex is both short on compute and overbuilt

    AIAhmad Al-Dahle argues that AI infrastructure faces both a compute shortage and overbuilding, with the four largest hyperscalers planning roughly $725 billion of capex in 2026, up 77 percent from last year. He describes a "mutually assured construction" dynamic in which every well-capitalized player buys the same insurance against falling behind, so the industry overbuilds by construction.

  3. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  4. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  5. Berkeley AI ResearchAI score44

    K-Search Adapts CUDA Kernel Expertise to Apple Silicon MLX Backend

    AIBerkeley AI Research extended the K-Search evolutionary kernel framework with an MLX backend and a CUDA-to-MLX translation layer, letting it adapt existing CUDA kernels for Apple Silicon. The team reports a 0.97x speedup relative to the native MLX Attention kernel and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel. The method uses Gemini 3.5 Pro Preview to both reason about optimizations and write candidate kernels.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

  2. Intern Large ModelsAI score62

    Intern Large Models introduces Visual Pretraining learned from visual documents

    AIIntern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents. The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

Jul 27

Jul 27Mon