Skip to content

#Data/Training

Oct 8

TodayOct 8Thu84 items
  1. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

  2. Claude BlogAI score67

    Claude adds live dashboards and animated explainers, Docs and Slides leave beta

    Claude now turns company data into dashboards that stay current, and it can build animated explainers from a prompt. Dashboards connect to BigQuery, Databricks, Snowflake, and Salesforce in beta on paid plans, while Motion is in beta on Team and Enterprise. Docs, Slides, and Design are out of beta and available on every plan, including Free.

    AIWhy it matters: The post specifies which data platforms connect, which features move out of beta, and where admins control access, clarifying what changes for enterprise workflows.

  3. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    Johns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    AIWhy it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

Oct 7

Oct 7Wed
  1. The Next PlatformAI score46

    Memory Now Drives the IT Industry as DRAM and Flash Prices Surge

    Memory has overtaken compute as the central control point in IT, according to The Next Platform, as generative and agentic AI drive demand for DRAM, HBM, and flash. Server DDR5 memory now sells for roughly 9X to 13X its November 2022 street price, while a 30 TB enterprise SSD costs 6X to 7X more. HBM pricing has risen only about 1.6X since the GenAI boom began, the article says.

  2. François CholletAI score44

    Chollet: Programming and math training don't boost general intelligence

    François Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  3. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  4. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    AIWhy it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. Hugging Face BlogAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    A Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    AIWhy it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  6. DatabricksAI score34

    Databricks adds Workday Data Connect federation to Unity Catalog in Beta

    Databricks has put Workday Data Connect federation into Beta in Unity Catalog, letting teams query Workday HR and finance data without copying it. Workday Data Cloud customers get zero-copy, read-only access to the shared tables, with Databricks running queries and Unity Catalog governing access, lineage, and auditing. Teams can combine current people and financial data with other enterprise data for analytics and AI, including Genie-powered natural-language exploration.

  7. Nathan LambertAI score13

    If you want people to do open research on the impact of distillation to frontier progress, we happily would at @trillium_labs, but just need to scale up funding/compute before getting that done. Happy to make this happen sooner - you know where to find me!

    If you want people to do open research on the impact of distillation to frontier progress, we happily would at @trillium_labs, but just need to scale up funding/compute before getting that done. Happy to make this happen sooner - you know where to find me!

  8. Sara HookerAI score26

    Generated videos are getting so good. One of the @adaption_ai team put this together as an explainer of invent a dataset. 🔥 The biggest hurdle in frontier AI is curating high-quality, diverse data. Invent solves for this zero data regime.

    Generated videos are getting so good. One of the @adaption_ai team put this together as an explainer of invent a dataset. 🔥 The biggest hurdle in frontier AI is curating high-quality, diverse data. Invent solves for this zero data regime.

  9. Google ResearchAI score23

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

  10. Google ResearchAI score42

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

  11. NVIDIA AIAI score26

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

  12. Google ResearchAI score10

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

  13. Microsoft Foundry BlogAI score22

    Azure Document Intelligence vs. Content Understanding: Choosing the Right Document Service

    Microsoft's Foundry blog guide advises keeping existing Azure Document Intelligence workloads that meet production requirements. It recommends evaluating Azure Content Understanding for high-variation, unstructured, reasoning, RAG, or multimodal document scenarios, and for new cloud OCR or layout workloads.

  14. DatabricksAI score22

    What if AI could figure out the right business context without waiting for a fully built ontology? Genie Ontology does that by inferring relevant context across your data and assets to answer open-ended questions. Advancing Analytics’ @MrSiWhiteley digs into how OntoRank chooses what to trust, why certified assets still matter, and what that means for data governance. Watch the full video: https://www.youtube.com/watch?v=voZXeHP2qWQ

    What if AI could figure out the right business context without waiting for a fully built ontology? Genie Ontology does that by inferring relevant context across your data and assets to answer open-ended questions. Advancing Analytics’ @MrSiWhiteley digs into how OntoRank chooses what to trust, why certified assets still matter, and what that means for data governance. Watch the full video: https://www.youtube.com/watch?v=voZXeHP2qWQ

  15. Epoch AIAI score22

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

  16. Epoch AIAI score50

    AI developers aim to create an automated AI researcher. How close are they? To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming.

    AI developers aim to create an automated AI researcher. How close are they? To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming.

  17. Lucas BeyerAI score38

    Another cool update from the robotics world! While we're all wow-ed in mental land, there's a ton of progress and i feel like acceleration happening in physical land too now. My guess is in good part also (but not only) thanks to the coding models progress over the last year.

    Another cool update from the robotics world! While we're all wow-ed in mental land, there's a ton of progress and i feel like acceleration happening in physical land too now. My guess is in good part also (but not only) thanks to the coding models progress over the last year.

  18. GitHub Blog · AI & MLAI score57

    GitHub argues secret protection must scale with AI-driven code growth

    GitHub reports that one in three pull requests now involves an AI agent, and that public secret exposures rise with the volume of pushes rather than from declining developer care. It introduces a ModernBERT-based classifier with Microsoft Applied Sciences that evaluates candidate secrets in under two milliseconds and could more than double the secrets prevented at push time. The feature is in private preview, with availability for GitHub Secret Protection customers later this month.

  19. Unsloth AIAI score23

    With just 2.5GB VRAM, you can train your own Decision Model using small models like Laya. Training is simply done through a UI interface. Video tutorial and analysis are in our guide. GitHub repo: https://github.com/unslothai/unsloth

    With just 2.5GB VRAM, you can train your own Decision Model using small models like Laya. Training is simply done through a UI interface. Video tutorial and analysis are in our guide. GitHub repo: https://github.com/unslothai/unsloth

  20. GoogleAI score42

    Google's Project Suncatcher tests TPUs in orbit on a satellite

    Google launched its first test satellite carrying four TPUs into orbit last week as part of Project Suncatcher, a moonshot exploring whether machine learning infrastructure could one day operate in space. The test aims to determine whether Google's AI hardware can withstand the physical stress of spaceflight and the radiation and thermal extremes of orbit.

  21. Daniel HanAI score48

    We made it possible to train your own Decision model locally on just 3GB VRAM! We converted Qwen, Gemma, Llama all into decision models by fine-tuning using a Clef head, boosting accuracy from 30 to up to 78%. You can try it yourself via Unsloth Desktop!

    We made it possible to train your own Decision model locally on just 3GB VRAM! We converted Qwen, Gemma, Llama all into decision models by fine-tuning using a Clef head, boosting accuracy from 30 to up to 78%. You can try it yourself via Unsloth Desktop!

  22. Unsloth AIAI score40

    Unsloth lets users train local decision models on 4GB VRAM

    Unsloth released an open-source method to fine-tune LLMs into decision models that run locally, lifting Qwen3.5 0.8B's aggregate accuracy from 20.7% to 74.3% across three decision benchmarks. The team used a Clef head with LoRA (r=64) for one epoch on just 4GB VRAM, with the approach applicable to models such as Qwen3.8 and Gemma 4. A guide and notebooks are available on the Unsloth documentation site and GitHub.

  23. SantiagoAI score22

    This is very clever: It's a model that "extracts" values that don't exist in a document but can be derived from the values that do exist. For example, you can ask for the yearly cost of using a product, and it will automatically compute it from the monthly cost. The video shows a bunch of cool examples. It looks like it can infer really convoluted and sophisticated formulas.

    This is very clever: It's a model that "extracts" values that don't exist in a document but can be derived from the values that do exist. For example, you can ask for the yearly cost of using a product, and it will automatically compute it from the monthly cost. The video shows a bunch of cool examples. It looks like it can infer really convoluted and sophisticated formulas.

  24. PerplexityAI score41

    Both sizes are distilled token by token from one 18B teacher, so they share one embedding space. A corpus indexed with the 9B model can be searched with 0.6B queries. That lifts ViDoRe v3 from 62.3% to 63.5% with no added query cost.

    Both sizes are distilled token by token from one 18B teacher, so they share one embedding space. A corpus indexed with the 9B model can be searched with 0.6B queries. That lifts ViDoRe v3 from 62.3% to 63.5% with no added query cost.

  25. PerplexityAI score36

    Dense embedding models compress each document into one vector, which loses detail as pages get longer or more visual. pplx-embed-v2-late keeps a 128-dimensional vector per token and scores with MaxSim, so each query token is matched to its closest token in the document.

    Dense embedding models compress each document into one vector, which loses detail as pages get longer or more visual. pplx-embed-v2-late keeps a 128-dimensional vector per token and scores with MaxSim, so each query token is matched to its closest token in the document.

  26. LlamaIndexAI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    LlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

  27. Microsoft ResearchAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    Microsoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    AIWhy it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

  28. Ai2AI score36

    We're making new Stage 1 checkpoints available with the byte-level components already trained & the original model's weights unchanged. Researchers can build on these to test new architectures & train the full system without repeating that initial stage.

    We're making new Stage 1 checkpoints available with the byte-level components already trained & the original model's weights unchanged. Researchers can build on these to test new architectures & train the full system without repeating that initial stage.