Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Alexander DoriaAI score54

    SYNTH paper proposes fully synthetic single-stage training for reasoning models

    AIThe SYNTH paper, titled It's All Training, presents a fully synthetic single-stage pipeline for training workable reasoning models with high data efficiency. The authors argue this approach does not separate training into pretraining, mid-training, or post-training stages. The image shows the paper's abstract, which describes a pipeline built from a 58,000-article Wikipedia-based synthetic corpus and models named Baguettotron-600M and Baguettotron-MoE.

    Image from @Dorialexander's post
  2. DatabricksAI score18

    Databricks launches ai_decide for fast, governed AI decisions

    AIDatabricks has introduced ai_decide, a new AI Function for fast, structured decisions over governed data. It classifies, scores, and chooses next actions in a fraction of a second, with lower latency and cost than an LLM on similar tasks. It is suited to model routing, document processing, agent evaluations, and real-time app logic.

    Video from @databricks's post
  3. Ai2 (Allen Institute for AI)AI score62

    Ai2 releases Olmo-core 3, an open framework for training large MoE models

    AIAi2 released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expert count rose from 8 to 128 with about 3.2B active parameters per token, total capacity grew from 4.6B to 47B, and throughput fell by less than 5%. The framework is fully open, so researchers can train their own MoEs and experiment with routing and parallelism.

    Why it matters: The release documents concrete MoE scaling results and reported failure modes, useful for teams weighing training-stack tradeoffs before adopting an open framework.

  4. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  5. Mastra BlogAI score26

    Mastra Platform Adds VPC-Isolated Postgres Databases for Same-Network Access

    AIMastra platform now lets users attach a VPC-isolated Postgres database to any environment, restricting access to resources on the same network. The database cannot be reached from outside the network, so psql connections from external clients return an error. VPC Postgres joins Turso and Neon as managed database options, with MongoDB and Redis coming soon; it requires mastra@1.32.0 or later.

Sep 30

Sep 30Wed
  1. Apple Machine Learning ResearchAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  2. LlamaIndex 🦙AI score14

    LlamaIndex hosts document-processing events for AI agents in New York and San Francisco

    AILlamaIndex held Tuesday-night events in New York and San Francisco on document processing for AI agents, with the New York room filling a waitlist and San Francisco drawing almost 600 attendees. The talks focused on the problem that agents often receive document text without its layout, so they must guess which figures, such as a monthly rate versus a total on an invoice, mean what.

    Image from @llama_index's post
  3. O'Reilly RadarAI score45

    The Agentic Data Science Playbook: Delegating Analysis to AI Agents

    AIAgentic data science has AI agents explore datasets, choose modeling approaches, run analyses, and explain findings while data scientists frame questions and verify evidence. In an experiment, Claude Opus 5.0 given the vague prompt "Build me a model to detect fraudulent nodes" on a modified Elliptic Bitcoin dataset reported F1 0.87 and ROC AUC 0.99 using a random split that leaked a planted label proxy.

  4. Microsoft ResearchAI score46

    Machine learning system forecasts space-weather grid risk for 66,935 U.S. substations

    AIMicrosoft Research intern-developed machine learning pipeline forecasts location-specific geomagnetic risk for 66,935 substations in the continental United States. It combines solar-wind observations, AE and Dst forecasts, geological conductivity and grid data to estimate risk 30 to 60 minutes ahead. The pipeline detected nearly 80% of major space-weather events during the evaluation period.

  5. Liquid AIAI score42

    LongevityBench: Liquid AI's compact LFMs beat frontier models on aging tasks

    AILiquid AI and InSilicoMeds released LongevityBench, an aging benchmark with 17 tasks spanning clinical records, DNA methylation, transcriptomics, proteomics, and genetics. On several tasks, Liquid AI's compact LFMs outperformed every frontier model the team evaluated. The team plans to present the work to the longevity research community at ARDD this week.

    Video from @liquidai's post
  6. Azure BlogAI score36

    Azure Circular Centers recover value from retired hyperscale hardware

    AIMicrosoft says its Circular Centers now operate eight facilities across North America, Europe, and Asia Pacific to decide the next life of decommissioned Azure hardware. Last year, Microsoft achieved a 92% reuse and recycling rate for decommissioned servers and components. Since 2014, Azure cores per rack have increased about 13-fold while power for the same task fell roughly 90%.

  7. OpenBMBAI score42

    Diffusion Reward Models learn full human preference distributions, not single scores

    AIOpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score. The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs. DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.

    Image from @OpenBMB's post
  8. X.PINAI score72

    DeepSeek releases Ascend versions of its core kernel toolkit

    AIDeepSeek has released an Ascend toolkit that mirrors its Nvidia components, including TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect. It says every TileLang kernel used in its training now has a high-performance Ascend implementation. The post also reports that a 128-card Ascend 950 supernode, jointly optimized with Huawei, has key compute and communication tests approaching hardware limits.

    Image from @thexpin's post

Sep 29

Sep 29Tue
  1. Jerry LiuAI score20

    Jev, a System One model, tops OSS rivals on document tasks

    AIJerry Liu says Jev, a System One model, outperformed other open-source classifiers and document-specific models on orientation detection, language detection, classification, and splitting. The benchmark measured accuracy, cost, and latency across these fast document decisions, with Jev leading most comparisons. The benchmark code is available in the run-llama/jev_vs_oss repository.

    Video from @jerryjliu0's post
  2. Jerry LiuAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post
  3. Fireworks AI BlogAI score51

    Fireworks explains how numerical mismatch and MoE routing can derail RL training

    AINumerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

  4. Microsoft Foundry BlogAI score30

    Why content extraction still matters in the GenAI era

    AIMicrosoft's Azure AI team argues that better models do not eliminate the need for a dedicated content extraction layer, since agents need trustworthy, structured, and auditable inputs. The post notes that building extraction directly on an LLM quickly demands chunking, layout parsing, grounding, normalization, and evaluation infrastructure. Microsoft positions Azure Document Intelligence and Azure Content Understanding in Foundry Tools as managed options for that layer.

  5. Azure BlogAI score40

    SQL Server on Azure Local Becomes Generally Available for Connected and Disconnected Use

    AIMicrosoft has made SQL Server on Azure Local generally available for connected and disconnected deployments, letting organizations run SQL Server in their own datacenters and edge locations. Disconnected operations continue locally where external connectivity is restricted or unavailable. Eligible existing SQL Server licenses can be used, and Foundry Local on Azure Local, currently in preview, brings AI inference alongside SQL Server data.

  6. Microsoft ResearchAI score34

    Microsoft Research unveils Quine, an early multimodal world model of biology

    AIMicrosoft Research has introduced Quine, an early-stage research effort to build a multimodal world model of biology that connects insights across biological scales and modalities. The system is designed to help scientists computationally search a space far larger than intuition allows and prioritize hypotheses before lab testing. Experimental results are meant to feed back into the model and sharpen future research directions.

    Video from @MSFTResearch's post
  7. Microsoft ResearchAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

  8. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

    Image from @OpenBMB's post
  9. Ahead of AI (Sebastian Raschka)AI score43

    Language Models for Text Classification: From Bag-of-Words to Jev

    AISebastian Raschka traces text classification from bag-of-words models such as naive Bayes and logistic regression through pre-transformer neural networks, then sets up an analysis of the recently released Jev AI model. The article frames Jev as a general-purpose classifier that trades specialized accuracy for speed, cost, and breadth of tasks.

  10. Azure BlogAI score46

    Microsoft Fabric and Copilot Integration: New Data Foundation Features for Agents

    AIMicrosoft is bringing business context from Fabric IQ into Microsoft Copilot, with Fabric IQ in Copilot Chat and Cowork generally available and integration into the new Code experience coming soon through the Frontier program. Power BI is also gaining agentic app creation in Power BI Desktop, letting users generate applications from trusted semantic models and publish them to Microsoft Fabric.