Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. GitHub Blog · AI & MLAI score57

    GitHub argues secret protection must scale with AI-driven code growth

    AIGitHub reports that one in three pull requests now involves an AI agent, and that public secret exposures rise with the volume of pushes rather than from declining developer care. It introduces a ModernBERT-based classifier with Microsoft Applied Sciences that evaluates candidate secrets in under two milliseconds and could more than double the secrets prevented at push time. The feature is in private preview, with availability for GitHub Secret Protection customers later this month.

  2. GoogleAI score42

    Google's Project Suncatcher tests TPUs in orbit on a satellite

    AIGoogle launched its first test satellite carrying four TPUs into orbit last week as part of Project Suncatcher, a moonshot exploring whether machine learning infrastructure could one day operate in space. The test aims to determine whether Google's AI hardware can withstand the physical stress of spaceflight and the radiation and thermal extremes of orbit.

    Image from @Google's post
  3. Unsloth AIAI score40

    Unsloth lets users train local decision models on 4GB VRAM

    AIUnsloth released an open-source method to fine-tune LLMs into decision models that run locally, lifting Qwen3.5 0.8B's aggregate accuracy from 20.7% to 74.3% across three decision benchmarks. The team used a Clef head with LoRA (r=64) for one epoch on just 4GB VRAM, with the approach applicable to models such as Qwen3.8 and Gemma 4. A guide and notebooks are available on the Unsloth documentation site and GitHub.

    Image from @UnslothAI's post
  4. SantiagoAI score22

    Model infers derived values from document data, computing yearly costs from monthly figures

    AIA new model extracts values absent from a document by computing them from figures that are present, such as deriving a yearly product cost from a monthly price. Santiago says the video shows examples of inferring complex formulas. The background post describes this as Higher-Order Extraction, which deterministically computes needed numbers from raw page values.

  5. LlamaIndex 🦙AI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

    Video from @llama_index's post
  6. Microsoft ResearchAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    AIMicrosoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    Why it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

  7. 🚨 AI News | TestingCatalogAI score41

    Google releases Foresight macOS app using Gemma 4 for voice notes

    AIGoogle released the Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma 2. The app can connect to Google Drive to build a knowledge graph and, when transcription is active, uses local Gemma 4 E4B or Gemma 4 12B models to transcribe voice notes into new documents. EmbeddingGemma 2 is an open-weight, Apache 2 licensed 740M-parameter multimodal embedding model with an 8K context window.

    Video from @testingcatalog's post
  8. elvisAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    AINVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

    Image from @omarsar0's post
  9. Google GemmaAI score46

    Gemma 4 E4B helps find climate-resilient crop mutations faster

    AIAI lab Living Models pairs Gemma 4 E4B with BOTANIC-1, a genomic language model, to speed up identifying DNA that makes crops climate-resilient. Gemma prepares genomic data and filters candidates, while BOTANIC-1 scores evolutionary impact to pinpoint the causal mutation. In a recent test, the system ranked a target melon yield mutation first out of 2,494 possibilities after an afternoon of computation.

    Video from @googlegemma's post
  10. NVIDIA Technical BlogAI score25

    NVIDIA cuPhoton Speeds Up Scientific Image Analysis for High-Throughput Instruments

    AINVIDIA's cuPhoton targets the computational bottleneck in scientific image pipelines, where data from observatories, telescopes, lasers, and X-ray light sources arrives faster than CPU-bound processing can handle. The source says the bottleneck is usually the whole path from raw sensor data to decision, not one slow kernel. The available text does not give benchmark figures, pricing, or availability details.

  11. IEEE Spectrum · AIAI score32

    HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object Interaction

    AIHiPHI is a 617.5-hour whole-body human motion dataset captured with optical motion capture at sub-millimeter accuracy, including 245.7 hours of human-object interaction with synchronized object trajectories and meshes. The dataset organizes coverage using FrameNet, a linguistic framework for human action. The white paper also reports results from policies trained on HiPHI and deployed on a physical Unitree G1 humanoid robot.

  12. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  13. O'Reilly RadarAI score42

    Build Your Own Post-Training Pipeline: SFT, Reward Model, and PPO

    AIThe final post in O'Reilly Radar's four-part post-training series walks readers through implementing the classic ChatGPT pipeline on Qwen2.5-1.5B, covering SFT, reward model training, and PPO. The walkthrough uses torchtune for SFT and verl, a Ray-based RL framework from ByteDance's team, for reinforcement learning. The author says the goal is hands-on understanding rather than reproducing InstructGPT, which took a large team and thousands of GPU-hours.

  14. Ai2 (Allen Institute for AI)AI score57

    Ai2's Bolmo byte-level language models are published in Nature

    AIAi2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

Oct 6

Oct 6Tue
  1. meng shaoAI score52

    xAI Cookbook adds five apps, expanding Grok API examples to ten

    AIThe xAI Cookbook now has ten runnable Grok API examples across three tracks: real-time voice agents, multimodal generation, and live X data analysis. The author says four voice examples show the same Realtime Voice API across WebSocket, WebRTC, Twilio phone, and mobile transports. The four multimodal examples chain understanding, image generation or editing, video, and TTS, with Grok making creative decisions and Imagine models executing them.

    Image from @shao__meng's post
  2. 👩‍💻 Paige BaileyAI score20

    Google launches ContentPilot to license specialized data for its products

    AIPaige Bailey, a Google and Gemini figure, invited holders of high-quality, specialized data to license or sell it to improve Google products through a new portal, contentpilot.google.com. The post frames data as the most important asset and welcomes such partnerships, but gives no terms, pricing, or eligibility details.

    Image from @DynamicWebPaige's post