Skip to content

#Data/Training

Aug 29

Aug 29Sat
  1. Chips and CheeseAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    Samsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Daniel HanAI score50

    We quantized GLM-5.3 to dynamic 1-bit (217GB) vs BF16 (1.5TB) and it retains ~76% top-1% accuracy whilst being 83% We did a small basic snake game in Unsloth Desktop using the UD 1-bit and it worked well! zai is on a roll with GLM-5.3-Flash and now GLM-5.3!

    We quantized GLM-5.3 to dynamic 1-bit (217GB) vs BF16 (1.5TB) and it retains ~76% top-1% accuracy whilst being 83% We did a small basic snake game in Unsloth Desktop using the UD 1-bit and it worked well! zai is on a roll with GLM-5.3-Flash and now GLM-5.3!

  2. Meituan LongCatAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    Meituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

Aug 27

Aug 27Thu
  1. Thinking MachinesAI score40

    Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task. Guest post by researchers at UIUC and Bridgewater, in collaboration with our team. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

    Cleaning data and aligning the reward function for RLVR takes expertise and effort upfront, but the result is a model that's state-of-the-art on a complex task. Guest post by researchers at UIUC and Bridgewater, in collaboration with our team. https://thinkingmachines.ai/news/putting-task-expertise-into-rl

  2. Ali GhodsiAI score22

    Branch your Neon database to protect against agent wipes

    To avoid this scenario where agents wipe everything out permanently, just branch your database, it's super easy to do on Neon Lakebase: ๐š—๐šŽ๐š˜๐š—๐šŒ๐š๐š• ๐š‹๐š›๐šŠ๐š—๐šŒ๐š‘๐šŽ๐šœ ๐šŒ๐š›๐šŽ๐šŠ๐š๐šŽ --๐š—๐šŠ๐š–๐šŽ ๐š—๐šŽ๐š ๐š‹๐š›๐šŠ๐š—๐šŒ๐š‘

  3. OpenBMB (MiniCPM) ยท new models on Hugging FaceAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    OpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

Aug 26

Aug 26Wed
  1. Tencent ยท new models on Hugging FaceAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    Tencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Jazzyear ยท Articles (็”ฒๅญๅ…‰ๅนด)AI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    In an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.

  3. Google Developers BlogAI score42

    Google Developers Blog explains deep learning with Keras for astroparticle physics data analysis

    The Google Developers Blog post describes how deep learning can analyze the large, image-like sensor data from astroparticle observatories such as the Pierre Auger Observatory and IceCube. The author argues these methods could improve instrument sensitivity and reveal patterns in cosmic-ray and neutrino signals that traditional analysis techniques miss.

  4. Amazon ScienceAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    Amazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

  5. Ai2 ยท new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    Ai2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  6. Ai2 ยท new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    Ai2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  7. Ai2 ยท new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    Ai2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  8. Ai2 ยท new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    Ai2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Google Developers BlogAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    Google Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.

  2. Daniel HanAI score34

    Yes, you can fine-tune Qwen3.8-27B completely for free by just having a Google account! ๐Ÿฆฅ Kaggle, like Colab, provides 30 hours of free GPU with 2ร— Tesla T4s. Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.

    Yes, you can fine-tune Qwen3.8-27B completely for free by just having a Google account! ๐Ÿฆฅ Kaggle, like Colab, provides 30 hours of free GPU with 2ร— Tesla T4s. Using QLoRA and our kernels, Unsloth fits the 27B model in 24 GB VRAM with no accuracy loss.

  3. Unsloth AIAI score34

    You can now fine-tune Qwen3.8-27B for free with our notebook! ๐Ÿ”ฅ Local training works on 24GB VRAM. Unsloth trains Qwen3.8-27B 1.5x faster with 50% less VRAM. GitHub: https://github.com/unslothai/unsloth Qwen3.8-27B Notebooks + Guide: https://unsloth.ai/docs/models/qwen3.8/train

    You can now fine-tune Qwen3.8-27B for free with our notebook! ๐Ÿ”ฅ Local training works on 24GB VRAM. Unsloth trains Qwen3.8-27B 1.5x faster with 50% less VRAM. GitHub: https://github.com/unslothai/unsloth Qwen3.8-27B Notebooks + Guide: https://unsloth.ai/docs/models/qwen3.8/train

Aug 24

Aug 24Mon
  1. Google ยท new models on Hugging FaceAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    Google Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Epoch AI ยท The Epoch BriefAI score58

    Epoch AI says US GDP underestimates AI growth by missing Nvidia's value

    Epoch AI argues US GDP growth over the last year was underestimated by about 0.3 percentage points because value from fabless chipmakers like Nvidia goes unrecorded. The report says no goods export, IP export, service export, or merchanting category captures Nvidia's value-add, and the Bureau of Economic Analysis confirmed the analysis. If Nvidia's growth continues, the gap could reach almost two percentage points per year by 2028.

  3. Engineering at MetaAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    Meta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    AIWhy it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.

  4. Microsoft ResearchAI score34

    Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5

    Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. https://msft.it/6011aOXe5

  5. GeneralistAI score38

    We've reduced the time it takes to go from physical prompt โ†’ robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.

    We've reduced the time it takes to go from physical prompt โ†’ robot behavior. The faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.

Aug 23

Aug 23Sun

Aug 21

Aug 21Fri
  1. Jim FanAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    NVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

  2. Amazon ScienceAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    Amazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

Aug 20

Aug 20Thu
  1. Ali GhodsiAI score20

    This is true. It wasn't actually possible before 2010 because datacenter networks would bottleneck. We used to design coupled storage/compute where you brought compute close to "big data". But research on "full bisection bandwidth" networks made it possible to essentially just talk from any machine to the storage system at full speed. The disaggregation started then! Databricks and Snowflake started soon after many others followed. Now "Put it on the object store" is the way to go.

    This is true. It wasn't actually possible before 2010 because datacenter networks would bottleneck. We used to design coupled storage/compute where you brought compute close to "big data". But research on "full bisection bandwidth" networks made it possible to essentially just talk from any machine to the storage system at full speed. The disaggregation started then! Databricks and Snowflake started soon after many others followed. Now "Put it on the object store" is the way to go.

  2. Mistral AIAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    Mistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Ali GhodsiAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    Databricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

  2. GeneralistAI score42

    To us, GEN-1.5 represents a new frontier of generality โ€” one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog: https://generalistai.com/blog/gen-1.5

    To us, GEN-1.5 represents a new frontier of generality โ€” one that challenges our own understanding of how these models behave when pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog: https://generalistai.com/blog/gen-1.5