Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)AI score38

    Ai2 Panel Identifies Five Hard Challenges for AI-Assisted Science

    AIAt an August 27 Ai2 event on expanding its work with the Paul G. Allen Research Center at Providence Swedish Cancer Institute, panelists identified five persistent challenges for scientific AI. The main ones are keeping AI steerable as research evolves, deciding which tasks to delegate, and avoiding the amplification of weak study design or bad data.

  2. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  3. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

Aug 31

Aug 31Mon

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Jazzyear · InsightsAI score40

    Helical Fusion's Stellarator Design Uses AI to Cut Parameters to Three

    AIHelical Fusion CTO Wei Xishuo told the NFEC2026 AI-for-fusion forum that the company uses an autoencoder to compress hundreds of stellarator shape parameters into three. The company says this lets it predict zonal flow residuals and turbulent transport from more than 15,000 global simulations and generate new configurations with up to 100x better confinement in simulation.

  3. Fireworks AI BlogAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

Aug 29

Aug 29Sat
  1. Chips and CheeseAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    AISamsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Meituan LongCatAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. RadixArkAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

    Image from @radixark's post
  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    AITencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Jazzyear · ArticlesAI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    AIIn an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.

  3. Google Developers BlogAI score42

    Google Developers Blog explains deep learning with Keras for astroparticle physics data analysis

    AIThe Google Developers Blog post describes how deep learning can analyze the large, image-like sensor data from astroparticle observatories such as the Pierre Auger Observatory and IceCube. The author argues these methods could improve instrument sensitivity and reveal patterns in cosmic-ray and neutrino signals that traditional analysis techniques miss.

  4. Amazon ScienceAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    AIAmazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

  5. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  6. Ai2 · new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  7. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  8. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Google Developers BlogAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    AIGoogle Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.

  2. Daniel HanAI score34

    Fine-tune Qwen3.8-27B free on Kaggle with Unsloth QLoRA

    AIDaniel Han says users can fine-tune Qwen3.8-27B for free on Kaggle with a Google account, which provides 30 hours of GPU time on 2× Tesla T4s. Using QLoRA and Unsloth's kernels, the 27B model fits within 24 GB VRAM with no accuracy loss, according to the post. The background post from Unsloth adds that its notebook trains Qwen3.8-27B 1.5x faster with 50% less VRAM.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging FaceAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    AIGoogle Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Epoch AI · The Epoch BriefAI score58

    Epoch AI says US GDP underestimates AI growth by missing Nvidia's value

    AIEpoch AI argues US GDP growth over the last year was underestimated by about 0.3 percentage points because value from fabless chipmakers like Nvidia goes unrecorded. The report says no goods export, IP export, service export, or merchanting category captures Nvidia's value-add, and the Bureau of Economic Analysis confirmed the analysis. If Nvidia's growth continues, the gap could reach almost two percentage points per year by 2028.

  3. Engineering at MetaAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    AIMeta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    Why it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.

  4. Microsoft ResearchAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post