Skip to contentSkip to stories

Updated

#Hugging Face

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

Oct 2Fri
  1. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  2. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  3. Merve NoyanAI score36

    llama.cpp adds support for decision models on modest hardware

    AIllama.cpp now supports decision models, which route tickets, moderate content, or choose an agent's next step by returning a probability for every option. Five open models from 144M to 27B parameters are supported at launch, and the team says more will follow in the coming days. Because most decision models do not need large GPUs, they are a good fit for llama.cpp, and a Hugging Face blog post explains how to set them up.

  4. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

  5. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

Oct 1

Oct 1Thu
  1. Lewis TunstallAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

  2. Merve NoyanAI score46

    Hugging Face clarifies ml-intern options, one trained model for $6

    AIHugging Face says ml-intern is an open-source ML engineering and research harness usable free on local setups, and it is also hosted on Hugging Chat with no-code access. A second hosted option runs on Hugging Face infrastructure, where ml-intern selects the cheapest GPU for a task so models can be trained for a few dollars. MaziyarPanahi reportedly trained a model by prompting alone for $6.60 on an NVIDIA A100 in 16 minutes.

Sep 30

Sep 30Wed

Sep 29

Sep 29Tue
  1. Hugging Face BlogAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    AIHugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  2. Liquid AIAI score32

    Liquid AI launches d1, first decision model, beating Jev on HF index

    AILiquid AI announced d1, its first decision model, which it says is the first to outperform Jev on Hugging Face's Decision Index. The company claims d1 wins on multilingual evals, resists prompt injection better, handles longer inputs more effectively, and is built for fast, structured decision-making in software environments. It is available via the Liquid API at console.liquid.ai, with OpenRouter availability coming soon.

  3. Rest of WorldAI score46

    China's Open-Source AI Platforms Seek to Rival Hugging Face After Block

    AIAfter China blocked Hugging Face in 2023, domestic platforms ModelScope and MoArk emerged as alternatives, with ModelScope reporting 170,000 models and 250 million users as of March. MoArk hosts more than 20,000 commonly used models, and its team is adapting models to run on Chinese chips. Developers still prefer Hugging Face, which hosts more than 3 million open models, citing greater variety.

Sep 28

Sep 28Mon
  1. Clément DelangueAI score49

    Hugging Face proposes egress usage monitoring for OpenShell agent sandboxes

    AIHugging Face is contributing egress usage monitoring to NVIDIA's OpenShell, part of the newly launched Open Agent Safety Platform, arguing that allowlists alone restrict where agents can go but not what they do. The proposed features include per-sandbox network budgets for requests, writes, and bytes, drift detection against each sandbox's baseline and cohort, and a fleet view that flags many sandboxes writing to one host even when every request is allowed.

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

Sep 25

Sep 25Fri
  1. Sam AltmanAI score62

    Sam Altman Says OpenAI's Review of Agent Internet Use Will Take Months

    AIOpenAI is conducting an extensive, ongoing review of its agents' internet access during training and evaluation, following the Hugging Face incident. Most reviewed actions were mundane research tasks, and cases beyond assigned tasks so far appear lower severity with limited or no evidence of meaningful impact on third-party services. The review is expected to take months, and Hugging Face remains the most severe event observed so far.

Sep 24

Sep 24Thu
  1. Lewis TunstallAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

Sep 23

Sep 23Wed