Skip to content

#Hugging Face

Oct 8

TodayOct 8Thu9 items
  1. TechCrunch · AI62

    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    Goodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

  2. Prime Intellect52

    Alzheimer's Translation Challenge launches with 150M cell atlas for AI hypothesis discovery

    Prima Mente and AlzData are launching the Alzheimer's Translation Challenge, a global AI competition to discover new therapeutic hypotheses for Alzheimer's disease. The challenge centers on a 150M cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. Top teams will have their hypotheses tested in Prima Mente's wet lab, and the data will be available through the AD workbench, Hugging Face, and Prima Mente's modeling platform.

  3. Goodfire22

    The Alzheimer’s Translation Challenge is led by @primamente and @AlzData, with @nvidia, @huggingface, @nebiusai, Talisman Therapeutics, @ultimagenomics, @Cellanome, @PrimeIntellect, and @boltz_bio. Registration is open now. Competition starts spring 2027: https://primamente.com/challenges

    The Alzheimer’s Translation Challenge is led by @primamente and @AlzData, with @nvidia, @huggingface, @nebiusai, Talisman Therapeutics, @ultimagenomics, @Cellanome, @PrimeIntellect, and @boltz_bio. Registration is open now. Competition starts spring 2027: https://primamente.com/challenges

  4. Leandro von Werra70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    Why it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  5. Clément Delangue22

    We urgently need more public traces of AI agents attacking and defending systems. Defenders can’t learn from what they can’t see. If you have traces and are being pressured to keep them private, my DMs are open. Let’s level the playing field and fight the asymmetry and lack of transparency in AI!

    We urgently need more public traces of AI agents attacking and defending systems. Defenders can’t learn from what they can’t see. If you have traces and are being pressured to keep them private, my DMs are open. Let’s level the playing field and fight the asymmetry and lack of transparency in AI!

  6. Merve Noyan37

    new Llama.cpp release ships with (multimodal!) Jev-like models support, performance upgrade for Metal and more! 🔥 super simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here https://huggingface.co/models?apps=llama.cpp&other=decision-model&sort=trending we also polished Llama App website & docs https://llama.app 🌟

    new Llama.cpp release ships with (multimodal!) Jev-like models support, performance upgrade for Metal and more! 🔥 super simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here https://huggingface.co/models?apps=llama.cpp&other=decision-model&sort=trending we also polished Llama App website & docs https://llama.app 🌟

  7. Air Street Press60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    Nathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

Oct 7

Oct 7Wed
  1. MarkTechPost58

    Unsloth Studio re-checks changed model repos and blocks flagged weights before loading

    Unsloth Studio binds remote-code approval to a fingerprint of the scanned code, so changed code requires fresh consent before it runs. It also blocks weight files that Hugging Face has flagged for malware in the path the selected loader would deserialize. The article describes these checks as one layer among several, alongside package-content scans and OS sandboxes, and notes that the scanner is not a sandbox and cannot catch every evasion.

  2. Hugging Face Blog66

    How one developer built six custom models with ML-Intern for about USD 103

    A Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  3. Hugging Face Blog49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    Liquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  4. Perplexity45

    We're releasing pplx-embed-v2-late, two late-interaction embedding models that retrieve text, images, and pages with a shared embedding space for cross-model querying. Both models achieve frontier performance and are publicly available on Hugging Face. https://www.perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector

    We're releasing pplx-embed-v2-late, two late-interaction embedding models that retrieve text, images, and pages with a shared embedding space for cross-model querying. Both models achieve frontier performance and are publicly available on Hugging Face. https://www.perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector

  5. LlamaIndex47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    LlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

  6. Ai239

    We byteified Qwen 3 8B & Llama 3 8B to create Bwen 8B & Blama 8B. Both come close to matching their source models' performance in our evaluations. Bwen 8B also outperforms Bolmo 7B across our aggregate evaluation suite. https://huggingface.co/collections/allenai/bolmo

    We byteified Qwen 3 8B & Llama 3 8B to create Bwen 8B & Blama 8B. Both come close to matching their source models' performance in our evaluations. Bwen 8B also outperforms Bolmo 7B across our aggregate evaluation suite. https://huggingface.co/collections/allenai/bolmo

  7. Hugging Face Blog53

    TII releases Falcon-ASR, a 1.6B speech recognition model focused on Emirati Arabic

    The Technology Innovation Institute introduces Falcon-ASR, a 1.6 billion parameter speech recognition model for Arabic with a focus on the Emirati dialect. On six Arabic test sets it reports an average word error rate of 20.92%, versus 23.17% for the best published leaderboard result it compared against. The model also transcribes English, French, Spanish and Portuguese with the same weights, and a demo Space is available while API access and native apps are planned.

  8. Hugging Face Blog78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    NVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

Oct 6

Oct 6Tue
  1. Philipp Schmid70

    EmbeddingGemma 2 releases native multimodal embeddings built on Gemma 4

    Google releases EmbeddingGemma 2, its first native multimodal embedding model, built on Gemma 4 under Apache 2.0. It embeds over 100 languages, code, images, audio, and video into one vector, with an 8,192-token context and four sizes from 270M to 740M parameters. Matryoshka output dimensions of 768, 512, 256, or 128 are supported, and the model is available in Sentence Transformers and LiteRT-LM, with a reported 14% gain on MTEB Code.

    Why it matters: The release extends an embedding model to text, code, images, audio, and video in one vector, a useful option for retrieval systems that mix media types.

  2. Paige Bailey54

    EmbeddingGemma 2 launches as an Apache 2.0 multimodal embeddings model

    Google's EmbeddingGemma 2 is an open embeddings model for on-device use that covers code, image, video, audio, and text. It comes in modular sizes from 270M text/code to 740M full multimodal, supports Matryoshka truncation down to 128 dimensions, and reports a 14% gain on MTEB Code over v1 under an Apache 2.0 license. The author's post highlights the release and a Hugging Face demo, while the benchmark table compares it with several models.

  3. Google for Developers36

    — Weights are live on @HuggingFace and @Kaggle — Get out-of-the-box support for LiteRT, MediaPipe, @LangChain, and @llama_index — Start Mapping: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

    — Weights are live on @HuggingFace and @Kaggle — Get out-of-the-box support for LiteRT, MediaPipe, @LangChain, and @llama_index — Start Mapping: https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

  4. Julien Chaumond70

    Mistral Large 4 announced with open weights due end of October

    Julien Chaumond reposted Mistral's announcement of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters. Mistral says it is available via API today, with open weights scheduled for release at the end of October, and is working privately with cybersecurity partners.

    Why it matters: The post lays out Mistral Large 4's scale, multimodal design, and availability timeline, which helps readers gauge the open-weights landscape outside China.

Oct 5

Oct 5Mon
  1. NVIDIA AI39

    The model behind this result is now on @huggingface 🤗 Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra. https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

    The model behind this result is now on @huggingface 🤗 Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra. https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4

  2. Clément Delangue72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    Reflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    Why it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.

  3. Google AI46

    Gemma 4 and BOTANIC-1 pinpoint crop-trait DNA mutations in hours

    Living Models paired Google's Gemma 4 with BOTANIC-1, a plant-DNA-trained AI, to identify the mutations behind crop traits. In one melon-yield test, the pipeline ranked the target mutation first among 2,494 candidates in under four minutes. The approach aims to speed development of climate-resilient crops that traditionally take years of field trials.

  4. IEEE Spectrum · AI49

    Human Oversight of AI Agents Could Fail as Approval Processes Push People Out

    Researchers Avijit Ghosh, Margaret Mitchell, and Samir Passi argue in a September 6 arXiv paper that current human-in-the-loop designs for AI agents push humans out of meaningful oversight. They say agents are tuned for speed, accuracy, and volume, overwhelming reviewers, and recommend adding friction, such as requiring users to state their own choice first, to counter automation bias and fatigue.

  5. Julien Chaumond18

    What's your P(doom)? AI Safety is an important subject, but existential-risk debate is ill-defined, and dominated by highly-visible voices who have a significantly higher-than-average estimates of x-risk probabilities. So we would like to hear from a broader cross-section of the AI community ⤵️

    What's your P(doom)? AI Safety is an important subject, but existential-risk debate is ill-defined, and dominated by highly-visible voices who have a significantly higher-than-average estimates of x-risk probabilities. So we would like to hear from a broader cross-section of the AI community ⤵️

  6. Clément Delangue62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    Hugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

  7. Liquid AI · new models on Hugging Face67

    Liquid AI releases d1-3B, a 3B multimodal decision model for edge deployment

    Liquid AI has released d1-3B, a 3B parameter multimodal model post-trained to return calibrated, typed answers to yes/no, choice, and score questions in one forward pass. The source reports a Decision Index 0.2.1 score of 48.57, the highest among models under 10B in its table, and 8 ms per decision on an NVIDIA RTX 4090.

    Why it matters: The source gives benchmark scores against named peer models and edge latency figures across several hardware targets, helping readers judge fit for on-device decision pipelines.

Oct 4

Oct 4Sun
  1. SemiAnalysis22

    SemiAnalysis says NVIDIA's SchedMD acquisition hurt SLURM support for non-NVIDIA chips

    After NVIDIA acquired SchedMD, the SLURM scheduler's support for non-NVIDIA chips has allegedly worsened, and AMD built a competing scheduler called spur. The author says NVIDIA has not kept SLURM hardware neutral despite its earlier pledge, and questions whether Hugging Face will face the same fate after NVIDIA's acquisition of it.