Skip to contentSkip to stories

Updated

All AI news

Sep 27

Sep 27Sun
  1. Xiaomi MiMoAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. Sebastian RaschkaAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

Sep 25

Sep 25Fri
  1. LMSYS OrgAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

  2. GitHub Blog · AI & MLAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

  3. Google Cloud · AI & Machine LearningAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  4. Amazon ScienceAI score38

    Amazon and Reactor build kernel path to real-time video generation on Trainium

    AIUsing the Neuron Kernel Interface, Reactor and Amazon's Neuron Science team built a kernel-centric path to real-time autoregressive diffusion video generation on Trainium. They addressed dynamic shapes, memory access patterns, and cache management, which are hard for generic compilers, and developed techniques intended to generalize across models.

Sep 24

Sep 24Thu
  1. LlamaIndexAI score17

    LlamaIndex Explains Using Confidence Scores to Control Document Extraction Automation

    AILlamaIndex argues that extraction confidence scores are useful only when they help decide what can be automated and what needs human review. Using ExtractBench, the post compares extraction systems after confidence filtering, reporting that LlamaParse Agentic Plus reached 66.48% recall on expected fields at a 97% precision target. The post covers confidence cutoffs, precision versus recall, score coverage, score granularity, and human review volume.

  2. Microsoft Foundry BlogAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  3. Lovable BlogAI score80

    How Lovable's Chats connect conversations to agent work on projects

    AILovable describes how its Chats feature lets a workspace-level chat agent hand work to project builder agents and receive progress back. The design records each agent's history as an append-only, forkable trajectory, and passes messages through durable inboxes that activations wake. Agents can suspend at iteration boundaries and resume on freshly deployed nodes without killing long-running runs.

    Why it matters: The post details how trajectories, inboxes, and activations let agents share work and resume after deploys, useful for designing comparable agent systems.

  4. Kling AI BlogAI score12

    Kling AI outlines six AI video limitations and workarounds for consistency and control

    AIKling AI's blog identifies six limitations of current AI video generation, including temporal consistency, character consistency across shots, unrealistic physics, long-form generation, fine details and text, and prompt control. It recommends workarounds such as reference images, shorter single-action clips, storyboards, and adding text or logos in post. The article says Kling VIDEO 3.0 and VIDEO 3.0 Omni offer reference-based subject consistency to help reduce these problems.

  5. Kling AI BlogAI score8

    Six Best Watermark Remover Tools for Cleaner Photo Edits Compared

    AIThis guide compares six watermark removal tools, including Kling AI, HitPaw Watermark Remover, Picsart, Adobe Photoshop, Fotor, and Cleanup.pictures, based on mark type and editing control. Kling AI's IMAGE 3.0 uses natural-language prompts and annotated images to rebuild marked areas in context, while IMAGE 3.0 Omni adds refinement with native 2K/4K output.

Sep 23

Sep 23Wed
  1. Boris ChernyAI score30

    More details for the formal methods people -- what's happening is Claude is doing something like: 1.

    AIBuilding a model of the program, targeting a tricky state machine or race-prone part of the code 2. Finding counter-examples in the model. These are suspected bugs 3. Reproducing the bugs 4. Fixing the bugs in the code It's not that the whole codebase is formally verified (yet!..), more that the hairiest parts of the code are modeled, checked for counter-examples, and fixed.

  2. eric zakariassonAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  3. GitHub Blog · AI & MLAI score46

    Copilot app rebuilds pull request view to render a 2,200-file diff smoothly

    AIGitHub rebuilt the pull request view in the GitHub Copilot app to keep review fast on very large diffs, testing it on an open source pull request with 2,200 files, over a million changed lines, and more than 400 inline review comments. The core difficulty is that review comment heights can only be measured at render time, which breaks the fixed-geometry virtualization used for code-only diffs. GitHub split the document height into a deterministic code domain and a separately measured domain for comment blocks.

  4. Microsoft ResearchAI score60

    Microsoft Research shows offloading robot AI inference improves performance and battery life

    AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.

    Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.

Sep 22

Sep 22Tue
  1. Together AI BlogAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  2. Alex AlbertAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  3. Unsloth AIAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

  4. OpenBMBAI score20

    Thanks so much for sharing this!

    AIReally cool to see MiniCPM5-2B being used in a practical multi-agent workflow like this — especially with the workers actually handling matching, short payments, duplicate references, and disputes through tool calls. Appreciate all the work you put into testing and documenting this. Such a nice case for the community 🙌