Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Sep 27

Sep 27Sun
  1. Xiaomi MiMoOfficialAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi releases MiMo-V2.6-Pro-MOPD, a 1.02T-parameter sparse MoE model

    AIXiaomi has released MiMo-V2.6-Pro-MOPD, an upgrade of the MiMo-V2.6-Pro-RL checkpoint that fuses several domain-specialized teachers into one model via MOPD2 and targets tool-call repetition. The sparse MoE model has 1.02T total and 42B activated parameters, a 1M-token context length, and accepts text, image, video, and audio inputs. Weights are available on Hugging Face and ModelScope, with deployment recipes for SGLang and vLLM.

  2. Sebastian RaschkaXAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

    Video from @rasbt's post

Sep 25

Sep 25Fri
  1. LMSYS OrgOfficialAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

    Image from @lmsysorg's post
  2. eric zakariassonXAI score8

    Grok Bot helps developers build apps on the X API

    AIEric Zakariasson highlights a range of apps that can be built on the X API and says the Grok bot makes getting started easy. The developer exhibit at offers inspiration, and the Grok-hosted X API Engineer can help build, test, and deploy projects.

  3. GitHub Blog · AI & MLOfficialAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

  4. Google Cloud · AI & Machine LearningOfficialAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  5. Amazon ScienceOfficialAI score38

    Amazon and Reactor build kernel path to real-time video generation on Trainium

    AIUsing the Neuron Kernel Interface, Reactor and Amazon's Neuron Science team built a kernel-centric path to real-time autoregressive diffusion video generation on Trainium. They addressed dynamic shapes, memory access patterns, and cache management, which are hard for generic compilers, and developed techniques intended to generalize across models.

  6. LlamaIndex 🦙OfficialAI score18

    LlamaParse preserves complex Fed forecast tables for AI analysis

    AILlamaIndex tested LlamaParse on the Fed's September 2026 projections PDF, where a 2029 column's June comparison cell is blank. The company says all nine GDP median figures on page 2 matched the original PDF after parsing, with alignment kept in the returned HTML.

    Image from @llama_index's post

Sep 24

Sep 24Thu
  1. MidjourneyOfficialAI score26

    Midjourney releases tutorial video on its new edit model

    AIMidjourney has published a video explaining how to use its new edit model for maintaining consistent characters and other edits. The post offers no further details on features, pricing, or availability.

    Video from @midjourney's post
  2. LlamaIndex 🦙OfficialAI score17

    LlamaIndex Explains Using Confidence Scores to Control Document Extraction Automation

    AILlamaIndex argues that extraction confidence scores are useful only when they help decide what can be automated and what needs human review. Using ExtractBench, the post compares extraction systems after confidence filtering, reporting that LlamaParse Agentic Plus reached 66.48% recall on expected fields at a 97% precision target. The post covers confidence cutoffs, precision versus recall, score coverage, score granularity, and human review volume.

    Video from @llama_index's post
  3. Baseten BlogOfficialAI score44

    LangSmith Fine-Tuning Trains Open Models on Agent Traces via Baseten Loops

    AILangChain launched LangSmith Fine-Tuning, which lets users fine-tune open models on their LangSmith agent traces using the open-source smithtune CLI. Training runs on Baseten Loops in the user's own workspace, and smithtune deploy places the evaluated checkpoint on a Baseten Dedicated Inference deployment. Loops is in early access, so users may need to request access for their workspace.

  4. GitHub Blog · AI & MLOfficialAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

  5. Microsoft Foundry BlogOfficialAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  6. Philipp SchmidXAI score56

    Gemini 3.8 TTS adds custom voice creation from a short recording or prompt

    AIGemini 3.8 TTS lets users replicate their own voice or design a custom voice from a text prompt. The workflow is to record about 20 seconds of speech with a consent sentence, create the voice through an API call, then use it in any request with styles set in speech_metadata. The author also points readers to a guide for setting up and testing the process with an agent.

  7. Philipp SchmidXAI score62

    Gemini 3.8 Flash TTS adds custom voice creation from recordings or a sentence

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new option to replicate a user's own voice from two recordings or design one from a sentence. The guide says the reusable voice ID can be passed in later requests, or an encrypted voicekey that expires after 7 days can be used if nothing is stored server-side. Prompting changed from gemini-3.1-flash-tts-preview: input text is spoken word for word, delivery goes in speech_metadata.style, and non-streaming responses are now real WAV.

  8. Lovable BlogOfficialAI score80

    How Lovable's Chats connect conversations to agent work on projects

    AILovable describes how its Chats feature lets a workspace-level chat agent hand work to project builder agents and receive progress back. The design records each agent's history as an append-only, forkable trajectory, and passes messages through durable inboxes that activations wake. Agents can suspend at iteration boundaries and resume on freshly deployed nodes without killing long-running runs.

    Why it matters: The post details how trajectories, inboxes, and activations let agents share work and resume after deploys, useful for designing comparable agent systems.

  9. Kling AI BlogOfficialAI score12

    Kling AI outlines six AI video limitations and workarounds for consistency and control

    AIKling AI's blog identifies six limitations of current AI video generation, including temporal consistency, character consistency across shots, unrealistic physics, long-form generation, fine details and text, and prompt control. It recommends workarounds such as reference images, shorter single-action clips, storyboards, and adding text or logos in post. The article says Kling VIDEO 3.0 and VIDEO 3.0 Omni offer reference-based subject consistency to help reduce these problems.

  10. LangChain BlogOfficialAI score50

    LangSmith Fine-Tuning and smithtune Turn Agent Trajectories Into Custom Models

    AILangChain launched LangSmith Fine-Tuning and smithtune, a CLI that turns LangSmith agent trajectories into fine-tuned models through dataset creation, training with Fireworks or Baseten, and evaluation in LangSmith. smithtune currently supports supervised fine-tuning, training models on recorded examples of good agent behavior by updating model weights. The tool lets teams train specialized models without building the data pipeline by hand.

  11. Kling AI BlogOfficialAI score8

    Six Best Watermark Remover Tools for Cleaner Photo Edits Compared

    AIThis guide compares six watermark removal tools, including Kling AI, HitPaw Watermark Remover, Picsart, Adobe Photoshop, Fotor, and Cleanup.pictures, based on mark type and editing control. Kling AI's IMAGE 3.0 uses natural-language prompts and annotated images to rebuild marked areas in context, while IMAGE 3.0 Omni adds refinement with native 2K/4K output.

Sep 23

Sep 23Wed
  1. OpenClaw🦞OfficialAI score12

    OpenClaw adds live meeting notes and saved transcript tabs

    AIOpenClaw lets users follow notes while a Google Meet, Teams, Zoom, or voice capture is still running. The saved transcript can be opened in its own tab, though transcription may lag. Generated notes use the user's configured model and incur its usual charges.

  2. OpenClaw🦞OfficialAI score19

    OpenClaw lets agents transfer files to a paired computer and back

    AIOpenClaw lets users send files to a paired computer where their agent works, then return finished files in chat. Memory and supported Skills can also be stored on that paired computer. Setup requires host configuration, permissions, and matching OpenClaw versions.

  3. Simon WillisonXAI score34

    Datasette blog backup now queryable by voice via ChatGPT iPhone app

    AISimon Willison had a voice conversation with his blog's Datasette backup through datasette-mcp, which is now accessible from the ChatGPT iPhone app. The quoted post says ChatGPT Voice can now use plugins and run on GPT-6 Astra, Sol and Luna in ChatGPT Work on web and mobile.

  4. Philipp SchmidBlogAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  5. vLLM BlogOfficialAI score54

    vLLM adds distortion-free Gumbel-max watermarking for text provenance

    AIvLLM now supports Gumbel-max watermarking, which embeds a keyed signal into generated text without changing the expected token distribution. Detection requires the secret key and tokenizer, and the signal accumulates over longer outputs. Benchmarks on Qwen3.5-27B with MTP-3 show throughput changes between -1.1% and +2.0% across batch sizes, with no consistent slowdown.

  6. Boris ChernyXAI score30

    Claude models tricky code states to find and fix bugs

    AIClaude builds a model of a program's most complex parts, such as state machines or race-prone code, and searches that model for counterexamples that signal suspected bugs. It then reproduces those bugs and fixes them in the code. The post clarifies that the whole codebase is not formally verified, only the riskiest sections are modeled and checked.