Skip to contentSkip to stories

Updated

#Tutorial/How-to

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 17

Sep 17Thu
  1. Gemini NotebookAI score37

    Mariposa Museum exhibit shows town in 1859, a decade after Gold Rush

    AIA Mariposa Museum exhibit photographed by writer Steven Johnson, shared by Gemini Notebook, documents the Sierra Nevada town in 1859, ten years after the Gold Rush began. Johnson says he used the Gemini Notebook mobile app's camera feature to generate a detailed report from photos of the display, which he says was 99% accurate on fact-checking.

  2. Z.aiAI score40

    GLM-5.3 helped build the inference stack serving GLM-5.3-Flash

    AIZ.ai reports that GLM-5.3 helped build and optimize the inference infrastructure for GLM-5.3-Flash. The system went from first successful run to production readiness in under two weeks, with end-to-end throughput tripling over the initial baseline. The team credited dense feedback from local correctness tests, execution traces, microbenchmarks, and end-to-end measurements for enabling targeted hypothesis testing.

Sep 16

Sep 16Wed
  1. TinkerAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  2. Google for DevelopersAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

Sep 15

Sep 15Tue
  1. Google · Innovation & AIAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

Sep 14

Sep 14Mon
  1. Google Developers BlogAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  2. vLLM BlogAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  3. LlamaIndex 🦙AI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

    Image from @llama_index's post
  4. Kilo (acq. by Anaconda)AI score20

    Hands-on guide to writing evals that catch false agent claims

    AIA hands-on guide by @pandemicsyn walks through writing evals that detect when an AI agent claims to have completed a task it never did. Working through a demo agent that fails on purpose, the author refines the checks until they can distinguish real work from mere claims of work. The post includes a coding agent skill that can guide readers through the exercise.

  5. Tencent · new models on Hugging FaceAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

Sep 13

Sep 13Sun
  1. Sebastian RaschkaAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post

Sep 11

Sep 11Fri
  1. Augment Code BlogAI score80

    Augment Code details how its software factory raised output per developer 4.5×

    AIAugment Code reports that size-adjusted output per active developer rose from 12.3 to 55.7 between November 2025 and July 2026, while median time to merge fell from 11.2 to 3.1 hours. The post says the company added specialized agents wherever work was piling up, across planning, review, verification, feedback, and incident response, and kept engineers responsible for product decisions, architecture, and production risk.

    Why it matters: The post pairs internal productivity and quality metrics with the order in which agents were added, showing how review and verification bottlenecks shaped a software delivery pipeline.

Sep 10

Sep 10Thu
  1. Google Developers BlogAI score55

    Google details autonomous LLM post-training loops using Tunix on TPUs

    AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.

Sep 9

Sep 9Wed
  1. Fireworks AI BlogAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  2. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  2. InferactAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

Sep 6

Sep 6Sun
  1. Sebastian RaschkaAI score22

    Raschka's Reasoning From Scratch video covers LLM text generation and KV caching

    AISebastian Raschka released a video in his Reasoning From Scratch series covering text generation in LLMs and KV caching. The walkthrough uses a pretrained Qwen3 model from the Reasoning From Scratch package, covering tokenization, greedy decoding, end-of-sequence handling, and a benchmarked KV caching speedup. It prepares the base model for reasoning techniques in later episodes.

    Video from @rasbt's post

Sep 4

Sep 4Fri
  1. Andrew NgAI score42

    Andrew Ng maps key skills for using AI coding agents effectively

    AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.

Sep 3

Sep 3Thu
  1. TinkerAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  2. Google Developers BlogAI score23

    Google's Gemini Enterprise DevEx sprint fixes governance setup friction for agents

    AIGoogle's Gemini Enterprise developer experience team tested agent governance workflows without internal shortcuts and fixed friction points across its agent governance products. Fixes included documentation stating that enabling the Identity-Aware Proxy API is a hard requirement, auto-allowing essential Google-managed platform APIs in the Agent Gateway, and adding Private Service Connect and Cloud DNS setup guidance for Semantic Governance. The team also published ready-made Logs Explorer queries for monitoring Agent Gateways and Content Security.

  3. Matei ZahariaAI score34

    Databricks uses Unity AI Gateway traces to cut AI waste fast

    AIDatabricks used Unity AI Gateway tracing and Genie One to find seven small MCP-server bugs and eliminate an estimated $1.2M in annual wasted AI spend and lost productivity within an hour. The bugs drove about $499K per year in wasted tokens, roughly 12,000 engineering hours per year in agent wait time, and 1,409 tool errors in a single 24-hour window. Matei Zaharia argues that analyzing tracing data for AI workloads will become a routine form of operational data analysis across companies, much like finance and security.

  4. Prime Intellect BlogAI score59

    Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds

    AIPrime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting. The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally. Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.

Sep 2

Sep 2Wed
  1. Daniel HanAI score34

    Stanford's Modern Software Developer course adds AI-native engineering curriculum

    AIMihail Eric announced the 2026 edition of his Stanford course "The Modern Software Developer," with 85% of the Fall 2025 material replaced by AI-native topics such as agent skills, context engineering, and agentic code review. Students will ship pull requests to real open-source AI repositories, with partners including Browserbase, HeyGen, and CopilotKit offering mentorship.