Skip to contentSkip to stories

Updated

#Tutorial/How-to

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

Oct 7

Oct 7Wed

Oct 2

Oct 2Fri
  1. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

Oct 1

Oct 1Thu
  1. OpenRouter BlogAI score52

    How agent frameworks handle tool-calling schemas across model providers

    AITool definitions and tool-call responses differ between OpenAI, Anthropic, and Google, so a tool that works on one model may fail on another. The article compares six agent frameworks, including LangChain, CrewAI, and the OpenAI Agents SDK, by where each performs schema translation. It also describes OpenRouter's API-layer normalization, which accepts an OpenAI-style tools array and returns a standard tool_calls response for tool-capable models.

Sep 29

Sep 29Tue
  1. Google Developers BlogAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.

Sep 25

Sep 25Fri

Sep 20

Sep 20Sun

Sep 3

Sep 3Thu
  1. TinkerAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

Jun 25

Jun 25Thu
  1. PaddlePaddleAI score38

    PP-OCRv6 recognition uses CTC and NRTR heads to curb hallucination

    AIPP-OCRv6's recognition module uses a CTC plus NRTR dual-head design so text is decoded from visual features rather than language priors, reducing hallucination. In hallucination tests, PP-OCRv6_medium reaches 93.2%, versus 85.0% for the best VLM, and recognition accuracy across 15 scenarios is 83.2%, above PP-OCRv5_server's 78.1%. NRTR is used only during training, adding language regularization at no inference cost, and it contributes +1.16% accuracy.

    Image from @PaddlePaddle's post

Jun 23

Jun 23Tue
  1. Lil'Log (Lilian Weng)AI score40

    Scaling Laws, Carefully: Early Empirical Power-Law Studies of Loss, Data and Model Size

    AILil'Log examines early empirical work showing that deep learning generalization error follows power-law curves as training data and model size grow. Hestness et al. (2017) found the exponent reflects the problem domain rather than the architecture, while Rosenfeld et al. (2020) modeled loss jointly as a function of model size N and data size D, fitting parametric forms on small configurations to extrapolate to larger ones.

  2. PaddlePaddleAI score38

    PP-OCRv6 lightweight OCR model challenges large VLMs with 34.5M params

    AIPaddlePaddle introduced PP-OCRv6, a lightweight OCR architecture built on the LCNetV4 backbone, in the first episode of its tech deep dive series. The post says PP-OCRv6_medium reaches 86.2% detection Hmean and 83.2% recognition accuracy, surpassing PP-OCRv5_server while running faster. Three model specs—Tiny, Small, and Medium—target edge CPU devices, balanced deployment, and industrial high-accuracy pipelines.

    Image from @PaddlePaddle's post

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    AIAi2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Oct 27, 2025

Oct 27, 2025Mon

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.