Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. eric zakariassonAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  2. GitHub Blog · AI & MLAI score46

    Copilot app rebuilds pull request view to render a 2,200-file diff smoothly

    AIGitHub rebuilt the pull request view in the GitHub Copilot app to keep review fast on very large diffs, testing it on an open source pull request with 2,200 files, over a million changed lines, and more than 400 inline review comments. The core difficulty is that review comment heights can only be measured at render time, which breaks the fixed-geometry virtualization used for code-only diffs. GitHub split the document height into a deterministic code domain and a separately measured domain for comment blocks.

  3. eric zakariassonAI score36

    Optimizing reading for AI agents cuts context-gathering costs

    AIEric Zakariasson argues that agents spend heavily on reading context before and after work, so optimizing that reading makes a major difference. He recommends the linked guide to builders, or handing it to an agent to implement its findings. Cursor's related post reports 7% lower token costs with no drop in agent quality, achieved through tighter prompts, selective tool loading, better caching, and compressed file reads.

    Image from @ericzakariasson's post
  4. Microsoft ResearchAI score60

    Microsoft Research shows offloading robot AI inference improves performance and battery life

    AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.

    Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.

Sep 22

Sep 22Tue
  1. TinkerAI score25

    Tinker fine-tunes Qwen3.6 for Jev-style probability prompts in 10 minutes

    AITinker says an open LLM can serve a Jev-like interface that takes discrete options and returns fast probabilities, since next-token prediction is already a probabilistic classifier. A post by @ekzhang1 reports that a $5, 10-minute supervised fine-tuning run on Tinker improved Qwen3.6-35B-A3B's handling of Jev-style prompts, with +8% on GPQA Diamond and +12% on MMLU-Pro.

  2. Together AI BlogAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  3. Alex AlbertAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  4. Unsloth AIAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  5. OpenBMBAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

Sep 21

Sep 21Mon
  1. Tencent HyAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  2. xAI News (Grok)AI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  3. Mike KnoopAI score38

    Mike Knoop says LLM logprobs are vanishing, yet they enable useful new patterns

    AIMike Knoop notes that logprobs used to be widely exposed by LLM inference APIs and sees the market maturing so that parts of the LLM stack can be packaged in new, useful ways. He links this to Bryan Helmig's post on prompting with max_tokens: 1 plus logprobs for fast, parallel judgments, which Helmig says has a lot more depth than he expected.

Sep 20

Sep 20Sun