Skip to content
Read the original: Baseten Blog· Published Pick70/100AI score70/100

Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

Original titleAgentic inference optimization: 50-90% faster engines

AISummary

Baseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs.

943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

AIWhy it matters

The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

Read the original baseten.co

Source: Baseten Blog · baseten.coPublished · added here