Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support
Original titleEnterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU
AISummary
Google Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs.
The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.
Source: Google Developers Blog · developers.googleblog.com