By spending lots of compute at indexing time, you can reuse that computation to get better performance without increasing inference-time compute!
Embeddings are the simplest mechanism that achieve this!
Aman Sanger of Cursor argues that heavy compute spent at indexing time can be reused to improve performance without raising inference-time compute, with embeddings as the simplest mechanism. Cursor's background post says semantic search improves its agent's accuracy across frontier models, especially in large codebases where grep alone falls short.
By spending lots of compute at indexing time, you can reuse that computation to get better performance without increasing inference-time compute!
Embeddings are the simplest mechanism that achieve this!
Semantic search improves our agent's accuracy across all frontier models, especially in large codebases where grep alone falls short. Learn more about our results and how we trained an embedding model for retrieving code.View quoted post on X
Source: Aman Sanger · x.comPublished · added here