K-Search Adapts CUDA Kernel Expertise to Apple Silicon MLX Backend
Original titleFrom CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
AISummary
Berkeley AI Research extended the K-Search evolutionary kernel framework with an MLX backend and a CUDA-to-MLX translation layer, letting it adapt existing CUDA kernels for Apple Silicon.
The team reports a 0.97x speedup relative to the native MLX Attention kernel and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel.
The method uses Gemini 3.5 Pro Preview to both reason about optimizations and write candidate kernels.
Source: Berkeley AI Research · bair.berkeley.eduPublished · added here