Tri Dao Praises Etched's Fast Inference Chip Design for LLM Serving
Original titleIt's wild how quickly Etched designed and got the chips out, all within 2 years. They went deep, hardcoding attention into silicon and ge...
AISummary
Tri Dao says Etched designed and produced its chips within two years by hardcoding attention into silicon and reaching high MFU. He expects hardware built for LLM inference to cut the cost of intelligence by 10x. The quoted Etched post says it has built its first racks after an A0 tapeout, raised $800m, holds $1B+ in customer contracts, and plans to ship the racks this summer.
Source: Tri Dao · x.comPublished · added here