Dust pretrains transformers with zeroth-order optimization, approaching backprop results
Original titleCool research from @industriaalist. It's related to one of my favorite papers from @guyvdb https://arxiv.org/abs/2205.11502 which shows ...
AISummary
Dust is a zeroth-order method that pretrains transformers and sometimes matches or exceeds backprop given large compute. The authors report it is about 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. The post also cites the gradient-alignment result up to 1B tokens and the virtual population idea for scaling.
Source: Mike Knoop · x.comPublished · added here