The same recency bias also appears in models using Kimi Delta Attention, which controls how much earlier information a model keeps. As it mixes new information into that memory, nearby words carry more shared information, creating clues about how far apart they are.
Kimi Delta Attention's recency bias shapes how models retain earlier information
AISummary
Models using Kimi Delta Attention show a recency bias, as the mechanism controlling how much earlier information is kept shapes memory. As new information mixes into that memory, nearby words share more information, which gives the model clues about how far apart words are.
Post on XView on X
Source: Zyphra · x.comPublished · added here
