Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi
Original titleSpeed, Scale, & Privacy
AISummary
Liquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens.
The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages.
LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.
Source: Liquid AI Newsletter · liquidai.substack.comPublished · added here