Byte-level language models process raw bytes instead of subword tokens
Original titleMost language models split text into subwords—words or fragments of words from a fixed vocabulary. This can obscure spelling details acro...
AISummary
Most language models split text into subwords from a fixed vocabulary, which can obscure spelling details and fragment meaningful units in code or math. Byte-level models instead work directly with the bytes computers use to represent text.
Source: Ai2 · x.comPublished · added here