Cohere explains how hard training samples improved North Small Translate
Original titleWhen preparing the data, we ran into a problem: after the first step of training, the model could translate over 90% of the training docu...
AISummary
Cohere reports that after the first training step, its model could already translate over 90% of the training documents, which created a data problem. Kocmi describes how the most difficult samples were used to strengthen North Small Translate's capabilities. This post is part 3 of a six-part thread.
Source: Cohere · x.comPublished · added here