Cohere's North Small Translate evaluated on post-release WMT26 benchmarks
Original titleBenchmarking: We used WMT26 benchmarks, which were released after we created the model, ensuring we couldn't train North Small Translate ...
AISummary
Cohere evaluated North Small Translate on WMT26 benchmarks, which were released after the model was created, so the model could not have been trained on them. This timing is presented as a safeguard against training contamination in the reported measurements.
Source: Cohere · x.comPublished · added here