Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model
Original titleCohereLabs/North-Micro-Vision-Instruct
AISummary
Cohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face.
The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic.
It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.
Source: Cohere · new models on Hugging Face · huggingface.coPublished · added here