Apple releases LensVLM-9B, a vision-language model for compressed text images
Original titleapple/LensVLM-9B
AISummary
Apple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form.
The repository provides a demo script and supports compression settings of 5x, 10x, and 15x.
Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.
Source: Apple · new models on Hugging Face · huggingface.coPublished · added here