Tech News
← Home  ·  All topics

Visual Tokens

1 GoKawiil brief on this topic

Apple researchers unveil LensVLM-9B, a vision-language model for reading compressed text images

Apple has introduced LensVLM, built on Qwen3.5-9B-Base, which lets vision-language models scan compressed text-as-image inputs and selectively expand only relevant portions using learned tools rather than reading everything at full resolution. The model reportedly matches full-text accuracy at 4.3x compression and beats retrieval and compression baselines up to 10.1x compression across seven QA benchmarks, plus document and code tasks.