OpenCode Senses
Vision for text-only OpenCode models — fully local, private, free, and slightly smug about it.
Senses adds a vision layer to OpenCode so any image becomes useful to your text-only coding model: screenshots get exact OCR, objects get located, colors get measured, and everything comes back as structured evidence the model can reason over. No API keys. No hidden cost. No pictures of your desktop leaving the machine — we promise, scouts' honor, we checked the source.
The text model reasons. Senses perceives, grounds, verifies, and fetches snacks.
Table of contents
Open source stays alive on donations, not vibes. Every coffee funds the next vision feature.
Features
Sight for text-only models — attach any image and the model sees it: a structured scene read, a caption, and exact OCR are auto-injected into your message before the model responds.
— attach any image and the model sees it: a structured scene read, a caption, and exact OCR are auto-injected into your message before the model responds. 13 grounded tools — inspect, OCR, detect, point, segment, crop, zoom, colors, diff, annotate, metadata, reverse search, status. All return normalized, source-grounded evidence, the way the vision gods intended.
— inspect, OCR, detect, point, segment, crop, zoom, colors, diff, annotate, metadata, reverse search, status. All return normalized, source-grounded evidence, the way the vision gods intended. Web images supported everywhere — any tool accepts an https:// URL as a path ; the image is downloaded verbatim (original type and bytes preserved) and cached locally.
... continue reading