PerchAI
How it works

Images and scans

Every model in Perch can work with images, including the ones that cannot see. Share a screenshot, a scanned invoice, or a photo and Perch reads it for the model doing your task.

Some of the strongest models available are text-only. They are excellent at reasoning, code, and long documents, and they cannot open an image at all. That is normally a hard wall: pick the better model, lose the ability to hand it a screenshot.

Perch removes the choice. Every model in Perch can work with images, including the ones that cannot see, and it works on every plan including the free one.

The short version

Share a screenshot, a scanned page, or a photo with any model. If that model cannot see, Perch reads the image for it and passes along what is there.

How it works

Image or screenshot
A scan, screenshot, or photo you share.
Vision lane
A vision-capable model reads the image and extracts the text and visual context.
Text-only model does the task
It receives the extracted text, so models like DeepSeek can work with images they cannot see directly.

When you share an image, Perch sends it to a model that can see. That model does one job: describe what is visibly present, including the text, the layout, and anything that looks uncertain. The description comes back as plain text and goes into the context of the model actually doing your work.

So a text-only model never sees the image. It reads a careful description of it, which is enough to answer questions about a screenshot, pull figures off a scanned statement, or tell you what a chart shows.

What this is, and is not

The model doing your task stays in charge from beginning to end. The reading step only describes what it sees. It does not decide what happens next, and it never takes over the work.

Screenshots, photos, and scans

Two different things happen depending on what you share.

Screenshots, photos, and images are read directly, and the description covers both the text and the visual context: which control is selected, what the error says, how a chart trends, what looks off.

Scanned PDFs go through text extraction first. Most PDFs already carry a text layer, and when they do Perch reads it straight through, which is faster and exact. When a PDF turns out to be a scan with no usable text, Perch renders the pages and reads them optically instead. You do not have to tell it which kind you have.

What it costs

Nothing extra. Reading images is part of Perch on every plan, including Starter, and it is not a premium feature or an add-on. The reading step does not count against your included monthly usage, so sharing screenshots and scans costs you nothing beyond the work the model does with them. There is no setting to turn on and nothing to configure.

Limits

Reading an image is a fast step by design, so it stays out of the way of the actual work. That comes with a few edges worth knowing:

  • Images up to 10 MB.
  • Scanned PDFs up to 80 pages in a single read.
  • The reading step is quick rather than exhaustive. It reliably catches the text, the structure, and the obvious visual facts. Very dense tables, small print, and handwriting are where it is most likely to miss something.
  • Because the model doing your task is working from a description, it can only tell you what the description captured. If a detail matters and you do not see it reflected, ask about it directly and Perch will look again.

When a scan comes back empty

Occasionally Perch will tell you it could not pull text out of a document. The usual causes:

  • The file is a photo of a page taken at an angle or in poor light. Retaking it flat and in better light usually fixes it.
  • The scan is very low resolution. Text that is unreadable to you is unreadable here too.
  • The document is longer than the page limit. Split it and share the part you need.

Perch says so when this happens rather than guessing at the contents. An empty read is reported, not filled in.

Documents are data, never instructions

Anything Perch reads out of an image is treated as untrusted material. If a document contains text that looks like an instruction, telling the assistant to ignore its guidelines, email a file somewhere, or change what it is doing, that text is handled as content on a page and nothing more.

This matters more than it sounds. A scanned invoice is a file someone else produced, and it arrives in the middle of a task you asked for. Perch reads it as evidence about the invoice, not as a request from you.

Working in the browser

The same capability is what lets a text-only model drive the agent browser on Perch Terminal Desktop. After each action, the model can look at the current screen and see the result before deciding the next step. See The agent browser.