Image recognition
Choose between local OCR and online vision models and understand platform differences.
When image recognition is enabled, scan and image records can extract text or generate a description automatically.
OCR
OCR works best for clean printed text. It is fast, can run offline, and does not consume model credits.
- Desktop environments may use Tesseract with downloadable language packages.
- macOS, iOS, and Android builds can use native recognition; results depend on the OS version.
- Separate multiple Tesseract languages with commas, for example
chi_sim,eng. - Adding a language for the first time may download its data file.
VLM
A vision language model works better for complex layouts, tables, handwriting, and semantic image descriptions. Select a vision-capable chat model. The image is sent to that provider.
Choose a method
| Scenario | Recommendation |
|---|---|
| Clean screenshots, scans, bulk text extraction | OCR |
| Tables, posters, handwriting, complex layouts | VLM |
| Sensitive or offline images | Local OCR |
| Describe the scene rather than only extract text | VLM |
Recognition results are stored in the record. Verify names, numbers, and specialist terms before organizing them into a note.