NoteGenNOTEGEN.

Image recognition

Choose between local OCR and online vision models and understand platform differences.

When image recognition is enabled, scan and image records can extract text or generate a description automatically.

Image Recognition

Use system OCR by default and optionally enhance recognition with a VLM.

Image recognition

Enable image recognition
Recognize images in screenshot and illustration captures

System OCR

System OCR

Use OCR provided by the operating system

Available

VLM enhancement (optional)

Vision model
Falls back to system OCR on failure

OCR

OCR works best for clean printed text. It is fast, can run offline, and does not consume model credits.

  • Desktop environments may use Tesseract with downloadable language packages.
  • macOS, iOS, and Android builds can use native recognition; results depend on the OS version.
  • Separate multiple Tesseract languages with commas, for example chi_sim,eng.
  • Adding a language for the first time may download its data file.

VLM

A vision language model works better for complex layouts, tables, handwriting, and semantic image descriptions. Select a vision-capable chat model. The image is sent to that provider.

Choose a method

ScenarioRecommendation
Clean screenshots, scans, bulk text extractionOCR
Tables, posters, handwriting, complex layoutsVLM
Sensitive or offline imagesLocal OCR
Describe the scene rather than only extract textVLM

Recognition results are stored in the record. Verify names, numbers, and specialist terms before organizing them into a note.