Why photos are harder than they look
A camera snapshot brings skew, uneven lighting, and noise — and classic OCR engines respond by returning a jumble of characters. Modern AI OCR takes a different route: the Image & Scan to Word card runs IBM's Docling Granite vision model over your photo, finds the actual text blocks, and rebuilds them as a document instead of loose lines.
From snapshot to editable Word in one click
Upload a PNG, JPG, or even a TIFF straight from your scanner. The engine auto-handles rotation and noise, detects headings and paragraphs, and exports an editable Word (.docx) by default — or a searchable PDF, Markdown, or plain text if you prefer. For bound books, legal letters, and multi-page paper reports, the dedicated Scanned Document to Structured Doc card preserves the heading hierarchy so chapters stay chapters.
Both tools live inside the Docling Granite OCR suite, so you can switch export formats without re-uploading. Nine recognition languages are supported, from English and Spanish to Chinese and Japanese.
Getting better results from tricky shots
Three habits dramatically improve accuracy: shoot in even light (shadows and glare are the classic accuracy killers), hold the camera parallel to the page, and fill the frame with the document. The engine auto-corrects mild rotation and skew, but crumpled paper and cut-off margins are easier to avoid than to repair. For thick stacks — bound books, legal files, semester notes — let the Scanned Document to Structured Doc card handle the batch so chapter headings stay hierarchical instead of flattening into one long page. Non-English material is fully supported; pick the language before converting and the recognition model adjusts.
If a photo is too blurry or too dark to salvage, don't retry ten times — retake it; sharpness beats any amount of AI cleverness. And remember the output is fully editable: a quick spell-check pass in Word fixes the last one percent, which is exactly why an editable .docx beats a flat image of text every single time.