User PDF

How to review OCR in a Portuguese PDF

OCR recognizes characters in images and can make a PDF searchable. Recognition can be wrong even when the file opens normally. Review important passages before copying the text into another system.

Steps and checks

  1. In Recognize text, choose Portuguese as the document language. For genuinely bilingual pages, Portuguese and English is also available.
  2. Choose Searchable PDF to search the document, TXT text for plain text, or PDF and TXT to receive both.
  3. Consider Correct skew and Preserve pages that already have text for your file. These options cannot restore details the scanner did not capture.
  4. Search for a phrase in the PDF and compare the extracted text against the original image. Pay special attention to passages you will reuse.

Accents, similar letters, and digits

Check ç, ã, õ, and accented words. Watch for O versus 0, I versus 1, and decimal punctuation. A plausible word may still be wrong. Names, dates, codes, and numbers require direct comparison with the original.

Searchability does not mean table reconstruction

A searchable PDF adds text for search, but does not turn tables into Excel cells. To work with columns and records, use PDF to Excel and review its structure. To rewrite paragraphs, consider PDF to Word in Flowing text mode.

Columns and reading order

Newspapers, forms, and two-column pages can produce text in an unexpected sequence. Check a passage across a column transition. A TXT file contains recognized plain text and does not preserve the full page layout.

Improve the input before retrying

Prefer sharp scans without cropping, shadows, or strong perspective. Enlarging a small image does not recover missing characters. If the original is unreadable, rescanning may help more than repeating the task with the same settings.