OCR Free · no signup · no watermark

PDF to Text Converter with OCR

Extract text from PDF files of any kind, including scanned documents. Krawly reads the built-in text layer where it exists and runs OCR on the pages that are only images, all inside your browser.

Updated Krawly Editorial TeamIn-house engineers, writers & reviewers

or drop a file here

PDF · up to 200 MB

Processed in your browser — files are never uploaded

Quick answer

To convert a PDF to text, drop the file into Krawly's PDF to Text tool and choose the document's language (up to 3). Pages that already contain text are extracted instantly; scanned pages are rendered at 300 DPI and read with Tesseract OCR. You get plain text with page separators, ready to copy or download as .txt. The PDF never leaves your device.

How to use the PDF to Text (OCR) tool

  1. 1

    Open your PDF

    Drag a PDF onto the drop area or click to choose one. If the file is password-protected, the tool asks for the password; it is used only in your browser to open the document.

  2. 2

    Select the document language

    English is the default. Choose the language of the scanned text, or up to three languages for documents that mix them. The choice only affects pages that need OCR.

  3. 3

    Decide whether to force OCR

    Leave “Force OCR on every page” off for normal use. Switch it on if a previous extraction gave garbled characters, so every page is rendered and read by OCR instead of trusting the embedded text.

  4. 4

    Extract the text

    Start the conversion. Pages with a text layer finish almost immediately; image-only pages take longer because they are rendered and recognized on your device. The result marks which pages were OCR'd.

  5. 5

    Copy or download .txt

    The text appears with a --- Page N --- separator before each page. Copy it to your clipboard or download it as a .txt file.

Digital vs scanned PDFs: why the text layer is checked first

Not every PDF stores text the same way. A PDF exported from a word processor, a browser or an accounting system usually contains a text layer: the actual characters, their fonts and their positions on the page. A scanned PDF, or one made from phone photos, is different. Each page is just a picture, so there are no characters to copy, and selecting text in a PDF viewer does nothing.

Krawly handles both in one pass. For every page, it first uses PDF.js, Mozilla's PDF library, to look for an embedded text layer. If the page has one, the text is extracted directly. That is exact, because it reads the characters the document already contains, and it is nearly instant. Only pages without a text layer are rendered to an image at 300 DPI and passed to the Tesseract OCR engine in the language you selected.

This matters for mixed documents, which are more common than you might think: a contract with a scanned signature page, a report with photocopied appendices, or a form that was printed, filled in and scanned back. Reading the text layer where it exists avoids introducing OCR errors into pages that were already perfect, and saves time on long documents. The output tells you which pages went through OCR, so you know where to proofread.

When to use Force OCR for garbled PDF text

Sometimes a PDF has a text layer, but the text you get out of it is nonsense: random symbols, letters in the wrong order, missing spaces, or boxes instead of accented characters. This usually comes from how the PDF was produced. Some generators embed fonts without a proper mapping from glyphs back to characters, and some scanning software adds an invisible OCR layer of poor quality on top of the page image.

In those cases, switch on “Force OCR on every page”. The tool then ignores the embedded text, renders every page at 300 DPI and reads what is visibly printed on it. You trade speed for reliability: OCR on every page takes noticeably longer than reading a text layer, especially on a phone, but the result reflects what you actually see on screen.

Force OCR is not needed for ordinary digital PDFs. If the extracted text looks clean, the text layer is the more accurate source, because it contains the exact characters rather than an interpretation of pixels.

Extracting text from multi-language and non-Latin documents

OCR reads scanned pages with a language model, so the language setting is important for any page that needs OCR. Krawly supports English, Spanish, Portuguese, French, German, Italian, Dutch, Turkish, Polish, Russian, Ukrainian, Arabic, Hindi, Japanese, Chinese (Simplified and Traditional), Korean, Vietnamese and Indonesian, and you can combine up to three of them. A German contract with an English annex, or a Ukrainian document with Russian quotations, should have both languages selected.

Pages that already have a text layer are not affected by the language choice, because their characters are read directly. That means a digital PDF in any script, not only the languages in the list, can still be extracted, as long as the PDF stores its text correctly.

Each language model is downloaded the first time you select it, a few megabytes per language, and then cached by your browser. Selecting only the languages that actually appear in the document keeps the first run shorter.

Private PDF text extraction with no upload

PDFs that need text extraction are often the sensitive ones: bank statements, payslips, medical letters, signed contracts, tax forms and scanned passports. With Krawly, the whole process runs in your browser using PDF.js and Tesseract.js compiled to WebAssembly. The PDF is never uploaded to Krawly's servers, and neither is any extracted text.

Your browser only downloads program files, namely the OCR engine and the language models. If the PDF is password-protected, the password is entered and used locally to open the file. Nothing is stored, and closing the tab discards the document and the results.

The trade-off of local processing is that speed depends on your device. Text-layer pages are quick everywhere, while OCR on many scanned pages can take a while on an older phone. PDFs up to 200 MB are supported, and the tool works in current Chrome, Edge, Firefox and Safari on desktop and mobile. It is free, with no signup and no daily limit.

Tips for the best results

  • Try a normal extraction first. Turn on Force OCR only if the text comes out garbled, since OCR on every page is much slower.
  • Proofread the pages marked as OCR'd. Those are the ones that may contain misread characters; text-layer pages are exact.
  • For a long scan where you only need a few pages, use Split PDF to extract those pages first and run PDF to Text on the smaller file.
  • Select all languages used in the scanned pages, up to three, but no more than you need, so fewer models have to be downloaded.
  • Scanned PDFs with crooked or very faint pages give weaker OCR. If you can, rescan at around 300 DPI with the page lying flat.
  • The --- Page N --- separators make it easy to find where each page starts when you paste the text into another document.

Frequently asked questions

Can it OCR PDF scans and convert them to text?

Yes. Pages without a text layer, such as scans or photographed pages, are rendered at 300 DPI and read with the Tesseract OCR engine in the language you choose. Pages that already contain text are extracted directly. The result shows which pages were OCR'd.

How do I convert PDF to TXT?

Drop the PDF here, pick the document's language and start the conversion. The tool extracts the text layer, runs OCR on scanned pages and joins everything with --- Page N --- separators. Click Download .txt to save a plain text file, or Copy to paste the text elsewhere.

Is there a page limit?

There is no fixed page count. PDFs up to 200 MB are supported. In practice the limit is time: text-layer pages are almost instant, while each scanned page needs OCR on your own device, so a long scanned document can take a while, especially on a phone.

Why is the extracted text garbled?

Some PDFs contain a broken text layer, often because fonts were embedded without a proper character mapping or because a poor-quality OCR layer was added when scanning. Switch on Force OCR on every page and run the conversion again; the tool will then read what is visibly printed instead.

Which languages are supported for OCR?

English, Spanish, Portuguese, French, German, Italian, Dutch, Turkish, Polish, Russian, Ukrainian, Arabic, Hindi, Japanese, Chinese (Simplified), Chinese (Traditional), Korean, Vietnamese and Indonesian, with up to three selected at once. Pages with an embedded text layer are extracted directly, whatever their language.

Is my PDF uploaded anywhere?

No. Text extraction and OCR both run in your browser. Only the OCR engine and language model files are downloaded; the PDF, its password if it has one, and the extracted text stay on your device. Nothing is stored, and closing the tab clears it all.

Why is the first OCR run slower?

On first use, your browser downloads the Tesseract engine and the model for each selected language, a few megabytes per language, and caches them. Later runs reuse the cache and start faster.

Does it keep the layout, tables or images?

No. The output is plain text with a --- Page N --- separator for each page. Fonts, columns, tables and images are not preserved, and the tool does not create Word or Excel files. Columns and tables are flattened into lines of text that you may need to tidy up.

Can I extract text from a password-protected PDF?

Yes, if you know the password. The tool asks for it when you open the file and uses it only in your browser to decrypt the document. The password is not sent to Krawly or saved anywhere.

Private by design. Every Krawly converter runs inside your browser using JavaScript and WebAssembly. Your files are not uploaded, stored or seen by anyone — when you close the tab, they are gone.

Related converters

See all 21 free converters