Extract text from a PDF without uploading it
Pull the text out of a PDF into a plain .txt file, in a couple of clicks, without uploading the document anywhere. It reads the PDF’s real text layer — not a picture of it — so you get clean, copyable text.
No sign-up. No upload. Nothing to install.
How to extract text from a PDF
- Open your PDF. Drop the document onto PDFLight. It opens instantly, without being uploaded.
- Open the Convert menu. It sits in the top bar and is always available.
- Choose "To text". PDFLight reads the document’s text layer and assembles it, preserving line breaks.
- Download the .txt. The text file is saved to your disk. Everything was extracted on your machine.
A text layer, not OCR
PDFLight extracts the text that is actually stored in the PDF — the same text you can select with your cursor. That is fast and exact. It does not run OCR, so a PDF that is really just a scanned image (a photo of a page, with no text layer) will return little or nothing; for those you would need an OCR tool first.
The distinction is worth internalising, because it decides whether this tool will work at all for your file. A PDF exported from a word processor stores real text objects, and reading them back is exact — no recognition, no confidence scores, no errors. A scan stores a picture, and no amount of extraction will find words in it, because there are none.
To tell which you have before you start: try selecting a line of text in the document. If nothing highlights, extraction will return nothing and you need OCR instead.
The text of a contract, a report or an invoice is often the most sensitive part of it — it is the part that gets quoted, pasted into an email or fed into another system. Extracting it in your browser means it is never handed to a third-party server, which is exactly the situation where a server-side "PDF to text" converter would receive the complete contents of the document in readable form.
Frequently asked questions
Does it use OCR?
No. It reads the PDF’s embedded text layer, which is exact and instant. Scanned-image PDFs with no text layer will not yield text — run OCR on those first.
Is my document uploaded?
No. The text is extracted in your browser. Nothing is sent to a server.
What file do I get?
A plain .txt file with the document’s text and its line breaks preserved.
Why did I get an empty file?
The PDF is probably a scanned image with no real text layer. PDFLight cannot read text that is only a picture.
Can I also export the pages as images?
Yes. Convert → "To images" renders each page to a PNG and returns them as a ZIP.
Go deeper
If extraction comes back empty, or comes back garbled, the reason is in the format itself. These explain both cases: