What "extracting from a PDF" means and when you need it

Extracting from a PDF means pulling out content — text, images, tables, or embedded files — and saving it in a format you can edit or use separately. A PDF is designed to look the same on any device, which makes it hard to change. When you extract, you're breaking that lock to get the raw material underneath.

You might need to extract text from a scanned invoice to paste into a spreadsheet, pull an image from a PDF to use in a presentation, or recover a Word document that was saved as a PDF. The method depends on what you're extracting and what you want to do with it afterward.

Key Takeaways

  • Text extraction works best on PDFs made from digital documents; scanned PDFs (images of pages) need optical character recognition (OCR) software to read the text.
  • Built-in tools in Adobe Reader and Preview (Mac) can copy text and save images without installing anything new.
  • Free tools like ILovePDF and Smallpdf work in your web browser and can extract text, images, or embedded files in seconds.
  • If a PDF is password-protected, you need the password to extract anything; if it's locked for printing only, most extraction still works.

Extracting text from a digital PDF using built-in tools

If the PDF was created from a Word document, email, or other digital source (not a scan), the text is already there — you just need to copy it. Open the PDF in Adobe Reader (free, from adobe.com) or Preview on Mac, then click and drag to select the text you want. Right-click and choose Copy, or use Ctrl+C (Windows) or Command+C (Mac). Paste it into your document with Ctrl+V or Command+V.

This works for most PDFs you encounter. The text comes out in the order it appears on the page, though formatting like bold or italics may not carry over. If you need to extract all the text at once, Adobe Reader has a "Select All" option under the Edit menu, then copy the whole thing.

Preview on Mac has a similar feature: open the PDF, go to Tools, then select the text tool, and drag to highlight what you need. Copy and paste works the same way.

Extracting images from a PDF

To pull an image out of a PDF, right-click directly on the image in Adobe Reader or Preview and look for "Save Image" or "Export Image." The image saves as a separate file (usually JPG or PNG) that you can open in any photo editor or insert into another document.

If right-clicking doesn't work, or if you need to extract many images at once, use a free online tool like ILovePDF (ilovepdf.com). Upload your PDF, choose "Extract Images," and the tool pulls out every image and lets you download them as a ZIP file. No account needed, and the file is deleted from their server after a few hours.

Smallpdf (smallpdf.com) and PDF.io offer the same feature. These tools work in any web browser, so you don't install anything on your computer.

Extracting text from a scanned PDF (image-based)

If your PDF is a scan — a photograph of printed pages — the text isn't actually text yet; it's part of an image. Copying won't work. You need optical character recognition (OCR), software that reads the image and converts it to editable text.

Adobe Reader has a built-in OCR feature. Open the scanned PDF, go to Tools, then "Recognize Text," and choose "In This File." Adobe processes the page and makes the text selectable. You can then copy and paste as normal. This works well for clean, clearly printed pages.

Free alternatives include ILovePDF's OCR tool and Smallpdf's OCR feature — both are free for a few uses per month. Upload your scanned PDF, choose OCR, and download the result as a new PDF with searchable text, or as a Word document. Google Drive also has a hidden OCR feature: upload a PDF to Drive, right-click it, choose "Open With," then "Google Docs." Google converts it and you can copy the text out.

Extracting embedded files from a PDF

Some PDFs contain hidden files — spreadsheets, documents, or data files attached inside. To see if yours does, open it in Adobe Reader, go to File, then look for "Properties" or "Attachments." If there are attachments listed, right-click each one and choose "Save As" to extract it to your computer.

If Adobe Reader doesn't show an attachments panel, the PDF may not have any embedded files, or they may be in a format Reader doesn't display. Online tools like ILovePDF can sometimes extract these too — upload the PDF and look for an "Extract Files" or "Attachments" option.

What to do if the PDF is password-protected

If a PDF asks for a password when you open it, you need that password to extract anything. There is no legitimate way around this — password protection is intentional, and removing it without permission is illegal in most places.

If you've forgotten your own password, Adobe Reader can't help you recover it. You would need to contact whoever created the PDF or ask them to send an unprotected version. If the PDF is locked for printing only (not for opening), extraction usually still works — you can copy text and save images even if you can't print.

Protecting your privacy when using online extraction tools

Free online tools like ILovePDF and Smallpdf are convenient, but they do see your file. Both claim to delete files after a few hours, and both use HTTPS encryption (the padlock icon in your browser address bar) so the file is encrypted in transit. However, if your PDF contains sensitive information — medical records, financial data, or personal details — consider whether you're comfortable uploading it to a third-party server.

For sensitive documents, use Adobe Reader's built-in tools instead, which work entirely on your computer. If you need OCR for a sensitive scanned document, Google Drive is an option (Google has privacy policies covering what it does with files), or install free desktop software like Tesseract (an OCR engine) if you're comfortable with technical setup.

Frequently Asked Questions

Can I extract text from a PDF that's locked for editing?

Yes. PDFs locked for editing usually allow copying text and saving images. Open it in Adobe Reader or Preview, select and copy the text as normal. The lock prevents you from changing the PDF itself, not from extracting what's in it.

Why does extracted text look jumbled or out of order?

PDFs store text in the order it was created, not always the order it appears visually. Multi-column layouts, sidebars, and complex designs often extract in a confusing sequence. Paste into a text editor first to see the actual order, then rearrange if needed. OCR on scanned PDFs can also produce errors if the scan is blurry or the font is unusual.

Is it legal to extract text or images from a PDF?

Extraction itself is legal. What you do with the extracted content depends on copyright. If you own the PDF or created it, you can extract freely. If it's someone else's work, extracting for personal use is usually acceptable, but republishing or selling the content is not. When in doubt, ask the creator for permission.

What's the difference between extracting and converting a PDF?

Extracting pulls out specific pieces (text, images, files). Converting changes the entire PDF into a different format — like turning it into a Word document or Excel spreadsheet. Conversion tools try to preserve layout and formatting; extraction just grabs the raw content.

Do I need to install software to extract from a PDF?

No. Adobe Reader is free and handles most extraction without extra software. Online tools like ILovePDF work in your browser. The only time you'd install something new is if you want advanced features like batch processing or desktop OCR for many scanned documents.