The short answer
A Word file stores a document: paragraphs, headings and lists that flow from one page to the next. A PDF stores a finished page: where each word, picture and line sits, like a printout. Nothing in a PDF says "this line and the next one are one paragraph". That is why a PDF looks the same everywhere, and why it is hard to edit.
To turn a PDF back into a Word file, a converter has to work out the document again from the positions on the page. PDF to Word does that in your browser, and your file is never uploaded. This post explains what it works out well, what it has to guess, and why the result is close to the PDF but never identical.
How a PDF stores a page
Think of a PDF page as a sheet of paper that has been printed. For each piece of text, it records which letters, which font, what size and exactly where on the sheet, down to a tiny fraction of a millimetre. For each picture, it records the picture and its place. For each line or box, its path and colour.
It stores the fonts too, or the parts of them it needs, so that the page looks identical on any phone or computer, even one that doesn't have those fonts. That is the whole point of a PDF: the page looks the same for everyone. The price is that the file doesn't know it contains paragraphs. A line that ends early because the paragraph ended and a line that ends early because the page ran out of room look the same to the file.
It also doesn't know what a table is. A table in a PDF is a set of words placed in rows and columns, plus some lines drawn around them. The lines are drawings, not part of the text.
How a Word file stores a document
A Word file, a .docx, stores the opposite kind of information. It stores the paragraphs as running text, with a style for each, such as a heading or a bullet list. It stores a picture as sitting in a particular place in the text. It doesn't store where each line ends. Word works that out each time, from the page size, the margins and the font you have.
That is what makes it editable. Type a word in the middle of a paragraph and the lines re-wrap, the page breaks move, and the table of contents can update. It is also why the same Word file can run onto a different number of pages on two computers: if one has a slightly different font, the lines break in different places.
What a converter has to work out
PDF to Word reads each page with pdf.js, the PDF reader built into Firefox, looks at the page the way you would, and rebuilds the parts a Word file is made of. Here is what it does with each:
| In your PDF | What the converter works out | In the Word file |
|---|---|---|
| Lines of text | Which lines belong together | Paragraphs that re-wrap as you type |
| Big or bold titles | Which text is a heading | Word headings, which show in the navigation pane |
| Bullets | Which lines are a list | Word bullet lists |
| Two blocks of text side by side | Whether the page has columns | Word columns |
| Rows and columns of words | That it is a table | Lines with tab stops, not a Word table |
| Photos and drawings | Where they sit | Pictures in their places |
Most of this is a good guess. A line that is bold and bigger than the others is very likely a heading. A line that starts with a bullet and sits beside other bulleted lines is very likely a list. The tricky parts are the ones the PDF never stored. A table is the clearest case: the lines and shading of its boxes are drawings, and they are left out, so each row becomes a line of text with a tab stop where each column starts. The columns line up and the text reads row by row, but it isn't a real Word table until you turn it into one.
Why the look is close, but never identical
Fonts are the main reason. A PDF's fonts are not copied into the Word file. Text keeps its font when it is one that comes with Microsoft Office, like Calibri, Arial or Times New Roman. Other fonts become the closest common one: Arial for plain fonts, Times New Roman for fonts with serifs, which are the small strokes at the ends of letters, and Courier New for typewriter fonts. Bold and italic are kept.
Letters in a different font can be a little wider or narrower, so a page of text can run onto an extra page. Our tests show both outcomes:
| Words found in the Word file | Pages in LibreOffice | |
|---|---|---|
| Research paper in two columns, 14 pages | 10,451 of its 10,507 words | 14 |
| Maths textbook, 117 pages | 11,784 of its 11,936 words | 122 |
The paper kept its 14 pages. The textbook grew to 122, because its fonts were swapped and its text spread out. We counted pages by opening each Word file in LibreOffice, and counted words by reading each PDF's words with a different PDF reader, pdftotext, and looking for each one in the Word file. Most of the few words we didn't find were labels inside charts, which stay inside the chart's picture.
Each page of the PDF also starts a new page in Word, with its own margins and spacing. That keeps the Word pages close to the PDF's pages, but it means text you add on one page pushes only that page's text down.
Scans have no text to edit at all
A scanned PDF is a photo of paper. Its pages hold pictures of words, not words. There is nothing for a converter to read, so it can't re-wrap or edit those words. Reading the letters in a picture is called OCR, short for optical character recognition, and PDF to Word doesn't do OCR. Each scanned page goes into the Word file as one picture, sharp enough to read and print, and the result says which pages those are.
You can tell a typed PDF from a scan with a simple test: try to select one word with your finger or the mouse. If you can, the file is typed. If the whole page is selected like a photo, it is a scan.
Some scanner apps and copiers run OCR themselves and hide the recognised text behind the picture. When a PDF has such a hidden text layer, the Word file gets that text instead of the picture. Check names, numbers and amounts in it, because OCR makes mistakes, like a 5 read as an S.
Which PDFs convert well, and which don't
- Easy: a PDF saved from Word or Google Docs, with ordinary paragraphs, headings, lists and a few pictures. It has real text, so it converts best.
- Fair: a form, a brochure or a report with tables and coloured boxes. The words come across, and the boxes and shading are left out.
- Harder: maths. Formulas come across as letters and symbols, but fraction bars, root signs and tall brackets are drawn lines and are left out, so formulas need tidying by hand.
- Hardest: scans, and PDFs in Hindi or other Indian scripts whose joined letters can't be read back. The result warns you, and you check the words by hand.
The post Convert PDF to Word without losing formatting shows what to check and how to fix each thing, and Convert PDF to editable Word on phone covers doing it on a phone.
When you don't need to convert at all
Converting is not always the best way. Sometimes a smaller job is enough:
- Ask for the original. If someone sent you a PDF of a document they wrote in Word, they probably still have the Word file. It is the cleanest source.
- Change the pages, not the words. Swapping a page, deleting some or moving them around doesn't need Word. See How to replace a page in a PDF and How to delete pages from a PDF.
- Add, don't edit. Page numbers, a watermark or a password can be added to a PDF without changing its words. See Add Page Numbers, Add Watermark and Protect PDF.
- Keep the exact look. For a signed or official document, the PDF is the original, with the exact look, fonts and layout. Convert a copy and keep the PDF beside it.
NoUploadPDF has no tool that edits the words inside a PDF. Converting to Word, editing there and saving a new PDF with Word to PDF is the way to get an edited PDF here. Expect the layout to need a check, as the sections above explain.
Your document stays on your device
Most online PDF to Word converters work by uploading: your PDF goes to their server, is converted there, and you download the Word file back. NoUploadPDF works the other way round. Your own browser reads the PDF and writes the Word file. The tool has no upload step and never sends your file anywhere. The page contacts only this site and Google's ad service, which shows the ads that keep the tools free, and our code never gives your files to the ad code. That matters for a CV, a bank statement, a contract or a medical report. The privacy policy has the details.
Try it: convert a typed PDF and see which parts come across as editable text.
Open PDF to WordPDF to Word, free and private: turn a PDF into a Word file you can edit, with its text, lists and pictures.
Open PDF to Word