en

PDF to Word

Turn a PDF into an editable Word document (DOCX) locally - headings, lists, bold/italic, tables and images included. No upload.

Running locally on your device ...

Your files

    Running locally on your device ...

    0%

    Is my file uploaded?

    No. Everything runs in your browser - your file never leaves your device. How this is verifiable

    No upload100% local
    Your content stays with youno third-party access
    Hosted in GermanyGlobal Content Delivery
    Independently auditedTLS A+ · HTTP headers A+

    A PDF is a finished layout, not editable text: paragraphs are awkward to select, sentences hard to rewrite, a table impossible to keep calculating with. This tool reads your PDF's actual structure - headings, paragraphs, lists, tables, footnotes and images - and builds a real DOCX file in the Office Open XML format locally on your device, ready to open directly in Word, in LibreOffice Writer or in any other program that understands .docx files. No watermark, no account, no daily cap - the conversion happens in your browser tab instead of on a server.

    At paragraph and character level far more survives than just bold and italic. Headings become real Word headings on levels 1 to 6, complete with an outline level, so the navigation pane and the automatic table of contents in Word work straight away. Nested bullet lists keep their levels as a real multi-level Word list. Underlines and strikethroughs are recognised from the lines actually drawn on the page, superscript and subscript from their shifted baseline, and a coloured text highlight is mapped to the nearest of the 16 highlight colours the Word format knows at all. Links from the PDF become clickable Word hyperlinks, and notes at the foot of the page land as real Word footnotes in the footnote pane instead of as loose text at the end of the page.

    Tables are detected three separate ways: from real ruling lines, where the column widths come from the measured line geometry and simple merged cells stay merged; from a clearly recurring gap between the columns when the table has no lines at all; and from a run of at least three consecutive "field: value" lines of the kind form PDFs produce. Page size, portrait or landscape, margins and the number of text columns are taken from the page's real geometry instead of an A4 guess, a clean page break sits at every original page boundary, and images are embedded at their place in the reading flow with their aspect ratio intact - charts drawn purely as vectors are recognised as figures too. Recurring headers and footers become real Word headers and footers with tab stops for left, centre and right; a running page number becomes a live Word field rather than a frozen digit, and a recurring watermark comes along.

    With fonts the tool takes the direct route: the font programs embedded in the source PDF are carried into the Word file unchanged, separately for the regular, bold, italic and bold-italic face. A bold paragraph therefore gets the type designer's real outlines instead of a weight the word processor calculated afterwards. Only where the PDF does not ship a font at all does a metric-compatible, openly licensed substitute step in, which is embedded into the file as well - so the document still looks right on someone else's machine that does not have those fonts installed. Text outside the Latin alphabet additionally gets the matching script and language tagging, so Word shapes Chinese, Japanese, Korean, Arabic or Hebrew correctly, sets it right to left where that belongs, and the spell checker picks the right language.

    A lot of the work sits in the PDFs that fall out of line. A page stored sideways is analysed in its own reading direction, so a table on it does not arrive transposed and a heading does not suddenly read vertically. PDFs printed from a browser, for example via "Print to PDF" in Chrome or Edge, draw bullet points as tiny vector squares rather than as text characters, and glue a whole table row into one single line of text - both are recognised from the geometry and turned back into real lists and individual table cells. A page number only becomes a header or footer when it genuinely keeps counting across pages; a one-off line that merely looks like a page number stays in the text, and changing page identifiers such as "K-6" or "K-86" are preserved rather than thrown away as furniture.

    Two things the tool tells you openly instead of hiding them. If text in the source PDF sat behind an opaque black bar but was still readable as text - the classic mistake of redacting with a painted-on rectangle - that text is not carried into the Word file, and you are told which pages it happened on. And if characters in the source PDF carry no Unicode value at all, which is also what makes copying out of any PDF viewer fail, their share is quantified and you are pointed at "PDF OCR" instead of being left with a silent gap. Analysis and file assembly run entirely in your browser tab; your PDF does not leave your device, and after the first load the tool works offline too.

    Specifications

    Specifications
    Input formatsPDF
    Auto-conversionPNG, JPG, WebP, GIF, BMP, TIFF, ICO, SVG, HEIC, ZIP, DOCX, 7Z, DDS, HDR, PCX, WBMP
    Output formatDOCX
    Batch processingNo
    ProcessingLocally in your browser (JavaScript)
    File uploadNone

    In 3 steps

    1. Drop your PDF - nothing is uploaded.
    2. Structure, tables, fonts and images are detected locally and carried into a DOCX.
    3. Download the DOCX and keep editing in Word or LibreOffice.

    Limitations: Structured rendering, not a pixel-perfect copy of the PDF. Carried over: headings with an outline level, paragraphs, multi-level ordered and unordered lists, bold, italic, underline, strikethrough, superscript and subscript, text highlights (rounded to the Word highlight colours), clickable hyperlinks, real footnotes, images and vector-drawn charts in the reading flow, real headers and footers with a live page-number field, plus page format, orientation, margins and column count from the real page geometry. Tables are detected from ruling lines, from a recurring column gap or from "field: value" lines; only one contiguous table region per page is currently evaluated - several separately ruled tables stacked on one page deliberately flow through as body text rather than being assembled wrongly, and heavily nested or multiply merged tables stay weak too. Every table cell holds one paragraph. Further honest limits: a coloured phrase inside a contiguously set text run loses its own colour, and a highlight covering only part of such a run colours the whole run; a link with no target address stored in the PDF is skipped; a link inside footnote text does not stay clickable; a bullet list nested inside a numbered list takes the numbered marker on its level (content and nesting are right, only the marker glyph is not); if the running header changes per chapter it stays as ordinary text in the document instead of becoming a separate header per chapter. Very elaborately designed pages with text wrapping around figures are approximated, and on pages with several thousand line or vector elements the table and graphic detection is skipped. Scanned or image-only PDFs have no text layer - a note pointing at "PDF OCR" appears instead of an empty result. Open password-protected PDFs with "Unlock PDF" first.

    FAQ

    Is the PDF uploaded?

    No. The PDF is opened in your browser tab, analysed there, and the Word file is assembled there as well; there is no server that ever gets to see the file. After your first visit the tool keeps working even in flight mode - the simplest proof that nothing is uploaded.

    Does the Word file look exactly like the PDF?

    No, and that is deliberate. A PDF is a fixed layout, a Word document is flowing text - a point-for-point copy would no longer be sensibly editable in Word. What is carried over is the meaning: headings, paragraphs, lists, tables, footnotes, images, headers and footers, plus page format and margins. Very elaborately designed pages with text wrapping around figures are approximated.

    Do headings stay real Word headings?

    Yes. Lines set larger and heavier are recognised as headings on levels 1 to 6 and given the matching outline level. That way not only the styles work, but also the navigation pane and the automatically generated table of contents in Word, straight away, with no reformatting on your part.

    Is the original font from the PDF preserved?

    Yes, if the PDF ships it. The font files embedded in the source PDF are carried into the Word file unchanged, separately for the regular, bold, italic and bold-italic face. Only when a font is missing from the PDF does a metric-compatible, openly licensed substitute step in, which is embedded as well - so the document still looks right on a machine where those fonts are not installed.

    Are tables preserved?

    Tables with visible ruling lines become real Word tables, with column widths from the measured line geometry and simple merged cells preserved. Borderless tables with a clearly recurring column gap and form lines of the "field: value" kind are recognised too. Only one contiguous table region per page is currently evaluated: if several separately ruled tables are stacked, they deliberately flow through as body text rather than being assembled wrongly.

    Does this work with rotated or sideways pages?

    Yes. A page stored rotated in the PDF is analysed in its own reading direction rather than the rest of the document's. A table on such a page therefore keeps its rows and columns the right way round, and a heading is not wrongly treated as vertically set text.

    Are bullet lists from a browser-printed PDF recognised?

    Yes. When printing to PDF, Chrome and Edge draw the dots of a bullet list not as text characters but as tiny vector squares - which makes the list markers invisible to a plain text reader. The tool recognises them by their size and by the fact that they line up in one column, and restores the list along with its levels. From the same source, glued-together table rows are split back into individual cells as well.

    Are images, links and footnotes carried over?

    Yes, all three. Embedded images are inserted at their place in the reading flow, width-capped and with their aspect ratio intact; charts drawn purely as vectors are recognised as figures too. Links with a stored target become clickable Word hyperlinks, and notes at the foot of the page become real Word footnotes. A link with no target address is skipped, and a link inside footnote text does not stay clickable.

    What happens to text covered with a black box in the source PDF?

    It is not carried into the Word file. If redaction was done by painting a black rectangle over the text instead of actually removing it, the text remains fully readable inside the PDF - a common and consequential mistake. The tool detects text sitting completely behind an opaque dark area, leaves it out, and tells you which pages it occurred on.

    What happens with a scanned PDF?

    A scan consists of images and has no text layer. The tool says so honestly and points at "PDF OCR" instead of producing an empty Word file. If a file does carry text but its characters have no Unicode value stored, the affected share is quantified - copying out of a PDF viewer fails in that case too, and OCR is the way back to the text.

    Related tools

    DOCX to PDF · PDF to EPUB · PDF to Markdown · PDF OCR (searchable) · PDF to text · PDF to Excel · PDF to PowerPoint