Migrating from Word to Markdown: the complete guide
Migrating documents from Word to Markdown is mostly mechanical, but the failures are predictable and cheap to prevent. Every one of them traces back to treating the conversion as a single magic step instead of a small pipeline with a few decisions in it.
This guide walks through those decisions in order: how much you are actually migrating, what survives conversion and what does not, how footnotes, tables, and math travel, and how to finish with a file your static site or wiki accepts on the first try.
Decide the scope before converting
A single live document and a ten-year archive are different projects. For one document, the goal is a faithful conversion you can polish by hand in an afternoon. For an archive, the goal shifts to repeatability: one set of options applied consistently, with humans reviewing a sample rather than every file.
Either way, start with a sample. Convert two or three representative documents, the ones with the worst tables and the most figures, and inspect the output closely. The sample tells you which options matter and how much cleanup the rest of the pile will need, before you commit to the full run.
What converts cleanly, and what does not
Real Word heading styles become Markdown headings, simple tables become tables, footnotes become footnotes, and embedded images extract to real files. Content that lives in the normal document flow transfers well, and that is most of what a typical document actually is.
The exceptions cluster at the edges. Formatting applied directly, such as bold lines pretending to be headings, does not become structure, so apply real heading styles in Word first. Tracked changes convert as if every change were accepted, so resolve revisions before converting if the deleted text matters to you. Comments, text boxes, and floating shapes do not carry over reliably, and content inside them should be moved into the main flow or captured separately.
Footnotes come along
Word footnotes are one of the pleasant surprises. Pandoc reads them and Markdown output uses real footnote syntax: each marker in the text becomes a bracketed reference with a caret before its number, and all the definitions collect at the end of the document, linked back to their markers.
Two things still deserve a check. Footnotes that reference each other, or that contain tables or images, can come out in a different order or shape than the original. And endnotes convert the same way as footnotes, which is usually what you want in Markdown anyway.
Tables and their limits
Enable the GFM option when the destination is GitHub, Obsidian, or most wikis, and simple tables convert into pipe tables: one header row, a separator line, and one row per line. That format renders everywhere and stays readable as plain text.
The limit is structural. Merged cells, cells holding several paragraphs, and nested tables have no pipe-table representation, and pandoc falls back to Markdown dialects that those platforms ignore. The honest fix is upstream: unmerge and flatten complicated tables in Word, or accept that a few tables will need manual rework after migration.
Mathematics becomes LaTeX
Equations typed with Word's equation editor are stored as structured math, and pandoc converts them rather than discarding them. Markdown output wraps inline math in dollar-sign delimiters and puts display equations on their own lines; the LaTeX output format emits native math environments for rendering pipelines that expect them.
After conversion, scan for the two weak spots: inline math that Word displayed ambiguously, and anything pasted into Word as a picture years ago, which is an image now, not math. Those need re-entry, and knowing that before the migration starts keeps the schedule honest.
One predictable image strategy
Every embedded image extracts to a media folder next to the converted file, with names the docx package assigned: image1.png and its siblings. The Markdown references those relative paths, and a ZIP download delivers both together, which keeps every converted document a self-contained pair.
Decide before the run whether documents keep their own media folders or merge into a shared assets folder. Keeping them separate is safe and automatic; merging requires renaming with meaningful prefixes and updating the references to match. What you want to avoid is discovering the collision after several image1.png files have already overwritten each other.
Front matter for your static site
Static site generators such as Hugo and Jekyll expect YAML front matter at the top of each file: title, author, date. The converter can emit that block from the document properties stored in the docx, so files arrive already carrying metadata instead of needing it typed in afterward.
That only works if the properties are real. Word files converted wholesale from older formats often have an empty Title field and a creator listed as a former employee's account. Spending ten minutes in Word before the run, filling in File Info for each document, is the difference between useful front matter and placeholders.
The batch migration pass
Once options are calibrated on the sample, the bulk run is uneventful. Add the folder of docx files to the converter on this site, let the queue process each one with its own status, and download the ZIP. Batch output gives every document its own subdirectory, so images from one paper can never shadow the images of another.
Keep the original docx files until the migration is signed off. They are the source of truth for anything the conversion smoothed over or dropped, and re-running with adjusted options is always cheaper than reconstructing a lost detail from memory.
A checklist for the finished migration
For each document, verify six things: the heading hierarchy descends without jumps, every image reference resolves to a file that exists, tables render on the target platform, footnotes appear as real footnotes, math renders where it should, and front matter fields are populated rather than empty.
Finally, run one document end to end in the real destination: build the static site, preview the wiki page, open the note in Obsidian. A migration is done when the target system renders the files, not when the converter reports success. That last check takes minutes and catches everything the file-level checks missed.