Batch-converting academic papers to Markdown with APA citations
Drafting happens in Word; publishing increasingly happens in Markdown. This guide sets up a batch pipeline that converts a folder of docx papers in one pass, keeps every figure attached, and ends with citations that still format correctly under APA or any other style.
The pieces are a batch converter, pandoc's citeproc engine, and a structured bibliography file. Citations are the one part that needs preparation before conversion rather than after, so the middle of this guide spends most of its time getting that part right.
Word drafts, Markdown destinations
Most scholars still draft in Word: collaborators expect it, journals accept it, and track changes is buried deep in everyone's habits. But the places papers actually end up increasingly speak another language. Lab wikis, institutional repositories, preprint platforms with Markdown support, and version-controlled manuscript projects all want plain text that git can diff and templates can parse.
The workflow that scales is batch: drop a folder of drafts into a converter, let each file become a .md file with its images, then polish. Doing this one paper at a time by hand is how footnote numbers drift and figure paths rot. Automation keeps the pile of drafts and the pile of converted files in lockstep.
The batch workflow
The converter on this site accepts a stack of docx files at once. Each file becomes a job in a queue with its own status: waiting, converting, done, or failed with a reason. Nothing is uploaded, because the pandoc engine runs as WebAssembly inside your browser, so unpublished manuscripts never leave the machine.
When the queue finishes, download everything as one ZIP. Inside, each source document gets its own subdirectory containing the converted Markdown and a media folder with that document's images. The per-document isolation matters for papers, since every Word file names its images image1, image2 from scratch.
Two habits make batch runs painless. First, name the files before converting, because output names derive from the source names. Second, convert a small batch while you are still calibrating options, then unleash the full folder once a spot check looks right.
Why citations are the real problem
Headings, tables, and equations convert well. Citations do not, because of how Word stores them. A bibliography generated by Word's citation tool, or pasted in from a reference manager, is formatted text: it looks right in the document but carries no structured link back to the source record.
Convert that document and the bibliography arrives as plain paragraphs, frozen at whatever style was active the day the document was written. Change the target venue and every in-text citation has to be retouched by hand. The fix is to migrate citations to structured data before converting, not after.
Citeproc in ninety seconds
Pandoc's citation engine is called citeproc. The idea: your references live in a bibliography file, typically BibTeX with a .bib extension, where every entry has a key like smith2020field. In the text you write the at sign followed by that key, and wrap it in square brackets when you want the citation tucked in as a parenthetical.
At conversion time, citeproc replaces every key with a formatted citation and appends the formatted reference list at the end of the output. The formatting itself comes from a CSL file, a small XML style description, with ready-made files available for thousands of journals.
This is the same machinery reference managers like Zotero and JabRef feed when they export, and it is why the one-time effort pays off: one .bib file can drive a Word manuscript today and a Markdown wiki tomorrow without retyping a single entry.
APA output with a CSL file
Without a style file, pandoc falls back to Chicago author-date. For APA, download the APA 7th edition CSL file from the official citation-style-language styles repository, then feed both files to the converter. The Pro citation option on this site takes a bibliography plus an optional CSL style and runs citeproc in the worker, still entirely in the browser.
Check the result against the manual: author-date citations in the right places, a reference list at the end, page numbers where the style wants them. Small deviations usually trace back to incomplete .bib entries, missing DOIs, or an outdated APA style file instead of the current one.
Equations survive the trip
Word stores native equations as OMML, and pandoc reads them, so converted papers carry real math rather than screenshots. Markdown output wraps inline math in dollar-sign delimiters and puts display equations on their own lines; LaTeX output produces the corresponding math environments directly.
Check two things after conversion: inline math that Word rendered ambiguously, such as a stacked fraction squeezed into a line, and equation numbering, which Markdown has no native concept of. Numbered equations headed for a Markdown venue usually need labels handled by the publishing system, not the converter.
What to check in tables and figures
Tables deserve a slow pass. Simple grids convert into pipe tables that GitHub and Obsidian render, but cells with merged regions, nested paragraphs, or multi-line content cannot be expressed in that syntax, and pandoc falls back to Markdown dialects those platforms ignore. Flatten complicated tables in Word before converting.
Figures convert with their captions as image references plus a text line, which is usually fine, but verify each media file actually exists in the ZIP and that nothing arrived in EMF or WMF form, the formats browsers cannot display. Those need a quick conversion to PNG back in Word.
Where the automation ends
Three limits are worth naming. Legacy .doc files are not supported, so save old drafts as .docx first. Word-managed bibliographies stay plain text; migrating them means re-keying entries into a .bib file, or exporting from the reference manager that produced them in the first place.
And if a LaTeX journal wants natbib commands rather than formatted citations, citeproc output will not satisfy it; that calls for pandoc on the command line with the appropriate flag. For everything short of that, batch conversion plus a well-kept .bib file covers the round trip from Word drafts to Markdown with citations intact.