Writing a Pandoc Lua Filter: Adding Custom Marks to Headings Converted from Word
When a Word document becomes Markdown, styles carry over as structure: Heading 1 paragraphs turn into top-level headers, Heading 2 into second-level, and so on. Sometimes that is not enough. You may want every heading from the Word import to carry a marker your site template or build pipeline can recognize, and hand-editing hundreds of headings is not a plan.
Pandoc has a built-in answer: Lua filters, small scripts that transform the document as it flows through the converter. This guide explains how they work, builds one step by step for exactly this tagging job, and covers how to run and debug it.
What a Lua filter is, and why pandoc chose Lua
Pandoc parses every input into a common abstract syntax tree, an internal representation of elements such as headers, paragraphs, and tables, before writing any output. A filter is code that walks that tree and rewrites it. Because the filter operates on the normalized tree, it works the same regardless of whether the input came from docx, HTML, or Markdown.
Pandoc embeds a Lua interpreter, currently Lua 5.4, so filters run inside the pandoc process with no separate runtime and no compiled extension to build. You write a text file, hand it to pandoc with a command-line flag, and the transformation happens mid-conversion. Filters written this way also avoid the JSON round-trip that external filters pay, which keeps the overhead small.
The anatomy of a filter
A filter file defines plain functions named after AST element types. A function named for the Header type runs once per header in the document, one named for the Para type runs per paragraph. If the file does not return an explicit filter table, pandoc collects those top-level functions by name and uses them automatically.
The return value decides what replaces the element in the tree. Returning nil, which includes forgetting to return anything, leaves the element unchanged, so a forgetful function is harmless. Returning the element, modified or not, replaces it with the new version. Returning a list splices those elements in place of the original, and returning an empty list is how you delete an element from the document.
The Header element and its attributes
Every Header in pandoc's AST carries three things: a level as a number, a list of inline content, and an attribute set. The attribute set has three parts, an identifier used for anchors, a list of class names, and arbitrary key-value attributes. This is the same triple that surfaces in output as the id and class on an HTML heading tag.
Where those attributes end up depends on the output format. Pandoc Markdown can render them as curly-brace annotations after the heading text, and HTML output puts them on the tag itself. A plain GFM export has no syntax for attributes, so classes would be dropped there, which is worth knowing before choosing a target format.
The task: tag every Word heading
The goal is a filter that inspects each Header and appends a custom class when it matches the criteria, for example every header, or only second-level ones. The logic reads the level field, decides, and appends a string to the classes list, either with the list's insert method or a plain table insert. Nothing else changes.
Conditionals make it useful in practice. You can give top-level headers one class and deeper ones another, restrict tagging to headers whose text starts with a given word, or skip headers that already carry the class so the filter stays idempotent and can run twice without duplicating marks.
Walking through the code logic
Written out, the filter is four short moves. Define one function named after the Header type so pandoc routes every heading to it. Inside, test the level field against the threshold you care about. If it matches, append the class name to the classes list and, optionally, set a key-value attribute. Return the element at the end.
That final return deserves attention: mutating the element in place and returning it is the documented pattern, but a missing return means pandoc keeps the unmodified original, so the change is computed and then silently discarded. With the return in place, headings come out carrying the new class, visible on HTML tags or in pandoc Markdown brace attributes.
Running the filter
On the command line, pandoc accepts a --lua-filter flag pointing at your file, and the flag can be repeated to chain several filters in order. Filters run between reading and writing, so the transformation applies no matter which output format you choose.
On this site, the Pro options include a Lua filter upload: hand over the same .lua file next to your Word document, and the browser worker passes it to the same pandoc engine. The classes land in the output exactly as they would locally, with nothing sent to a server.
Debugging techniques
The quickest tool is the print function: output printed from a filter goes to stderr, so it never contaminates the converted document, and you can dump a header's level, classes, or text to see what the filter actually received. Pandoc also ships a logging module, pandoc.log, whose info and warning functions tag messages with severity.
A good habit is to develop against a tiny file. Build a docx with one heading at each level, run the filter on it, and inspect the output before pointing the filter at real archives. Most failures are typos in field names, and the error channel reveals them immediately.
Extensions and performance
The same pattern extends far beyond headings: drop paragraphs that contain no text, generate identifiers from heading text for stable anchors, or wrap tables in a Div so a site template can style them as a group. All of these are one function per element type in the same file.
Performance is rarely a concern. Filters run per element inside the conversion process, and the walk is fast even on very large documents. The one thing to avoid is heavy work repeated per element, such as reading a file inside the function, which belongs at the top of the script where it runs once.