Anydoc: Word, PowerPoint, Excel, OpenDocument, RTF, ePub, CSV, PDF to Markdown Firecrawl released anydoc, a Rust library that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into GitHub-Flavored Markdown in single-digit milliseconds, with bindings for Node.js, Python, and WebAssembly. The library powers Firecrawl Parse and is available as an Agent Skill for use with Claude Code, Codex, Cursor, and OpenCode. Fast Rust library that converts documents Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF into clean GitHub-Flavored Markdown. Includes bindings for Node.js /firecrawl/anydoc/blob/main/node/README.md , Python /firecrawl/anydoc/blob/main/python/README.md , and the browser /firecrawl/anydoc/blob/main/wasm/README.md WebAssembly . Built by Firecrawl https://firecrawl.dev to turn any office document into LLM-ready Markdown in single-digit milliseconds, with one consistent output no matter which format goes in. It powers Firecrawl Parse https://firecrawl.dev/parse , so if you'd rather not run it yourself, the hosted API gives you the same conversion plus our OCR models for the scanned pages anydoc can't read on its own. Try it in your browser : the demo page runs the library as WebAssembly, so files are converted locally and never leave your machine. anydoc ships as an Agent Skill https://agentskills.io , so your agent can read any document it runs into: npx skills add firecrawl/anydoc The skill /firecrawl/anydoc/blob/main/skills/convert-documents-to-markdown/SKILL.md teaches the agent to convert documents with the anydoc CLI. Works with Claude Code https://claude.ai/code , Codex https://openai.com/codex/ , Cursor https://cursor.com , OpenCode https://opencode.ai , and any other compatible agent https://agentskills.io/clients . npx @firecrawl/anydoc report.docx Markdown to stdout npx @firecrawl/anydoc slides.pptx -o slides.md or to a file npx @firecrawl/anydoc - --format csv < data.csv read stdin npx downloads the prebuilt binary for your platform on first run. For a permanent anydoc command, install globally with npm install -g @firecrawl/anydoc . Run anydoc --help for all options. npm install @firecrawl/anydoc js import { toDocument, toMarkdown, toMarkdownBytes } from '@firecrawl/anydoc'; // From a file path: const markdown = await toMarkdown 'report.docx' ; // From bytes, with the format detected from the content: const fromBytes = await toMarkdownBytes bytes ; // Or name it, which signature-less formats CSV need: const fromCsv = await toMarkdownBytes bytes, 'csv' ; // Or stop at the document model, which also carries embedded assets: const document = await toDocument bytes ; Full API reference: node/README.md pip install firecrawl-anydoc python import anydoc From a file path: markdown = anydoc.to markdown "report.docx" From bytes, with the format detected from the content: markdown = anydoc.to markdown bytes data Or name it, which signature-less formats CSV need: markdown = anydoc.to markdown bytes data, "csv" Or stop at the document model, which also carries embedded assets: document = anydoc.to document data Full API reference: python/README.md npm install @firecrawl/anydoc-wasm python import init, { toMarkdownBytes, toDocument } from '@firecrawl/anydoc-wasm'; await init ; // From bytes, with the format detected from the content: const markdown = toMarkdownBytes bytes ; // Or name it, which signature-less formats CSV need: const fromCsv = toMarkdownBytes bytes, 'csv' ; // Or stop at the document model, which also carries embedded assets: const document = toDocument bytes ; Full API reference: wasm/README.md cargo add anydoc js // From a file path: let markdown = anydoc::to markdown "report.docx" ?; // From bytes, with the format detected from the content: let markdown = anydoc::to markdown bytes &bytes, None ?; // Or name it, which signature-less formats CSV need: let markdown = anydoc::to markdown bytes &bytes, anydoc::Format::Csv ?; // Or stop at the document model, which also carries embedded assets: let document = anydoc::to document &bytes, None ?; One output for every format. Each format parses into a shared document model and renders through a single Markdown serializer, so escaping, tables, heading anchors, and footnotes behave identically whether the input was a .doc from 2003 or a .pptx from yesterday. Full document structure. Headings with anchors, bold/italic/strikethrough, inline code and code blocks, links and internal cross-references, bulleted/numbered/nested/task lists with the source's own numbering, tables with merged cells and header rows, block quotes, footnotes and endnotes, and speaker notes. Embedded assets. Images and embedded objects render as their alt text in the Markdown, and the raw bytes stay available on the document model, tagged with their media type. Images with an external URL become ordinary Markdown images. Content-based format detection. The format is read from the bytes themselves PDF header, RTF open group, OLE stream names, ZIP package mimetype , so mislabeled files still convert correctly. Fast. Pure Rust, no ML models, no external services. Median conversion time is under 5ms per document. Bindings that stay out of the way. Node.js conversion runs on the libuv thread pool and never blocks the event loop; Python releases the GIL so other threads keep running. TypeScript types and Python stubs ship with the packages. PDF support built in. Text-based PDFs convert locally through pdf-inspector https://github.com/firecrawl/pdf-inspector , no OCR service required. Agent ready. Ships as an Agent Skill agent-skill : one npx skills add firecrawl/anydoc and any agent can read office documents. | Format | Extensions | |---|---| | Word | .doc , .docx , .docm | | PowerPoint | .ppt , .pps , .pot , .pptx , .pptm , .ppsx , .ppsm | | Excel | .xls , .xlsx , .xlsm , .xlsb | | OpenDocument | .odt , .ods , .odp | | Rich Text Format | .rtf | | EPUB | .epub | | CSV | .csv | .pdf | anydoc is measured against six other converters on 100 real-world documents spanning fourteen formats. Scores run from 0 to 100, higher is better; speed is the median time to convert one document. | tool | formats | median ms | docs judged | score | completeness | structure | formatting | cleanliness | |---|---|---|---|---|---|---|---|---| | anydoc | 14/14 | 4.4 | 94 | 81 | 87 | 79 | 78 | 81 | | libreoffice | 12/14 | 1129.5 | 87 | 40 | 59 | 42 | 40 | 24 | | unstructured | 8/14 | 572.9 | 58 | 63 | 76 | 59 | 51 | 63 | | markitdown | 6/14 | 134.8 | 33 | 65 | 78 | 66 | 60 | 52 | | pandoc | 5/14 | 102.1 | 34 | 56 | 74 | 57 | 56 | 38 | | docling | 4/14 | 513.6 | 21 | 57 | 60 | 60 | 57 | 51 | | mammoth | 1/14 | 52.5 | 8 | 70 | 84 | 71 | 75 | 51 | Per format, like for like: | format | anydoc | libreoffice | unstructured | markitdown | pandoc | docling | mammoth | |---|---|---|---|---|---|---|---| | doc | 87 | 57 | 67 | - | - | - | - | | docm | 84 | 48 | - | - | - | - | - | | docx | 88 | 56 | 53 | 71 | 68 | 71 | 70 | | epub | 77 | - | 72 | 72 | 52 | - | - | | odp | 86 | 23 | - | - | - | - | - | | ods | 82 | 38 | - | - | - | - | - | | odt | 80 | 51 | 68 | - | 60 | - | - | | ppt | 80 | 26 | - | - | - | - | - | | pptx | 74 | 24 | - | 66 | - | 52 | - | | rtf | 88 | 53 | 46 | - | 45 | - | - | | xls | 80 | 38 | 66 | 62 | - | - | - | | xlsm | 76 | 32 | - | - | - | - | - | | xlsx | 72 | 30 | 66 | 55 | - | 47 | - | How quality was scored: an LLM judge Claude Sonnet 5 compares two tools' outputs blind against ground truth: the document's first six pages, rendered to images by LibreOffice. Each output is scored on completeness, structure, formatting, and cleanliness. Every pair is judged twice with the outputs swapped to cancel position bias, for 482 verdicts in total. Each tool's score averages its per-format scores over the formats it supports, so a corpus heavy in one format can't skew it. It also means each row averages a different set of formats mammoth's 69 is docx alone, while anydoc's 81 spans all fourteen , so the per-format table is the fair comparison. Speed is one warm conversion per document on a Ryzen 9 9950X3D Windows 11, 64 GB DDR5-6400 . anydoc and the Python libraries are timed with process spawn excluded; the CLI tools include it, since that is how they are used. The harness lives in bench/ /firecrawl/anydoc/blob/main/bench/README.md ; the corpus is not redistributable and is not in the repo. Best fit: pipelines that receive a mixed bag of office documents and need one consistent, structured Markdown output. In this comparison, anydoc was the only tool to cover all fourteen formats, scored highest on every judged format, and converted documents an order of magnitude faster than the next-fastest tool. The format is read from the file content, using the marker its specification designates: the PDF header, the RTF open group, OLE stream names, the ZIP package mimetype and content types. CSV has no such marker, so the extension or an explicit format names it instead. Format::from bytes &bytes ; // Some Format::Docx , or None when nothing matches Format::from extension "pptm" ; // Some Format::Pptx Format::from path Path::new "report.odt" ; // Some Format::Odt The same three functions exist in Node formatFromBytes , ... and Python anydoc.format from bytes , ... . A conversion returns Err only when no meaningful Markdown could come out of the file. ConvertError names what went wrong: js match anydoc::to markdown path { Ok markdown = Some markdown , // No document comes out of these, so record the file and take the next one. Err error @ ConvertError::Encrypted | ConvertError::Unsupported = { unconverted.push path, error ; None } Err error = return Err error , } | Variant | Meaning | |---|---| Unsupported | Unknown format, or one that cannot be converted an image-only PDF | Malformed | Structurally unusable: no meaningful content could be extracted | Encrypted | Encrypted or password-protected | ResourceLimit | Crossed a fixed safety limit decompression, nesting, node count | MissingPart | A part required for any meaningful output is absent | Io | The file could not be read, from to markdown only | Node and wasm publish the variant name on error.code ; Python raises one anydoc.ConvertError subclass per variant, or OSError when the file cannot be read. document bytes │ ├─► format detection → content markers, not the extension │ ├─► format parser → one per format doc, docx, ppt, pptx, xls, │ xlsx, odt/ods/odp, rtf, epub, csv │ │ │ └─► Document → shared model: blocks, inlines, tables, │ footnotes, assets │ │ │ └─► GFM serializer → Markdown │ └─► PDF → pdf-inspector → Markdown directly Because every format funnels through the same document model and serializer, output quirks get fixed once. A table-escaping fix for docx is automatically a table-escaping fix for rtf, odt, and everything else. cargo test cd node && npm install && npm run build && npm test cd python && pip install maturin && maturin develop && python -m unittest discover -s tests wasm-pack build wasm --release --target web --scope firecrawl && node --test wasm/test.mjs see wasm/README.md A committed fixture corpus under tests/fixtures/ is snapshot-tested, tests/robustness.rs mutation-tests every fixture, and fuzz/ carries cargo-fuzz targets per format. The speed and quality benchmark lives in bench/ /firecrawl/anydoc/blob/main/bench/README.md . Releases are tagged v