anydoc

A Firecrawl skill and CLI that lets your agent convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF to Markdown on your own machine

SkillEditors’ pick

Medium risk

We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

Why this level

  • The agent runs the @firecrawl/anydoc package through npx, which downloads a prebuilt binary on first run
  • The CLI reads whatever files the agent points it at and writes Markdown to the chosen file
  • With --ocr hosted the whole document goes to the Firecrawl Parse cloud
All reasons and checks

firecrawl/anydoc

Install

Manual install

npx skills add firecrawl/anydoc

Installs the convert-documents-to-markdown skill through the skills CLI into Claude Code, Codex, Cursor, OpenCode and other agents that support Agent Skills. The skill runs the converter through npx and needs Node.js 20 or newer.

Security check

How we review

Medium riskWe rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.

  • The agent runs the @firecrawl/anydoc package through npx, which downloads a prebuilt binary on first run
  • The CLI reads whatever files the agent points it at and writes Markdown to the chosen file
  • With --ocr hosted the whole document goes to the Firecrawl Parse cloud

This is third-party code. Review the repository files before installing.

What it does

anydoc turns office documents, ebooks and PDFs into Markdown that an agent reads as plain text. The convert-documents-to-markdown skill teaches Claude Code, Codex, Cursor and OpenCode to call the CLI through npx, which prints the result to stdout or writes it to a file with -o. The converter is written in Rust and runs without ML models or external services. It keeps headings, lists with their original numbering, tables with merged cells, footnotes and speaker notes, and turns equations into LaTeX. The CLI detects the format from the file content, so a document with the wrong extension still converts.

Who it is for. For developers, analysts and technical writers who hand their agent contracts, spreadsheets and slide decks in office formats, including legacy .doc, .ppt and .xls.

Good fit when

  • Your agent needs to read a .docx contract, a slide deck or an Excel sheet it cannot open on its own
  • A folder mixes legacy .doc, .ppt, .xls and OpenDocument files, and you need the same kind of Markdown for all of them
  • Documents must not go to an external service: office files and text-based PDFs convert on your machine
  • You need to build conversion into a Node.js, Python or Rust project through a library with the same formats as the CLI

Not a fit when

  • You need to turn an HTML page, an image, audio or a YouTube video into text: anydoc does not accept those, MarkItDown reads them
  • Most PDFs in your archive are scans and must not go to the cloud: anydoc has no OCR of its own
  • Your agent needs an MCP server with a conversion tool: the repository ships a skill, a CLI and libraries, with no MCP server

Example request

Read contract.docx and prices.xlsx in the docs folder and list the payment terms and the price for each item

Limitations

The skill and the CLI need Node.js 20 or newer. On first run npx downloads a prebuilt binary for your platform: macOS, Linux or Windows x64. The Python package needs Python 3.10 or newer. There is no built-in OCR: a scanned PDF fails with NeedsOcr, and with --ocr hosted the whole document goes to the Firecrawl Parse cloud. Parse works without signup at limited rates, and a FIRECRAWL_API_KEY raises the limits. Encrypted and password-protected files do not convert, and images inside a document reach the Markdown only as their alt text.

How to disable. Run npx skills remove convert-documents-to-markdown or delete the convert-documents-to-markdown folder from your agent's skills directory, for example ~/.claude/skills. Remove the global CLI with npm uninstall -g @firecrawl/anydoc and the Python package with pip uninstall firecrawl-anydoc.

FAQ

Do documents get sent anywhere?

Not by default. Office files and text-based PDFs convert on your machine, and the Rust crate never makes network calls. Only scanned PDFs go to Firecrawl Parse, and only with --ocr hosted.

How does the agent handle a large document?

SKILL.md tells it to write the result to a file with -o and read the parts it needs. A long spreadsheet or book never lands in the context in one piece.

How does the agent know a conversion failed?

From the exit code: 0 is success, 1 means the document could not be converted, 2 is a usage error and 3 means the PDF has scanned pages. The CLI prints the reason as one line to stderr and never prompts.

Skills
Official

Microsoft's utility and MCP server that turn PDF, Word, Excel, PowerPoint, HTML and other files into Markdown for models

MCP serverMedium risk187.8KRepository stars
Editors’ pick

Anthropic's official skill examples: docx, pdf, pptx and xlsx documents, design, MCP builder and the Agent Skills spec

SkillMedium risk179.2KRepository stars
Editors’ pick

A skill that builds editable PPTX decks from PDF, DOCX, URLs and Markdown with native shapes, charts and animations

SkillMedium risk57.1KRepository stars
Editors’ pick

A skill that draws editorial diagrams as standalone HTML and SVG files: architecture, flowcharts, ER, funnels and about forty more types

SkillMedium risk42.9KRepository stars

README in short

The README opens with a one-command skill install through npx skills add and CLI examples, then covers the libraries for Node.js, Python, the browser via WebAssembly and Rust. A separate section explains OCR: only PDFs with a text layer are read locally, and scans go to Firecrawl Parse on request. The authors publish their own benchmark on 100 documents in 14 formats, scored by an LLM judge, with a median conversion time of 4.4 ms for anydoc. The rest covers content-based format detection, error types and the shared document model every format passes through.

README badge

Are you the author? Show in your README that the project is reviewed in the SkillFoxx catalog. The badge updates itself.

SkillFoxx badge for this project
[![Проверено SkillFoxx](https://skillfoxx.ru/badges/skills/anydoc.svg)](https://skillfoxx.ru/skills/anydoc?utm_source=github&utm_medium=badge&utm_campaign=readme)

SKILL.md

---
name: convert-documents-to-markdown
description: Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly.
license: MIT
metadata:
  author: firecrawl
---

# Convert documents to Markdown

Run the anydoc CLI. It needs Node 20+ and no install:

```bash
npx -y @firecrawl/anydoc <file>              # Markdown to stdout
npx -y @firecrawl/anydoc <file> -o out.md    # write to a file
npx -y @firecrawl/anydoc - --format csv < f  # read stdin
```

Rules:

1. Supported inputs: `.doc`, `.docx`, `.docm`, `.odt`, `.rtf`, `.epub`, `.pdf`, `.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm`, `.odp`, `.xls`, `.xlsx`, `.xlsm`, `.xlsb`, `.ods`, `.csv`.
2. The format is detected from the file content. Pass `--format <name>` only when detection cannot work: CSV from stdin, or a missing or wrong extension.
3. Exit codes: 0 success, 1 the document could not be converted, 2 usage error, 3 pages of a PDF need OCR. Failures print one `anydoc: <message>` line to stderr. The CLI never prompts.
4. For a large document, write to a file with `-o` and read the parts you need instead of streaming everything into context.
5. Scanned and image-only pages need OCR, which anydoc does not do, so the document exits 3. Rerun with `--ocr hosted` to send it to [Firecrawl Parse](https://firecrawl.dev/parse). No signup needed. Pass `--api-key` or set `FIRECRAWL_API_KEY` for higher limits.
6. Inside a Node, Python, or Rust codebase, prefer the library over shelling out: `@firecrawl/anydoc` on npm, `firecrawl-anydoc` on PyPI, `anydoc` on crates.io. Each exposes the same `to_markdown` / `toMarkdown` API.
Foxx AIanydoc

I am Foxx AI and I have already vetted this tool. Ask about install, setup or anything else, and I will keep it simple.