← Archive

The Toolkit·October 9, 2026·Read by 100K+ subscribers

Stop Pasting Whole PDFs Into ChatGPT

Microsoft's local converter turns PDFs, Office files, and HTML into Markdown before they hit a paid model. Free, MIT, and easy to install wrong. Here is the install, the trap, and the 10-minute start.

Debrief · No. 33 · Fri Oct 9

MarkItDown: Stop Burning Tokens on PDFs

Microsoft's local converter turns PDFs, Office files, and HTML into Markdown before they hit a paid model. Free, MIT, and easy to install wrong. Here is the install, the trap, and the 10-minute start.

Friday is the tools half of The AI Debrief. One free Academy pick you can run today, one workflow to steal, and two more free tools if you have another 20 minutes.

In This Email

» MarkItDown: what it replaces, what it costs, and the [all] extras trap

» Steal this: a 3-step start that turns a PDF into clean Markdown locally

» Two more free tools: Creen AI and Ruflo

This Week's Tool From the Academy: Convert Locally Before You Paste

MarkItDown

MarkItDown is Microsoft's open-source converter for PDFs, Office docs (Word, PowerPoint, Excel), HTML, and more into Markdown you can feed a model without pasting a blob. It is MIT and free forever. Live GitHub stars sit around 188k, but the useful fact is the local convert, not the vanity number.

What it replaces: dumping whole PDFs into ChatGPT or Claude and watching the token meter climb for layout you never needed - headers, footers, page numbers, and whitespace that burn context before the model even sees the paragraph you care about.

Why local beats paste: the file never leaves your machine for the convert step. You get clean Markdown you can skim, trim, and chunk before a paid call. You stop paying for OCR-ish noise. You keep a reusable .md artifact for the next prompt instead of re-uploading the same PDF every time.

Cost, honestly: $0 for the tool. You still pay for whatever model you send the Markdown to. The win is fewer tokens per answer and fewer retries when the model hallucinates from a messy paste.

The [all] trap - before / after:

Before: you run pip install markitdown, point it at a PDF, and get a missing-converter error (or an empty shell of output). Bare install skips the optional extras that actually handle PDF and Office. You blame the repo. You go back to pasting the PDF into chat.

After: you run pip install 'markitdown[all]' (or install only the extras you need). The same PDF converts to Markdown on your laptop in one command. You open the .md, cut the boilerplate, paste the cleaned section into your model, and the answer lands on the content instead of the page chrome.

That install line is the whole lesson. Everything else is muscle memory.

Steal This

10-Minute Start

  1. Install with pip install 'markitdown[all]' so PDF and Office converters come along. Confirm with markitdown --help (or the Python import path you prefer) before you touch a real file.
  2. Pick one real file you would have pasted into a chat this week - a vendor PDF, a slide deck, a Word brief. Convert it locally to Markdown. Open the .md and delete headers, footers, and junk pages before any model sees it.
  3. Paste only the cleaned Markdown into your model. Ask the same question you asked last time with the raw PDF. Compare tokens used and answer quality side by side. Keep the .md in the project folder so the next prompt starts from clean text, not another upload.

If PDF support is missing, you skipped the [all] extras. Fix that before you blame the repo. If the Markdown still looks noisy, trim it yourself - the tool gets you out of binary land; you still own the edit pass.

Two More Free Tools

Creen AI

One tab for 40+ frontier image and video models instead of hopping dashboards. Free plan includes daily free credits that reset each day (media storage on free is 7 days). Pro from $29/mo. Trap: burning the day's free credits on a heavy video model, then thinking the free plan is broken until the daily reset. creen.ai.

Try this Monday: open Creen, spend one free credit on a still (not video) for a real asset you need this week - a thumbnail, a slide hero, a social crop. Note which model you picked and how many credits remain before you touch a video job.

Ruflo

A meta-harness that turns Claude Code into coordinated agent swarms with memory. MIT, free. Live GitHub stars sit around 74k. Install via npx ruflo init / npm ruflo. Trap: expecting 100 agents on first boot without Claude Code plus marketplace plugins. Init first, then /ruflo. github.com/ruvnet/ruflo.

Try this Monday: run npx ruflo init in a scratch repo you already use with Claude Code. Confirm Claude Code and marketplace plugins first. Start with one small chore (tests or lint cleanup) so you see memory and handoff without burning a morning on a 100-agent fantasy.

Go Deeper

The full MarkItDown walkthrough is already in the AI Resource Vault: MarkItDown - Stop Burning Tokens on PDFs. It covers the install, the extras trap, and when local convert beats pasting the blob.

The other two picks have their own vault lessons too: 40+ Frontier Video and Image Models in One Tab (Creen) and Turn Claude Code Into a 100-Agent Team (Ruflo).

Join the free AI Academy →

Free, no card. Guides, tools, and the people running them live here.


That's the Debrief.

See you Tuesday for the news.

From Drew and The AI Debrief team


P.S. Did markitdown[all] land clean on the first try, or did you hit the bare-install trap? Hit reply. We read every one.

P.P.S. New here, or skipped one? Every past edition is in the archive.