🗒️ PDF Comment & Markup Extractor

Turn a marked up PDF into a punchlist. Every sticky note, highlight, strikeout, text box and drawing markup comes out with its page, author and date, and for highlights, strikeouts and shapes the text they sit on comes with them. Download as Excel, Word or CSV.

Drop a marked up PDF here
or click to choose a file

PDF text is stored in runs, not in single letters, so a highlight over part of a line usually gives back the whole run. Treat the recovered text as a locator for the spot, not as an exact quote. Turning this off makes a big file scan faster.

Your PDF is never uploaded. It is read inside this browser tab. The page pulls pdf.js, its font maps, and the Excel and Word writers from public CDNs, and it records an anonymous page view for Castwright plus a count of how many markups were found and exported. Your file, its text and its comments never leave the tab.

This PDF is password protected

Enter the password to read its comments. It is used in this tab only and is not sent anywhere.

No comments or markup in this PDF

Nothing was found in . That normally means one of these:

Downloads carry every comment that matches the filters above, not only the ones drawn on screen. The Excel file adds empty Assigned to and Done columns so it works as a tracker straight away.

What comes out. Sticky notes, text boxes, highlights, underlines, strikeouts, squiggly marks, boxes, ovals, lines, polygons, freehand ink, stamps, carets and file attachment markers, each with its page, author, creation and modification date, colour, review status and any replies. Popup windows are dropped because they only repeat their parent comment, and form fields and hyperlinks are left out because they are not review markup.
🎬 Plus PDF signing, fillable forms, bookmarks & 300+ more free toolsOpen Castwright

The comment summary, without the subscription

Acrobat Pro can export a comment list, and Bluebeam Revu can build a markup summary. Both are paid seats, and both are usually held by one person on the team while everyone else waits for the file. This page does the same read on any PDF you already have, in the browser, for nothing.

It is aimed at the people who live in review cycles: architects and engineers turning a redlined drawing set into a punchlist, editors pulling an author's tracked notes out of a proof, lawyers collecting the other side's margin notes, and agencies who just got a proposal back covered in client scribbles and need it as a task list instead of a re-read.

How the text underneath is recovered

A highlight, underline, strikeout or squiggly mark stores the quadrilaterals it covers. This page reads those, then reads the page's own text layer and matches the two by position, so the marked words come back as text. For a box, oval, polygon or freehand scribble there is nothing marked as such, so the text sitting inside the shape is returned instead, and only when the shape covers less than half the page.

One honest limit worth knowing before you rely on it: a PDF stores text as runs, not as individual letters, and pdf.js exposes it at that granularity. A highlight over three words in the middle of a line will usually return the whole run those words belong to. It puts you on the exact spot, it is not a legally exact quote. Pages with no text layer at all, a pure scan for instance, return nothing here, because there is no text to find. Pages that store their text rotated return nothing either, and the status line says so when that happens.

Dates, authors and review status

Dates are stored inside the PDF as the commenter's own local wall clock, so a note written at 9am in Karachi reads as 9am even if you open it in Toronto. Both the creation and the modification date come through. Sort by creation date when you want the order the review actually happened in, because the modification date changes every time the file is saved again.

When a reviewer marks a comment Accepted, Rejected, Cancelled or Completed, Acrobat records that as a hidden reply rather than a field on the comment. Those replies are folded back into a Status column here instead of cluttering the list as empty rows.

What it will not do

It reads, it never writes, so your original PDF is untouched. Comments that were flattened into the page are no longer annotations, so no PDF tool can read them back as comments, and only OCR of the page image would recover the words, without the author, date or replies. Markup held in a separate FDF or XFDF file is not read, only what is inside the PDF itself. Password protected files work once you supply the password. Very large files are limited to 100 MB, 2,000 pages and 20,000 markups, because the whole thing is held in one browser tab. When a run hits one of those limits the page says so in red, since the download is then incomplete. In the CSV export a cell that begins with an equals sign, a plus, a minus or an at sign gets a leading apostrophe, so a comment cannot run as a formula when Excel opens it.