Free local text extraction as the foundation
Many PDFs already contain a text layer from the authoring app. Local extraction pulls that layer so CasperWasp can index and display content. Because this step is free, you can validate readability early — open a known paragraph, confirm names and numbers look right, then proceed to AI-assisted surfaces.
Scanned PDFs without a text layer are a different problem. Extraction cannot invent accurate text from pixels alone without OCR upstream. If chat citations look hallucinated or summaries feel generic, inspect the text layer first. Most “model problems” on PDFs are actually extraction problems.
Healthy text extraction also protects collaboration. When teammates see the same underlying content, disagreements focus on meaning instead of whether someone copied from a corrupted export.