Full Document Extraction
When you need the data—not just the text.
Nasaas combines recognized text, page structure, and visual regions into one ordered, typed document model built for people and systems.
17 page-layout classes26 figure subtypes1 structured document model
Explore Full Document Extraction- Input files
PDFs, office documents, spreadsheets, presentations, web and markup files, scans, and phone-captured document images.
Supported input extensions File category Supported extensions PDF Text documents - .doc
- .docx
- .odt
- .rtf
Spreadsheets - .xls
- .xlsx
- .ods
Presentations - .ppt
- .pptx
- .odp
Web & markup - .html
- .htm
- .md
Scans & document photos - .png
- .jpg
- .jpeg
- .tif
- .tiff
- .webp
- Document structure analysis
Extracted text, reading order, layout structure, structured values, and visual classifications.
Data recovered into the document model Data type Extracted output Text & reading structure Heading levels, sections, paragraphs, reading order, lists, captions, footnotes, formulas, code, and repeated page headers and footers. Tables & key-value pairs Rows, columns, cells, row and column spans, and table headers when cell structure is available, plus recognized key-value pairs. Selection marks & barcodes Selected and unselected states with associated labels when a reliable match is available, plus decoded barcode and QR values when returned. Figures & visual classes Figures classified across 26 subtypes, including signatures, stamps, charts, maps, diagrams, and screenshots. - Structured outputs
The canonical structured result preserves text, structure, geometry, and source evidence, with five additional exports for downstream systems and workflows.
Available output formats Output Technical format Canonical structured data JSON Formatted Markdown Markdown Web document HTML Plain text TXT Editable document DOCX Searchable PDF PDF