Extraction at the depth your work demands.

Information extraction platform

Read the page, recover the document, or teach a repeatable workflow. Nasaas brings each path into one focused platform.

Full Document Extraction

Service1 of 4

Current capability

When you need the data—not just the text.

Nasaas combines recognized text, page structure, and visual regions into one ordered, typed document model built for people and systems.

17 page-layout classes26 figure subtypes1 structured document model

Explore Full Document Extraction
From input files to structured outputs
  1. Input files

    PDFs, office documents, spreadsheets, presentations, web and markup files, scans, and phone-captured document images.

    Supported input extensions
    File categorySupported extensions
    PDF
    • .pdf
    Text documents
    • .doc
    • .docx
    • .odt
    • .rtf
    Spreadsheets
    • .xls
    • .xlsx
    • .ods
    Presentations
    • .ppt
    • .pptx
    • .odp
    Web & markup
    • .html
    • .htm
    • .md
    Scans & document photos
    • .png
    • .jpg
    • .jpeg
    • .tif
    • .tiff
    • .webp
  2. Document structure analysis

    Extracted text, reading order, layout structure, structured values, and visual classifications.

    Data recovered into the document model
    Data typeExtracted output
    Text & reading structureHeading levels, sections, paragraphs, reading order, lists, captions, footnotes, formulas, code, and repeated page headers and footers.
    Tables & key-value pairsRows, columns, cells, row and column spans, and table headers when cell structure is available, plus recognized key-value pairs.
    Selection marks & barcodesSelected and unselected states with associated labels when a reliable match is available, plus decoded barcode and QR values when returned.
    Figures & visual classesFigures classified across 26 subtypes, including signatures, stamps, charts, maps, diagrams, and screenshots.
  3. Structured outputs

    The canonical structured result preserves text, structure, geometry, and source evidence, with five additional exports for downstream systems and workflows.

    Available output formats
    OutputTechnical format
    Canonical structured dataJSON
    Formatted MarkdownMarkdown
    Web documentHTML
    Plain textTXT
    Editable documentDOCX
    Searchable PDFPDF

Fast Structure-aware OCR

Need speed and predictable cost?

Available now

Choose the fast mode when you need the same structured document model, recognized text, coordinates, page regions, and visual classifications—without table-cell, field-pair, or label-bound checkbox reconstruction.

See the detailed comparison
Full Document Extraction and Fast Structure-aware OCR comparison
CapabilityFull Document ExtractionFast Structure-aware OCR
Structured document modelIncludedIncluded
Recognized text, words, lines, and coordinatesIncludedIncluded
Page regions and visual classificationsIncludedIncluded
Table matrices, cells, and spansIncludedNot included
Field–value pairs and label-bound checkbox statesIncludedNot included
Credits per input page1–3 credits per page1 on ordinary pages; 3 when table or checkbox analysis triggers a layout rerun.1 credit per pageFlat rate with the layout rerun disabled.

Pipeline Factory

Service2 of 4

Limited access

Teach the format. Publish the Pipeline.

Bring one known format—one page or a multi-page package. Nasaas reads every sample page, selected language models propose page templates and fields, and your team refines the candidates into a tested Pipeline that runs through the Pipeline endpoint.

Explore Pipeline Factory
Pipeline Factory authoring and published runtime workflow

Build the Pipeline, page by page.

Model-assisted Pipeline authoring

Each selected language model produces a proposal. Your team compares, refines, tests, and approves what the Pipeline will return.

  1. Teach

    Bring the format as it really arrives: one page or a multi-page package. Full Document Extraction reads every page before each selected language model proposes its own candidate page templates, anchors, fields, types, and source regions.

    Start from a working map—not a blank setup.
  2. Prove

    Compare proposals, refine one with feedback, or edit the draft directly. Each page keeps its own template while its fields join one flat Pipeline field contract.

    Models propose. You set the contract.
  3. Test

    Challenge the draft with fresh documents, partial page sets, or pages in a different order. Inspect matched templates, missing pages, unmatched pages, and states and evidence for fields on matched pages.

    Prove page routing before integration.
  4. Publish

    Freeze the reviewed page templates and field contract as an immutable Pipeline version. New draft edits stay out of production until you publish again.

    Production runs the reviewed contract—not another model guess.

Integrate the Pipeline. Send documents. Get the fields you defined.

Published Pipeline runtime

Call the Pipeline endpoint from your system and send one document from the known format. The published Pipeline matches pages by anchors and relationships rather than position, routes each matched page to its page template, and returns one flat field result shaped by your reviewed contract.

  1. Pipeline endpoint

    Your system submits a document

  2. Page-template routing

    Matched by anchors and relationships—not page number

  3. Field-contract response

    Typed fields from every matched page

What the Pipeline returns

Page-aware outcome
Passed, partial, or mismatch—with explicit states for fields on matched pages
Source-linked evidence
Missing templates, unmatched pages, and source evidence for fields on matched pages while the result is retained
Versioned integration
Run the active Pipeline version or pin one for controlled reruns

Publish the Pipeline. Call its endpoint from your systems. Send the document; get the fields your contract defines.

Arabic-first Multilingual Transcription

Service3 of 4

Coming soon

Planned from Arabic outward, across multilingual recordings.

A planned multilingual path toward time-aligned text, distinct speaker turns, and structured transcripts prepared for later extraction workflows.

See the direction
The direction for Arabic-first multilingual transcription
  1. Audio & video input

    Planned multilingual recordings, beginning with Arabic

  2. Model evaluation

    Planned evaluation of advanced speech and language systems

  3. Timing & speakers

    Intended timing and distinct speaker turns

  4. Structured transcript

    Intended preparation for later extraction workflows

  • Arabic-first, multilingual direction

    Planned from Arabic outward without treating other languages as an afterthought.

  • Planned model evaluation

    We plan to evaluate advanced speech-recognition and language-model systems across multilingual recordings.

  • Planned timing & speaker turns

    A direction toward segment timing, word timing, and distinct speaker turns.

Document Reconstruction

Service4 of 4

Current capability

Bring the document back—not just the data.

Turn understood content into searchable, editable, or layout-aware document outputs without pretending every format is a pixel-perfect replica.

  • Searchable documents
  • Editable documents
  • Layout-aware web output
  • Structured and plain text
Explore Document Reconstruction
From understood content to reconstructed documents
  1. Structured content

    Text, hierarchy, geometry, and evidence

  2. Reconstruction

    Flow or layout-aware rendering

  3. Usable document

    Searchable, editable, or system-ready

Platform Side Services

Selected supporting services

Use one focused part of the platform when the complete document workflow is more than the job requires.

View all side services