Technical architecture

Document processing, retrieval, and recovery

GAIA Civil treats a project set as a chain of addressable artifacts and validated handoffs rather than one enormous prompt.

Reference version 1.1 · Updated September 4, 2026

System layers

System layers
LayerResponsibility
IntakeRegister project, tenant, document category, content hash, and processing state.
IngestionCreate extracted text, rendered images, single-page PDF assets, visual summaries, and localized components.
RetrievalStore scoped text and image representations with document, page, category, tenant, and provenance metadata.
AnalysisRun validated specialist stages against typed project needs.
OutputAssemble findings, quantities, citations, warnings, and report contracts.
ReviewHold external action behind named human ownership and approval.

Retrieval behavior

  • Filter by tenant, project, document, page, category, and authorized context.
  • Use the representation that matches the question: text, full page, drawing, table, chart, or localized region.
  • Expand, deduplicate, rerank, and preserve source addresses for the selected evidence.

Failure and recovery

  • Attempt ownership prevents ambiguous concurrent work.
  • Resumable provenance identifies which source and stage can continue.
  • Page-level isolation prevents one failed page from silently invalidating the entire set.
  • Marked fallbacks retain extracted text or raw page assets without claiming complete analysis.
  • Warnings survive into the customer-visible report.

Detail: eight ingestion stages

Detail: eight ingestion stages
StageWork performedArtifactFailure behavior
01 / Validate and admit the sourceThe upload path accepts one or more project PDFs, sanitizes filenames, validates file type, reads page count, calculates a SHA-256 content hash, checks for duplicate content, and links the document to a new or existing project.Document ID, filename, project relationship, tenant/category context, page count, content hash, and pending processing state.Non-PDF input, duplicate content, invalid project scope, missing permissions, or unmet credit/approval requirements are rejected before full processing.
02 / Reserve the processing pathFull ingestion estimates and reserves Analysis Credits before processing; fast screening uses its separate lightweight lane. Attempt IDs and ownership prevent an older worker from overwriting a newer run.Ingestion mode, credit reservation where applicable, attempt identity, queue/status metadata, and promotion state.Insufficient balance, missing confirmation, or required administrative approval keeps the document from entering full processing.
03 / Persist the original sourceThe worker resolves or restores the original PDF, records storage metadata, and carries the source hash into embedding provenance before parsing. The original is retained after temporary chunks are removed.Original PDF reference, storage metadata, source hash, and resumable provenance.A missing source stops the run. A required durable-source upload failure stops parsing rather than leaving an unrecoverable index.
04 / Split work into bounded page unitsLarge PDFs are divided into bounded temporary page ranges. Chunk jobs run concurrently under a semaphore, and each range retains an offset to the original document.Temporary page-range files and an exact mapping back to original page numbers.A failed range becomes a recorded partial-ingestion event while successfully processed ranges remain usable.
05 / Extract text and native page assetsThe parser extracts embedded text, generates Markdown when available, falls back to native page text when necessary, renders a plan-resolution PNG, and creates a single-page PDF.Markdown or text, rendered page image, one-page PDF, source page number, and page dimensions.Extraction failures are isolated to the affected page and added to the partial-coverage record.
06 / Build multimodal representationsText is split into overlapping chunks. The complete PDF/image/text page receives a native page representation, and the rendered sheet receives a native image representation in the same search space.Text-chunk records, full-page analysis record, full-page image record, asset IDs, and embedding model/schema metadata.If page analysis fails, GAIA Civil can retain the extracted text and raw page assets as a marked stub rather than calling the page fully analyzed.
07 / Read visual structure and localize componentsA structured visual pass classifies the page and can recover document context, OCR-like text, specifications, measurements, retrieval keywords, and components/materials. Valid normalized bounding boxes are cropped into component assets.Structured page summary, component text, normalized bounding boxes, cropped images, and component-level text/image embeddings.Invalid bounding boxes are skipped. Missing visual analysis never erases the source page or its text representation.
08 / Index, complete, and expose coverageRecords are batch-upserted with tenant, category, document, page, type, source hash, embedding model, and schema. Matching pages can be reused on resume; completion records partial failures and promotion state.Mixed text/image retrieval collection, companion asset collection, completion state, indexed-page coverage, and partial-failure details.A catastrophic failure marks the document failed. Page/chunk failures remain visible while successfully indexed evidence can still be reviewed with warnings.

Detail: addressable page artifacts

Detail: addressable page artifacts
ArtifactPurposeIdentity
Single-page PDFPreserves the page as a citation-addressable source fragment.document · page · source hash
Rendered page imagePreserves drawings, tables, callouts, and layout that extracted text may not express.document · page · image asset
Text chunksMakes clauses, notes, schedules, and written requirements retrievable at smaller semantic units.document · page · chunk index
Full-page structured summaryOrganizes page type, context, specifications, measurements, keywords, and commercial signals.document · page · processing status
Component records and cropsMakes a detected material, detail, or component searchable independently when a valid location exists.document · page · component · bounding box
Embedding provenanceDistinguishes artifacts created from a different source, model, or schema during resume and maintenance.source hash · model · schema · timestamp

Detail: retrieval sequence

Detail: retrieval sequence
StageBehavior
ScopeResolve active, processed project document IDs and the approved library categories available to the workflow.
ExpandCombine the user question with configured vector queries and tenant product/application language.
RetrieveSearch relevant project, product, and engineering-manual collections concurrently.
UnifyMerge duplicate nodes by stable ID so fan-out does not shift the evidence ledger.
RerankRerank text against the primary question while retaining a bounded set of strong native-image matches.
AddressCarry document, page, category, asset, and score metadata into the analysis context and report sources.
Need help applying this page to an evaluation?Ask a product question