| 01 / Validate and admit the source | The upload path accepts one or more project PDFs, sanitizes filenames, validates file type, reads page count, calculates a SHA-256 content hash, checks for duplicate content, and links the document to a new or existing project. | Document ID, filename, project relationship, tenant/category context, page count, content hash, and pending processing state. | Non-PDF input, duplicate content, invalid project scope, missing permissions, or unmet credit/approval requirements are rejected before full processing. |
|---|
| 02 / Reserve the processing path | Full ingestion estimates and reserves Analysis Credits before processing; fast screening uses its separate lightweight lane. Attempt IDs and ownership prevent an older worker from overwriting a newer run. | Ingestion mode, credit reservation where applicable, attempt identity, queue/status metadata, and promotion state. | Insufficient balance, missing confirmation, or required administrative approval keeps the document from entering full processing. |
|---|
| 03 / Persist the original source | The worker resolves or restores the original PDF, records storage metadata, and carries the source hash into embedding provenance before parsing. The original is retained after temporary chunks are removed. | Original PDF reference, storage metadata, source hash, and resumable provenance. | A missing source stops the run. A required durable-source upload failure stops parsing rather than leaving an unrecoverable index. |
|---|
| 04 / Split work into bounded page units | Large PDFs are divided into bounded temporary page ranges. Chunk jobs run concurrently under a semaphore, and each range retains an offset to the original document. | Temporary page-range files and an exact mapping back to original page numbers. | A failed range becomes a recorded partial-ingestion event while successfully processed ranges remain usable. |
|---|
| 05 / Extract text and native page assets | The parser extracts embedded text, generates Markdown when available, falls back to native page text when necessary, renders a plan-resolution PNG, and creates a single-page PDF. | Markdown or text, rendered page image, one-page PDF, source page number, and page dimensions. | Extraction failures are isolated to the affected page and added to the partial-coverage record. |
|---|
| 06 / Build multimodal representations | Text is split into overlapping chunks. The complete PDF/image/text page receives a native page representation, and the rendered sheet receives a native image representation in the same search space. | Text-chunk records, full-page analysis record, full-page image record, asset IDs, and embedding model/schema metadata. | If page analysis fails, GAIA Civil can retain the extracted text and raw page assets as a marked stub rather than calling the page fully analyzed. |
|---|
| 07 / Read visual structure and localize components | A structured visual pass classifies the page and can recover document context, OCR-like text, specifications, measurements, retrieval keywords, and components/materials. Valid normalized bounding boxes are cropped into component assets. | Structured page summary, component text, normalized bounding boxes, cropped images, and component-level text/image embeddings. | Invalid bounding boxes are skipped. Missing visual analysis never erases the source page or its text representation. |
|---|
| 08 / Index, complete, and expose coverage | Records are batch-upserted with tenant, category, document, page, type, source hash, embedding model, and schema. Matching pages can be reused on resume; completion records partial failures and promotion state. | Mixed text/image retrieval collection, companion asset collection, completion state, indexed-page coverage, and partial-failure details. | A catastrophic failure marks the document failed. Page/chunk failures remain visible while successfully indexed evidence can still be reviewed with warnings. |
|---|