CDT
Conduit
Read and normalise
What it does
Conduit is the entry point. It resolves heterogeneous formats, encodings and metadata schemas into one typed, addressable record, so every downstream module works against the same shape regardless of the source.
Partial inputs, inconsistent naming and duplicates across sources are detected at the point of entry, where they are a query, rather than downstream, where they are a correction.
What it accepts
Any source the pipeline is pointed at, each one mapped to a consistent typed output in code. One integration profile is live today; adding a new source means writing a new field map.
- Documents
- Files and their accompanying metadata, arriving singly or as mixed batches under one identifier. The specific formats supported are named in the agreed scope.
- Source mappings
- The mapping from each source to the typed record. The field map is defined in code, so a new source type is a code change.
What it emits
- Typed outputs
- One addressable output per input, carrying the resolved fields, the source identity and the ingest history.
- Entry findings
- Structured findings for incomplete inputs, naming conflicts and cross-source duplicates, raised at the point of entry.
How it is configured
Configuration belongs to the client. It encodes what your specialists already know, and it is versioned like code.
- Source mappings
- The mapping from each source to the typed record, defined in code.
- Identity rules
- What makes two inputs the same input.
- Completeness expectations
- What a complete input contains, per source.
What it produces as evidence
- Ingest record
- A versioned record of what arrived, from where, when, and how it was resolved, for every input.
What it is tested against
Conduit is tested against acceptance criteria agreed before build: named sources resolved to the typed output, defined completeness and duplicate cases detected, and the ingest record produced for every input.