Skip to content

CDT

Conduit

Read and normalise

1.0Purpose

What it does

Conduit is the entry point. It resolves heterogeneous formats, encodings and metadata schemas into one typed, addressable record, so every downstream module works against the same shape regardless of the source.

Partial inputs, inconsistent naming and duplicates across sources are detected at the point of entry, where they are a query, rather than downstream, where they are a correction.

2.0Input

What it accepts

Any source the pipeline is pointed at, each one mapped to a consistent typed output in code. One integration profile is live today; adding a new source means writing a new field map.

Documents
Files and their accompanying metadata, arriving singly or as mixed batches under one identifier. The specific formats supported are named in the agreed scope.
Source mappings
The mapping from each source to the typed record. The field map is defined in code, so a new source type is a code change.
3.0Output

What it emits

Typed outputs
One addressable output per input, carrying the resolved fields, the source identity and the ingest history.
Entry findings
Structured findings for incomplete inputs, naming conflicts and cross-source duplicates, raised at the point of entry.
4.0Config

How it is configured

Configuration belongs to the client. It encodes what your specialists already know, and it is versioned like code.

Source mappings
The mapping from each source to the typed record, defined in code.
Identity rules
What makes two inputs the same input.
Completeness expectations
What a complete input contains, per source.
5.0Evidence

What it produces as evidence

Ingest record
A versioned record of what arrived, from where, when, and how it was resolved, for every input.
6.0Acceptance

What it is tested against

Conduit is tested against acceptance criteria agreed before build: named sources resolved to the typed output, defined completeness and duplicate cases detected, and the ingest record produced for every input.