KOR IT LAB / 03
Data Lab
The Data Lab is where security data architecture is modelled end to end so that governance properties can be tested rather than asserted. It is deliberately early, and it reports structure rather than headline numbers.
- Data architecture
- Security data
- Governance
- Metadata
- Lineage
- Data quality
Governed pipeline reference model
Question
Can lineage, classification and quality be emitted as pipeline outputs rather than maintained as separate documentation?
Method
A layered pipeline is modelled end to end — raw ingestion through normalised and curated stages — with governance dimensions attached to each transformation rather than recorded alongside it. The question is whether the metadata a governance function needs can be produced as a by-product of the data movement itself.
- Pipeline stages modelled
- Bronze → Silver → Gold
Lab measurement from a tested configuration in an experimental environment. Not a benchmark, and not portable to other hardware, models or workloads.
Findings
This lab is early and deliberately reports structure rather than numbers. There are no headline throughput or coverage figures here, because any we produced would describe a modelled estate rather than a measured one. Numbers will be reported when there is something real to measure.
- Security telemetry and AI training or retrieval corpora are the same problem wearing different labels: both require provenance, classification and lineage, and both fail in the same way when those are absent.
- Lineage emitted by the pipeline is current by construction; lineage documented beside the pipeline is a snapshot that begins decaying immediately.
- Quality checks are only useful where a failure is routed somewhere a person will see it. A failing check that writes to a dashboard is indistinguishable from no check at all.
Limitations
Next experiments
Written down before they are run, so the results can be judged against what we set out to find rather than what we happened to notice.
Run the reference model against synthetic telemetry at volume to find where the governance metadata becomes the bottleneck.
Test whether classification survives a summarisation or aggregation step, or whether it has to be re-derived.
Model the cost of retaining lineage at record level versus at batch level.