Skip to content
KOR IT

Services / 02

Data Engineering

We treat security telemetry as an engineered data estate — sources, pipelines, schemas, lineage, retention and cost — rather than as whatever happens to reach the SIEM. The same discipline applies to the analytical and AI workloads that increasingly consume the same data.

Security Data

The problem

Security data estates are rarely designed. They accrete. A source is onboarded for one detection, a second team adds a parallel feed, a third builds a report on a field whose meaning changed two years ago. Nothing is wrong in isolation; collectively the estate becomes something no one can reason about — and every downstream capability, detection and AI alike, inherits that.

The symptoms are familiar: an analytic that silently stops matching after an upstream format change; a cost line that grows faster than coverage; a field that means three different things depending on which pipeline populated it; a compliance question that takes two weeks to answer because lineage exists only in people's heads.

Capabilities

What this pillar covers.

Grouped by the part of the system they act on.

Architecture

  • Data architecture
  • Security data engineering
  • Data platforms
  • Telemetry architecture

Movement

  • Data pipelines
  • Data integration
  • Normalisation & enrichment
  • Routing & tiering

Governance

  • Data governance
  • Metadata
  • Lineage
  • Data quality
  • Retention & classification

Consumption

  • Analytics
  • Observability
  • Feature and context stores

Approach

How an engagement runs.

The sequence matters more than the tooling. Skipping a stage moves its cost later, it does not remove it.

  1. 01

    Inventory

    Catalogue what is actually flowing, including the feeds nobody owns any more. Most estates contain both undetected gaps and expensive duplication.

  2. 02

    Model

    Design the schema, the classification and the routing rules. Decide deliberately what is kept hot, what is tiered, and what should never have been collected.

  3. 03

    Build

    Implement the pipelines with tests, quality gates and lineage emitted as a first-class output rather than reconstructed later from documentation.

  4. 04

    Govern

    Put the operating model in place: ownership, change process, quality SLOs and the metadata that makes the estate navigable by people who did not build it.

What you receive

Concrete artefacts, in your repositories and your platforms.

  • A source inventory with owner, schema, volume, fidelity and retention for every feed
  • Pipeline architecture: collection, normalisation, enrichment, routing and tiered storage
  • A schema and naming standard that survives contact with a second team
  • Lineage from source to consumer, so an analytic's dependencies are inspectable
  • Data quality checks that run continuously and fail loudly, not quarterly and quietly
  • A cost model — which data is expensive, which is load-bearing, and where those two differ

Expected outcomes

Properties of the resulting system — not performance figures, which depend on your estate rather than on us.

  • One inventory of sources, with owners, schemas and known fidelity limits
  • Pipelines with tests and quality gates, so upstream changes fail visibly
  • Lineage that answers 'what breaks if this field changes' in minutes
  • A deliberate cost position rather than an accidental one
  • Data that AI and analytics workloads can consume without re-deriving its meaning

Data Engineering, on your estate.