Skip to main content

Architecture

PDQ combines advanced automation with end-to-end active metadata management. It is the foundation for data operations that are efficient, transparent and fully automated.

By integrating PDQ into your infrastructure you ensure that modelling, quality, versioning, integration, code generation and orchestration hold together as one unit β€” not as a collection of loosely coupled tools.


From source to delivered value​

Ingest, Data Lake Service and Data Warehouse Automation in one coherent flow β€” with catalogue, orchestration, automation, observability and traceability across the whole chain.

PDQ architecture: data sources pass through Ingest into Data Lake Service (Landing, Raw Archive, Trusted, Profile) and on to Data Warehouse Automation (Published, Integrated, Business) and out to delivery.

PDQ’s reference architecture. Every step is driven by active metadata, which makes the flow traceable from data source to delivered report.

The same flow is rendered live on the PDQ dashboard, drawn from the systems, source files and models actually configured in your installation. Solid lines are deliveries that go via the Ingest Agent, dashed lines are direct transfers, and the icon on each source shows whether it is a database, an API, a file source, a custom connector or a manual delivery.

PDQ dashboard flow: configured data sources on the left, then Ingest, DLS (Landing, Raw Archive, Trusted, Profile) and DWA (Published, Integrated, Business)

Every stage owns one responsibility. That is what makes a fault locatable, and a layer replaceable without rewriting the rest. Metadata drives the run: lineage, versions and compliance fall out of the flow rather than being separate projects.

The stages in detail​

Sources β€” Data sources

What Files, APIs, streams and databases β€” the source systems as they actually are, with no preparation.

Why Sources change without asking. The platform has to absorb that without anyone rewriting a pipeline.

How Each source is described as metadata: connection, schema and load window. A new source is a configuration, not a project.

Ingest β€” Ingest Agent

What The agent pulls data and lands it unchanged.

Why What arrived has to stay separable from what we did to it, or no reconciliation is possible after the fact.

How The agent runs close to the source, writes files untouched, and stamps every delivery with its time and origin.

DLS β€” Data Lake Service

What Landing β†’ Raw Archive β†’ Trusted, with Profile alongside.

Why Raw data is preserved unchanged while normalisation happens on the way to Trusted. Two different jobs, two different layers.

How Profiling and quality rules run as part of the flow and record deviations against the data contract β€” see What the platform records, and what it stops.

DWA β€” Warehouse Automation

What Published β†’ Base β†’ Core β†’ Business. From the read-in delivery to an integrated, historised model and finally a business-facing delivery layer.

Why Core is the heart of it: a normalised, historised model where every value carries a validity period. That is what makes "what did we know then?" answerable.

How The model is the source. Load logic, historisation and point-in-time views are generated from it β€” nobody writes historisation logic by hand.

Delivery β€” Out to the business

What Reports and BI, APIs, AI and ML workflows, and direct SQL access.

Why Every consumer reads the same quality-assured base. One number β€” not one number per tool.

How The business layer exposes business-facing views over the normalised model. The complexity sits next to the model, not next to the user.

Active metadata β€” Capabilities around the whole chain

What Data catalogue, orchestration, automation, observability and traceability β€” observability covering both the live monitoring views and the record they leave behind.

Why They wrap the whole chain instead of living in six separate tools. The metadata does not just describe the system β€” it runs it.

How Active metadata drives every step, so lineage, versions and compliance fall out of the run rather than being documented beside it.


Technology-agnostic by design​

The platform sits on top of your stack β€” not the other way around.

Change the database, change the storage, change the cloud. The model, the logic and the metadata come with you. That is what turns vendor leverage from a talking point into a fact.

The layers are independent of each other: changing database or storage is a configuration change, changing cloud a migration.

LayerSupported options
Any databaseSnowflake, Databricks, SQL Server, Synapse, Fabric, PostgreSQL
Any storageAzure Blob, Amazon S3, CEPH, Swift, MinIO
Any cloudAzure, AWS, Cleura

What the platform actually does​

Modelling and code generation​

The data model is the source. Load logic, historisation and structures are generated from it β€” consistent no matter who builds it, and without hand-written ETL that drifts out of sync over time.

Governance in the target, not beside it​

The same generation run that builds the tables also documents and classifies them. Column and table comments, primary keys, foreign keys, unique constraints, category tags and sensitivity tags are emitted from the declarations already made in the data contract and the model β€” so a classification set once on a source field arrives on the target column, including on the normalised child tables that inherit the key.

β†’ What lands in the target environment

Quality and profiling​

Profiling and quality rules run as part of the flow, not as an afterthought. Anomalies are caught and recorded where they occur.

Orchestration and operations​

Dependencies, run order, re-runs and alerting are handled centrally. Production-ready from day one instead of months of scaffolding.

The orchestration is automatic, and it rests on three things:

  1. Late arriving dimension handling β€” key generation, so a transaction that references master data nobody has delivered yet still gets a key to hang on.
  2. Model dependencies β€” what has to be loaded before what, read off the model.
  3. Loosely coupled sources β€” a source that is late or absent does not hold up the sources that are ready.

Together those are why nobody writes a run order by hand. Additional Tasks and QPI checks chain onto that automatic flow rather than replacing it.

Load Group and Load Step fall out of the mappings; they are not configuration. The mappings resolve into a schedule, and that schedule has Load Groups and Load Steps in it β€” a Load Group being, for instance, all the attribute loads, and a Load Step a single action in Base. Neither has any machine function. They exist so that it is legible how the loads hang together and in what order they happen, and so that the console can be filtered by them.

A mapping group is a different thing entirely: a logical grouping of the rules for one particular load, from one source to some number of target tables.

Active metadata​

The metadata does not just describe the system β€” it runs it. Lineage, versions and compliance are by-products of execution, not separate projects. It is also portable: the REST API serves the whole repository, and every execution emits OpenLineage events, so lineage reaches the catalogue you already run rather than staying inside the console.


What the platform records, and what it stops​

PDQ works by discovery. Profiling and quality rules run as part of the flow and record deviations against the data contract β€” they do not block the delivery. A field that disappears, a field that appears, a value whose content changes: all of it passes, with the deviation logged, because which deviations matter is decided by the use case consuming the data, not by the platform.

Two things do stop, at different points in the chain. A corrupt file is stopped in Raw and never reaches Trusted; the original stays in the Raw Archive. A missing primary key stops the model load in DWA β€” by then the data has already passed Trusted and Published.

WhatWhereConsequence
Corrupt fileIn RawNever reaches Trusted. The original is kept in Raw Archive.
Missing primary keyAt the DWA loadThe data has already passed Trusted and Published. The model load does not happen.

Trusted is therefore not a quality gate. Because nothing is filtered out between the zones, a difference in the file counts means something is stuck, not that something was discarded.

Expected is an assumption made by an operator or a developer. It is a pointer, not a verdict β€” a deviation against Expected means the assumption or the delivery needs looking at, not that the system has failed.

Contract Validation​

Recorded deviations surface in Contract Validation, which is the discovery promise in concrete form. It reports three kinds of deviation:

DeviationWhat it means
Data Type MismatchValues no longer match the type the contract declares.
Missing In ContractThe source started sending something the contract does not describe.
Missing In DeltaThe source stopped sending something the contract does describe.

One limitation is worth knowing: the view compares today's profile data against the contract. It is a current-state view, not a history of drift.


The model chain​

The chain is Published β†’ Base β†’ Core β†’ DM. Each step has more than one name, because the layers correspond to patterns that were named elsewhere first:

StepAlso known as
PublishedBronze
BaseEnsemble Model
CoreIntegrated, Silver
DMData Mart, Gold, Business

Base is a real step, and the flow diagrams above still show three. Data is read into Published, integrated in Base using ensemble modelling, consolidated in Core, and shaped for a use case in DM.

The DataOps Console labels the last step Data Mart in Data Lineage. That is expected to change; until it does, this documentation writes Business (Data Mart in the console) where the console label matters.

Stage has been retired as a name. The read-in step it used to describe is now Published, loaded with the techniques described in Target platforms β€” COPY INTO from a named stage on Snowflake, read_files() on Databricks, OPENROWSET(BULK …) in the T-SQL family and a pg_lake foreign table on PostgreSQL. It is a rename, not a removal: the mechanism is the same, the layer is called Published.


DLS and DWA β€” what happens where​

StageLayerWhat happens
IngestInto the platformThe ingest agent pulls data from files, APIs, Message Queues and databases and lands it in Landing without altering it.
DLSData Lake ServiceLanding β†’ Raw Archive β†’ Trusted, with Profile alongside. Raw data is preserved unchanged; normalisation happens on the way to Trusted.
DWAData Warehouse AutomationPublished β†’ Base β†’ Core β†’ Business. From the read-in delivery to an integrated, historised model and finally a business-facing delivery layer.
DeliveryOut to the businessReports, APIs, dashboards, AI and ML workflows and direct SQL access β€” all from the same quality-assured foundation.

Data processing zones​

ZonePurpose
LandingThe entry point for data, where it is initially ingested.
Raw ArchiveStores raw data indefinitely, organised and tagged for audit purposes.
TrustedProcesses and publishes data in a standardised, easily consumable format. How long a delivery stays available here for a Published rerun is set by KEEPDATAINTRUSTEDFORDAYS β€” default 3 days.
ProfileAnalyses data to discover new and missing data points without interrupting the loading process.
PublishedProvides the selected table format for the agreed source delivery, facilitating SQL-based data access.
BaseThe ensemble integration step, where sources are merged on business keys. Also called the Ensemble Model.
CoreAn integrated core business concept data model using ensemble data modelling methodology. Also called Integrated or Silver.
BusinessUse case-driven business data sets designed for specific scenarios, accommodating unique business and domain-specific definitions and rules. Also called DM or Data Mart β€” the console’s Data Lineage view labels it Data Mart.

Next steps​