Platform guide
PDQ follows a layered architecture where data flows through three stages — from raw source extraction to a modelled warehouse. Each layer has a clear responsibility and a well-defined handoff to the next.
DATA SOURCES ──► INGEST ──► DLS ──► DWA
| Layer | Full Name | Responsibility |
|---|---|---|
| INGEST | Ingestion | Extract data from source systems (databases, APIs, files) and land it in the data lake |
| DLS | Data Lake Service | Organise, archive, profile and publish raw data through the storage zones |
| DWA | Data Warehouse Automation | Model, map and transform data into structured warehouse models |
QPI is not a fourth stage. It is the framework's data quality engine: SQL controls run against data in DWA, placed on Published, Core and Business. It reports, and cannot block a load or stop downstream work. It is documented with DWA.
The deep dive, layer by layer
| INGEST layer | Connection types, export types and strategies, scheduling, column selection, transformations, destination paths and the sourcefile structure. |
| DLS layer | The storage zones and their configuration, the DLS pipeline, publish patterns and the considerations behind them. |
| DWA layer | Models and attributes, source-to-target mappings, transformations, relationships, worked modelling scenarios — and QPI. |
This page used to carry all three, plus QPI, in a single tab set. The tabs hid their contents from the table of contents and from any deep link, which is why each layer now has its own page.
How the layers connect
Source Systems → INGEST
Source systems are registered with a connection type (Database, API, File, Custom, or Manual). Each system can have one or more source exports that define how data is extracted — what to query, which columns to include, how often to run, and what output format to produce.
INGEST → DLS
Exported data lands in the Landing Zone. From there, the DLS pipeline archives it to the Raw Zone, processes it through DLS Workers, standardises it into the Trusted Zone, and profiles it in the Profile Zone.
DLS → DWA
Trusted data is synced/published into the Published Zone, where it becomes available for warehouse modelling. The DWA layer maps source fields to model attributes along the model chain: Published → Base → Core → DM.
DWA and QPI
Quality checks are placed on the layers DWA loads. Mapping coverage percentages, field-level governance flags and data profiling results give continuous feedback on pipeline health — reported, never enforced.
Zone architecture
Data passes through a series of storage zones, each with a specific purpose:
┌──────────┐ ┌──────────────┐ ┌─────────────┐ ┌───────────────┐
│ Landing │ ──►│ Raw Archive │ ──►│ Trusted │ ──►│ Published │
│ Zone │ │ Zone │ │ Zone │ │ Zone │
└──────────┘ └──────────────┘ └─────────────┘ └───────────────┘
Temporary Immutable Standardised Query-ready
staging audit trail format tables
Profile sits alongside Trusted rather than after it. Nothing is filtered out between the zones — a deviation against the data contract is recorded, not blocked, so a shortfall in the counts means something is stuck rather than discarded. See What the platform records, and what it stops.
Zone names are configurable in Settings, so an installation may display its own labels. Paths follow a convention:
[ZoneName]/[System]/[Filename]/[YYYY]/[MM]/[DD]/
Landing takes no date tokens. The zone configuration in full is on the DLS layer page.
Model tiers (DWA)
The warehouse chain runs Published → Base → Core → DM:
| Tier | Also known as | Purpose | Example |
|---|---|---|---|
| Published | Bronze | The SQL-readable landing point for a delivery, read straight from the trusted zone | One table per sourcefile |
| Base | Ensemble Model | Integrate data from multiple sources with consistent rules | Merge CRM + ERP customer records |
| Core | Integrated, Silver | Final deduplicated business entities for consumption | Clean Customer, Order, Product tables |
| DM | Data Mart, Gold, Business | Models shaped for a specific analytical use case | A sales star schema |
The console still labels the last step Data Mart in Data Lineage. See the model chain for the full synonym table, and why Stage is no longer one of these names.
Terminology
Every term the platform uses is defined once, in the Glossary. This page used to carry its own table of eight of them, which is how they came to be defined twice and differently.
Next steps
- Architecture — the reference architecture and what the platform records versus stops
- Data management capabilities — the full capability list, with links into the detail
- Your first pipeline — the end-to-end path from a source file to a scheduled data product