Data management capabilities
A single list of what PDQ does for data management, grouped by the job it does rather than by the component that does it. Use it to check whether a capability exists, then follow the link to where it is documented in detail.
Each capability names the layer that owns it β INGEST, DLS, DWA, QPI, AME (Active Metadata Engine) or the DataOps Console.
1. Data acquisitionβ
Getting data out of source systems and into the platform, on a schedule, without hand-written extraction code.
| Capability | What it means | Owner |
|---|---|---|
| Database extraction | SQL sources pulled over a direct connection β PostgreSQL, SQL Server, Oracle and MySQL | INGEST |
| API extraction | REST endpoints configured as source exports | INGEST |
| File extraction | Files collected from object storage and network shares β S3, Azure Blob, mounted paths | INGEST |
| Custom and manual sources | Proprietary connectors, and sources delivered by hand where no automation exists | INGEST |
| Full export | The whole dataset re-extracted each run, for small reference and dimension tables | INGEST |
| Change data capture | Only new, modified and deleted records since the previous run | INGEST |
| Sliding window | A rolling time window re-extracted each run, for late-arriving data | INGEST |
| Scheduling | Frequency, interval multiplier and an earliest allowed start time, per export | INGEST |
| Column selection | Include lists, exclude lists, or everything β decided per export | INGEST |
| Export-time substitutions | String-level substitutions applied during extraction | INGEST |
| Unaltered landing | What arrives is written untouched and stamped with its time and origin, so a delivery stays separable from what was done to it | INGEST |
β INGEST layer Β· Configuring ingestion
2. Storage and lifecycleβ
Where data lives between arrival and consumption, and what each zone guarantees.
| Capability | What it means | Owner |
|---|---|---|
| Landing zone | The entry point β temporary staging for arriving deliveries | DLS |
| Raw archive zone | An immutable audit trail, kept indefinitely, organised and tagged | DLS |
| Trusted zone | Data in a standardised format. Deviations against the data contract are recorded, not blocked β see Architecture | DLS |
| Rerunnable window on Trusted | KEEPDATAINTRUSTEDFORDAYS, default 3 days β how far back Published can be rerun from without reprocessing from Raw | DLS |
| Profile zone | Profiling statistics written alongside the flow, without interrupting loading | DLS |
| Published zone | Query-ready tables in the agreed format, open to SQL access | DLS β DWA |
| Configurable zone names and paths | Zone names set centrally; paths follow a [ZoneName]/[System]/[Filename]/[YYYY]/[MM]/[DD]/ convention; Landing takes no date tokens | DLS |
| Date-token partitioning | [YYYY], [MM], [DD] and [HH] tokens for time-based partitioning of zone paths | DLS |
| Compression | Optional gzip, levels 1β9, per sourcefile | DLS |
| Encryption at rest | Enabled or disabled per sourcefile | DLS |
| Encoding support | UTF-8 (recommended), UTF-16 and Windows-1252 | DLS |
β DLS layer Β· Data lake service
3. Data contracts and profilingβ
Knowing what a delivery is supposed to look like, and noticing when it stops looking like that.
| Capability | What it means | Owner |
|---|---|---|
| Sourcefile definitions | The structure of an ingested file β fields, types, keys and governance flags | DLS |
| Automated profiling | Continuous checks against the data contract, enabled per sourcefile | DLS / QPI |
| Type conformance checks | Verification that values match their declared types | DLS |
| Null and empty ratios | How much of each field is actually populated | DLS |
| Value distribution | Distribution patterns per field, used to spot silent changes | DLS |
| Schema drift detection | New or missing columns flagged as the source changes | DLS |
| Contract validation | Where a source has drifted from its agreed contract: Data Type Mismatch, Missing In Contract (the source started sending something new) and Missing In Delta (the source stopped sending something). Compares todayβs profile against the contract. | DataOps Console |
| Per-field profiling exclusions | Individual fields skipped where profiling is not wanted | DLS |
β Data profiling Β· Contract validation
4. Modelling and warehouse automationβ
Turning source data into a model, and generating the code that loads it.
| Capability | What it means | Owner |
|---|---|---|
| Model chain | Published β Base (integrate) β Core (consume) β Business | DWA |
| Datamart models | Additional models configured for specific analytical use cases | DWA |
| Model definitions | Target tables with typed attributes, business keys, relationships and a business description | DWA |
| Specific and combined models | One source per model, or several sources integrated into one | DWA |
| Loading patterns | full, incremental, transaction, dayspan or none, chosen per model | DWA |
| Business keys | Single or composite keys with an explicit priority order, driving deduplication and upserts | DWA |
| Relationships | Foreign keys between models, including 1:N and M:M | DWA |
| Early-arriving data | A record whose master source has not loaded yet keeps its relationship. When the master arrives, its attributes complete the same record β no held-back loads, no reprocessing, no manual repair | DWA |
| Historisation | A normalised, historised core where every value carries a validity period β point-in-time questions stay answerable | DWA |
| SCD Type 2 | Slowly changing dimensions tracked through validFrom versioning fields | DWA |
| Code generation | Load logic, historisation and structures generated from the model, not written by hand | DWA |
| SQL preview | The generated loading SQL inspected before it runs | DWA |
| Target methods | TRANSACTION, APPEND, CHANGES ONLY, LATEST VERSION or OVERWRITE, per sourcefile | DLS β DWA |
| Target formats | Standard tables, Apache Iceberg, or semi-structured JSON | DLS β DWA |
| Hierarchical normalisation | Nested data flattened as a single table, split by lists, or fully normalised into a table set | DLS |
β DWA layer Β· Data warehouse automation
5. Transformation and mappingβ
The rules that connect a source field to a model attribute.
| Capability | What it means | Owner |
|---|---|---|
| Field-to-attribute mappings | The core mechanism linking sourcefile fields to model attributes | DWA |
| Key mappings | Mappings marked as business keys, with an explicit key order | DWA |
| Optional calculations | SQL expressions applied per mapping, for derived and computed values | DWA |
| Mapping filters | A WHERE condition scoped to a single mapping | DWA |
| Sort order | Sort priority per mapping, ascending or descending | DWA |
| Batch editing | Mappings edited in bulk rather than one at a time | DWA |
| Mapping diagram | A visual view of how sources reach models | DWA |
β Mappings and transformations
6. Data quality monitoringβ
Continuous measurement of whether the platform is complete, connected and healthy.
| Capability | What it means | Owner |
|---|---|---|
| Platform counters | Systems, sourcefiles, fields, mappings and models, tracked over time | DataOps Console |
| Mapping coverage | The share of a sourcefile's fields that reach a model, scored green / yellow / red | DataOps Console |
| Connection health | Connection type, whether a system has exports, and how many | DataOps Console |
| QPI checks | Quality checks created, edited, retired and triggered manually | DataOps Console |
| Quality dashboards | Architecture diagram, data lineage flow and source detail panels | DataOps Console |
β Quality monitoring (QPI) Β· Data quality pages
7. Governance, security and complianceβ
Classification and control that come out of the flow rather than sitting beside it.
| Capability | What it means | Owner |
|---|---|---|
| Sensitive-data flagging | Fields tagged as PII or otherwise sensitive | DLS |
| Field domains | Governance classification per field | DLS |
| Field exclusion | Fields removed from the pipeline entirely | DLS |
| Key field indicators | Explicit marking of primary and business keys | DLS |
| Generated constraints | Unique constraints from business keys, foreign keys from relationships, primary keys on publish tables β emitted as database objects, on by default | DWA |
| Generated comments | Table and column comments emitted inside CREATE TABLE, from the descriptions held in the model and the data contract | DWA |
| Classification tags in the target | fieldDomain and sensitive emitted as tags on the target columns, using each platform's own tagging facility | DWA |
| Column masking | Masking policy on Snowflake, dynamic data masking on the T-SQL family, SECURITY LABEL on PostgreSQL β Databricks emits a tag only, no mask | DWA |
| Classification inheritance | Domain tags follow inherited business keys onto normalised child tables β classify once, stay classified | DWA |
| Query tagging | Every statement carries its job, step and run into the target's own query history | DWA |
| Lineage | Where a column comes from and where it ends up, derived from metadata | AME / DataOps Console |
| Immutable audit trail | Raw deliveries preserved unchanged for audit and reconciliation | DLS |
| Role-based console access | admin, developer and reader roles, with every mutating action admin-only | DataOps Console |
| Action audit log | One structured log line per user action β who, role, page, action, target, timestamp | DataOps Console |
| Encryption and compression | Applied per sourcefile, at rest | DLS |
β What lands in the target environment Β· Field-level governance Β· Roles and audit log
8. Orchestration and operationsβ
Running the platform day to day, and fixing what did not run.
| Capability | What it means | Owner |
|---|---|---|
| Central orchestration | Dependencies, run order and re-runs handled centrally rather than per pipeline | AME |
| Task scheduling | A schedule per sourcefile, with its health visible | DWA |
| Pre- and post-processing tasks | Additional tasks chained onto the loads | DWA |
| Run tracing | Every delivery followed through Landing β Raw β Trusted β Published | DataOps Console |
| Publisher tracing | Confirmation that each publish landed in its destination table | DataOps Console |
| Loading statistics | Rows staged and inserted, by sourcefile and by object | DataOps Console |
| Platform health status | One indicator for the selected window: healthy, warnings, degraded or unavailable | DataOps Console |
| Corrective actions | Restart, skip and re-trigger for the work that failed | DataOps Console |
| Service health and logs | Whether the services themselves are running, and what they logged | DataOps Console |
| Time-windowed views | Every operational page scoped by one shared date range | DataOps Console |
β Daily operations Β· DataOps Console
9. Metadata and documentationβ
The metadata does not only describe the platform β it runs it.
| Capability | What it means | Owner |
|---|---|---|
| Central metadata repository | All logic, every operation that has run and every operation that is planned, in one place | AME |
| Metadata-driven execution | Every component reads its configuration from the repository | AME |
| API access to metadata | The REST API that the console and integrations read | AME |
| OpenLineage export | Start and end events for every execution, parent-job events from quality checks, and attribute-level lineage as an option β lineage lands in the catalogue you already run | AME / DWA |
| Definition export | The generated DDL returned from the SQL Generator API, so a definition can be diffed in CI before it is deployed | AME |
| Generated documentation | The full configuration and data contract behind a sourcefile, rendered from metadata | DataOps Console |
| Data model views | How the target model hangs together, drawn from the model itself | DataOps Console |
| Live architecture view | The configured flow drawn from the systems, sourcefiles and models actually in place | DataOps Console |
| Versioning and lineage as by-products | Produced by the run rather than maintained as a separate project | AME |
β Reference views Β· Getting metadata out β APIs and OpenLineage Β· Architecture
10. Portabilityβ
The platform sits on top of your stack, not the other way around.
| Capability | What it means |
|---|---|
| Any database | Snowflake, Databricks, SQL Server, Synapse, Fabric, PostgreSQL |
| Any storage | Azure Blob, Amazon S3, CEPH, Swift, MinIO |
| Any cloud | Azure, AWS, Cleura |
| Layer independence | Changing database or storage is a configuration change; changing cloud is a migration |
| Portable model and logic | The model, the logic and the metadata move with you when the underlying technology changes |
β Technology-agnostic by design
Next stepsβ
- Architecture β how the layers fit together
- Platform guide β how the layers connect, and the way into the INGEST, DLS and DWA deep dives
- What lands in the target environment β the constraints, comments, tags and masking the platform emits into your database
- Your first pipeline β the capabilities applied end to end