Skip to main content

PDQ 3.2 β€” What changes for your model

Version: 3.2 Release Date: December 1, 2025

Version numbers now track the AME API contract a component requires. This release adds two new build targets, reworks the normalisation model, emits governance from the same metadata that builds the tables, and records every quality check and every statement individually.

Read this page before upgrading, then the per-component notes for the detail: Config UI Β· Data Operations Console Β· Active Metadata Engine (AME) Β· INGEST Β· DLS Β· DWA

Release highlights​

Data pipeline​

  • Two new build targets: Databricks on Delta and Unity Catalog, and PostgreSQL β€” published, core and publish.
  • Profiling now runs after a file lands; the file is handed on instead of stopping at organised.
  • A failed audit log is retried, then parked β€” a failed load is visible, not silently absent.
  • Every quality check records its own outcome; a batch no longer reports only its last check.

Data governance​

  • Constraints, comments and classification tags come out of the same metadata that builds the tables.
  • Sensitive-data classification is emitted per platform β€” a masking policy on Snowflake, dynamic data masking on the T-SQL family, SECURITY LABEL on PostgreSQL, and SET TAGS on Databricks, which emits no mask.
  • Domain tags follow inherited business keys down the normalisation hierarchy.
  • Every statement carries its job, step and run into the target's own query history.

Data management​

  • Hierarchical sources normalise differently: one table per object, children folded into their parent.
  • Technical columns are __-prefixed and protected from casing and aliasing; one __checksum per table.
  • BOOLEAN is a supported source datatype; core business keys may be NULL again.
  • Column naming is configurable, and key order is preserved end to end.

Versioning​

Before: independent calendar tags​

Each component carried its own date-based tag β€” the generator and listener on release/v25.10.1, the worker on release/v25.6.2, the file mover on release/v25.10.4. Nothing in the tag said which backend API the component spoke, so compatibility was established by reading release notes, or by deploying and finding out.

Now: one number, tied to the AME contract​

Every component in this release carries the same number: 3.2. The number states the AME API contract the component speaks β€” the 3.2 line requires the v3 API family, with object-level generation on /api/v3.2. A component and a backend are compatible when their numbers agree, and that is now readable from the version alone.

Why it was changed​

The versioning exists to make AME compatibility answerable without reading anything. A 3.2 component will not run against a v2-only backend β€” that is now stated by the version rather than discovered at runtime, and it is why bringing the backend to v3 is the first step of any upgrade.

Where a contract can still vary inside the major version, the generator negotiates rather than assuming: the sourcefile data contract is tried at v3.2, then v3.1, then v3, with a 404 falling through to the next version and a 422 stopping immediately. The internal worker build string is separate and unchanged β€” it identifies the build, not the contract.

Build targets​

Databricks​

  • Published reads with read_files() over abfss:// and Unity Catalog volumes β€” no named stage, no COPY.
  • Delta GENERATED ALWAYS AS IDENTITY instead of sequences; MERGE instead of UPDATE…FROM.
  • VARIANT columns supported, with LATERAL VIEW over VARIANT.
  • Classification via SET TAGS. Databricks emits no mask β€” a sensitive column is tagged, not masked, so enforcement has to come from catalogue policy.
  • No enforced UNIQUE, PK or FK β€” those templates are deliberately omitted. Plan quality checks accordingly.

PostgreSQL​

  • Modelled on the SQL Server package: relational, enforced constraints, sequences, no stored procedures.
  • Source files are read through a pg_lake FOREIGN TABLE over the object store, created once in the Published definition and reused by loader and publisher.
  • JSON navigation uses native jsonb operators: ->, ->>, #>>, jsonb_path_query_first.
  • Case-insensitive column handling and quoted aliases, required when the repository itself is Postgres.

Object storage is no longer Azure-shaped​

OBJECTSTOREURL replaces AZUREBLOBACCOUNTURL and is scheme-neutral: abfss://, s3://, gs://, dbfs:/, wasbs:// and absolute paths such as /Volumes/… are used verbatim, and a bare host gets https://. PUBLISHEROBJECTSTOREURL lets the publish layer target a different scheme than Published and Base. The old name still works; if both are set, the new one wins.

The model​

Hierarchical sources normalise differently​

  • One table is created per object; a level's children are folded into their parent and grandchildren are discarded.
  • The root now behaves as expected for LISTs, and the self level is included in the output.
  • Parent paths are matched on path segments instead of a substring LIKE, so similarly-named levels no longer match each other.
  • Field keys sharing a name across different levels are de-duplicated.

Columns are predictable again​

  • Technical columns are __-prefixed and treated as protected, so casing and aliasing no longer rewrite them.
  • One __checksum per table, positioned last, at table level only β€” it was previously emitted on several levels.
  • Column sort order corrected across generated scripts; key order preserved end to end; load_date propagated through.
  • Naming is configurable β€” proper, pascal, camel, init or upper, with spacing and prefix. levelAlias prefixes nested fields; root-level fields stay unprefixed.

BOOLEAN is now a source datatype​

Mapped to BOOLEAN on Snowflake and Databricks, BIT on SQL Server, Synapse and Fabric. Sources that were previously coerced to another type will change shape β€” check any source carrying true/false values before re-generating.

Loading​

Core loading​

  • RelValidFrom is included in the day-span and incremental patterns, with null defaulting added.
  • A portable lowest-date default replaces the non-portable GREATEST β€” Fabric keeps GREATEST.
  • Per-attribute X_HasUpdate gating: an attribute is only overwritten when a source row actually landed in the delta window. Propagating NULL remains the intended "clear the value".
  • Snowflake incremental core loads use MERGE again.

Publish and schema drift​

  • The v3.2 publisher is rebuilt: schema drift handled via MERGE, primary keys on publish tables with __fileKey included, and version-based ALTER generation.
  • method=cleanup added; method=overwrite now drops the table first β€” it previously prevented table creation.
  • CREATE OR REPLACE is restricted to regular tables, not Iceberg tables.
  • Categories are added to published tables, referencing only the parent table.

Structural and decorative DDL are now separated​

The DEVELOPMENT path emits only structural column changes β€” CREATE TABLE IF NOT EXISTS and ALTER ADD/DROP COLUMN. Comments, unique and foreign-key DDL are definition-only, and duplicate executable statements are de-duplicated in the audit trace. Core business keys may also be NULL again; existing tables created with NOT NULL are not migrated automatically.

Governance​

Constraints​

  • Unique constraints from business keys β€” inline where the platform requires it, ALTER ADD elsewhere.
  • Foreign keys from relationship metadata, child β†’ parent, de-duplicated.
  • Primary keys on publish tables.
  • All three default to on.

Documented at birth​

  • Column and table comments are emitted inside CREATE TABLE on Snowflake and Databricks.
  • Staging landing tables carry fixed technical descriptions on rowId, fileName, the VARIANT column and loadedTs.
  • Apostrophes in descriptions are escaped on every comment path β€” they previously broke the statement.

Classification​

  • Category tags via TAG_NAME, default category.
  • Sensitive-data tagging via SENSITIVE_TAG_NAME and SENSITIVE_TAG_VALUE.
  • On Databricks, emitted as SET TAGS only β€” there is no generated mask.

Domain tags now follow the keys they belong to​

On a normalised child table, fieldDomain tags now also include the domains of business keys inherited from ancestor levels. Those key columns were already carried onto the child table, so their tags belong there too β€” previously only fields declared on the table's own level were tagged.

Ancestors are walked nearest-first and only fill gaps: a domain declared on the table's own level always wins, each alias is emitted once, and only key fields are inherited. Publish definitions for child tables will therefore carry additional tag statements for key columns that were previously left untagged.

Data quality​

Per-check execution records​

  • Each check is keyed by (qpi, runDttm) and records its own state and result through its own start and finish calls.
  • The batch status previously reflected only the last check in the batch.
  • runDttm is taken from the message and written back onto it, so a retried batch reuses the same run key and start/finish stay idempotent.
  • On completion the API rolls MDQpi.LastRunTm forward β€” there is no separate last-run write.

A failing test is not a broken check​

  • A check whose SQL ran but whose test case failed is recorded Completed with resultState=fail, and reported.
  • A check whose SQL raised is recorded Failed with the real error and traceback, and the batch is retried.
  • The report header reflects the check outcomes, not the retry decision.
  • validate() now raises on any failure β€” no SQL, no rows, database error β€” instead of returning a partial result.

A quality run is now readable as data rather than as a single pass/fail. You can see which checks failed, distinguish a rule that caught something from a check that was itself broken, and trust that a retry does not create a second run record. Definitions are no longer written by the worker β€” that endpoint is reserved for editing them.

Traceability​

The unit of work travels with the SQL​

Every statement the worker sends now carries the job it belongs to β€” app, worker version, kind (dwa / dm / qpi), job, step, run, queue, retry, plus the step's own object, attribute and relationship identifiers. Nothing needs configuring. A static QUERY_TAG or APP value already set is preserved, not discarded.

TargetNative facility the tag lands in
SnowflakeQUERY_TAG, as JSON
Databricksuser_agent_entry
PostgreSQLapplication_name (63-byte limit)
SQL Server / Synapse / FabricODBC APP β†’ program_name
All targetsLeading sqlcommenter comment

Why this matters beyond debugging​

A query in the target's own history can now be joined back to the pipeline run that issued it. That is the join key cost attribution has been missing: usage recorded by the platform on one side, the declaration that caused it on the other.

The comment is emitted in sqlcommenter form β€” sorted, percent-encoded key='value' pairs that database observability tooling already recognises. It leads rather than trails, because statements are split on ; / GO; / EXECUTE IMMEDIATE and a trailing comment would attach itself to the next statement.

Landing​

Profiling runs by default​

  • After a successful organise with an accepted audit log, the file is posted to the modify queue and handed on.
  • It is additionally posted to the profile queue when targetData.enableProfiler is 1 β€” the default when the field is absent. Set it to 0 to opt a source out.
  • The forwarded message copies the incoming one and overrides only what the audit log resolved; retries are dropped so the next worker starts with a fresh budget.

A failed load is now visible​

  • Audit logging in the landing zone retries five times, then pushes the message to the <zone>-parked queue and reports failure.
  • Parked messages are not retried automatically β€” monitor those queues.
  • A message with no objectUrl reports failure; it previously reported success.
  • Audit logging is now limited to the landing zone.

Event times are ordered, and objectUrl is a path​

Event times are normalised once, as early as possible: cloud times are converted to the configured local timezone and the tzinfo dropped, while times parsed from file names are already local and are left unshifted. This removes the double-conversion that produced offset times in the audit log. Timestamps now carry milliseconds, so files landing inside the same second can be ordered.

objectUrl is now a store-relative path rather than a fully-qualified URL, and the .gz suffix is applied in exactly one place. Anything reading objectUrl must prepend the host itself.

Modelling​

Mapping is visual​

  • Source fields and target attributes are rendered as a graph, now the default view mode.
  • Mappings are created and deleted by drag-and-drop, with search, subject-area filter, show only mapped and image export.
  • A mapping lines overlay in the tree views links source to target on hover.
  • An additional Tasks page manages pre- and post-processing extra processors.

Lineage and loose ends​

  • A lineage graph built from live metadata: systems β†’ source files β†’ models β†’ exports.
  • Orphaned mappings β€” whose target attribute or model was renamed or removed β€” are surfaced with a Remove button. They were previously invisible and could be neither seen nor cleaned up.
  • Field classifications are managed from Settings; Postgres is selectable as a data platform.

The key mapping dialog now fixes the problem instead of describing it​

It previously listed your mappings, stated a rule, and offered a single destructive button β€” Remove Mappings deleted your work, while the correct fix had no button at all. It now names what is missing, explains what a business key is, and lets you pick the source field and press Map and save. Removal drops to a text link that names its consequence and count, and the dialog deep-links to the specific model with a return link that restores your system, source file and group.

Impact on an existing warehouse​

ChangeWhat you do about it
Published table shapes change for hierarchical sourcesRe-generate definitions and diff them before running any load
Technical columns __-prefixed, one __checksum lastColumns generated by v25.10.1 will not line up with the new naming and ordering
Constraint generation defaults to onSet ENABLE_CORE_UNIQUE / FOREIGN_KEY / PUBL_PRIMARY_KEY to 0 to keep the previous output
Core business keys allow NULL againExisting tables created with NOT NULL are not migrated automatically
Snowflake and Databricks CREATE TABLE carry COMMENTRe-baseline any test that compares generated DDL byte-for-byte
Domain tags include inherited business keysChild publish definitions carry additional tag statements
BOOLEAN is a supported source datatypeSources previously coerced to another type will change
rundate default is now resolved per requestPass rundate explicitly if you relied on the previously frozen value

Also worth knowing​

Databricks emits no enforced UNIQUE, PK or FK β€” uniqueness has to be a quality check, not a database guarantee. Field categorisations cannot yet be deleted from the UI.

Adoption β€” do this, in this order​

  1. Bring the backend to v3. The 3.2 line requires it: /api/v3, with /api/v3.2 for object-level generation. Nothing else can move until it does.
  2. Move every worker to REDIS. SQS and Azure Storage Queue backends are removed; a deployment left on one falls through to a null queue silently.
  3. Start the modify and profile consumers. They must be running before the file mover is deployed, or messages accumulate behind it.
  4. Re-generate definitions and diff them. Compare old and new generated SQL before running a load. Table shapes and column naming have changed.
  5. Decide on constraints and classification. Constraint generation is on by default. Set TAG_NAME, SENSITIVE_TAG_NAME and SENSITIVE_TAG_VALUE to your own scheme.
  6. Update objectUrl consumers and monitor parked queues. objectUrl is a path with no host, queue:qfileready is gone, and parked messages are not retried automatically.
caution

Steps 4 and 5 are the ones most often skipped. Neither fails loudly β€” a stale definition loads into the wrong shape, and a default tag scheme quietly becomes the one you live with.