DLS layer
The DLS layer manages how data flows through the storage zones β from raw landing to standardised, query-ready tables. It handles archiving, processing, profiling and publishing.
Deviations against the data contract are recorded here, not blocked. See What the platform records, and what it stops.
For the screen-by-screen guide to configuring a sourcefile, see Data lake service.
The DLS layer manages how data flows through storage zones β from raw landing to validated, query-ready tables. It handles archiving, processing, validation, and publishing.
Zone architectureβ
ββββββββββββ ββββββββββββββββ βββββββββββββββ βββββββββββββββββ
β Landing β βββΊβ Raw Archive β βββΊβ Trusted β βββΊβ Published β
β Zone β β Zone β β Zone β β Zone β
ββββββββββββ ββββββββββββββββ βββββββββββββββ βββββββββββββββββ
Landing Zone
- Purpose: Temporary staging area where exported data first arrives
- Lifecycle: Files are processed and moved to Raw Archive β not retained long-term
- Path format:
[landingZoneName]/[system]/[filename]/ - No date tokens β landing is a transient buffer
Raw Archive Zone
- Purpose: Immutable, permanent record of every file that was ingested
- Lifecycle: Write-once, never modified β serves as the audit trail
- Path format:
[rawZoneName]/[system]/[filename]/[YYYY]/[MM]/[DD]/ - Date-partitioned for efficient historical queries
Trusted Zone
- Purpose: Standardised data ready for modelling
- Lifecycle: Written by DLS Workers as each delivery is processed
- Path format:
[trustedZoneName]/[system]/[filename]/[YYYY]/[MM]/[DD]/ - Deviations against the data contract are recorded here, not blocked β see Architecture
Published Zone
- Purpose: SQL-optimised tables ready for direct querying and warehouse consumption
- Lifecycle: Synced from Trusted via the Publish step
- Controlled by:
publishedZoneNameandpublishedModelSchemain Settings
Zone configurationβ
All zone names are configurable in Settings β Data Lake Settings:
| Setting | Description | Example |
|---|---|---|
landingZoneName | Landing zone label | landing |
rawZoneName | Raw archive label | raw |
trustedZoneName | Trusted zone label | trusted |
publishedZoneName | Published zone label | published |
Storage Settings
| Setting | Description |
|---|---|
storageIntegration | Snowflake storage integration name (required for external stages) |
externalVolume | External volume for data staging |
formatFile | Format definition file (JSON/text) |
tableCatalog | Snowflake catalogue specification |
publishedModelSchema | Target schema for published models |
The DLS pipelineβ
Landing βββΊ Raw Archive βββΊ DLS Workers βββΊ Trusted + Profile
β
βΌ
Sync/Publish
β
βΌ
Published Zone
DLS Workers β parallel worker processes (minimum six containers) handle data transformation and validation:
- Parse incoming file formats (CSV, JSON, XML)
- Apply data type coercion and normalisation
- Run profiling checks when
enableProfileris active - Write data to Trusted Zone, recording any deviation against the contract
- Write profiling statistics to Profile Zone
Sync/Publish Step β bridges the gap between DLS (Trusted Zone) and DWA (Published Zone):
- Generates loading SQL for the target format (TABLE or ICEBERG)
- Applies the configured
targetMethod(TRANSACTION,APPEND,CHANGES ONLY,LATEST VERSIONorOVERWRITE) - Creates or updates tables in the Published schema
Publish patternsβ
Target Methods
The targetMethod on a sourcefile controls how data is written to the trusted zone path, and from there into the published tables:
| Method | Behaviour | When to Use |
|---|---|---|
| TRANSACTION | Appends all incoming data to the target table without any controls | Event and log data where every delivery is kept as delivered |
| APPEND | Appends changes and new records without comparing against what is already there, so duplicate change records may occur | Append-only patterns where the source has already deduplicated |
| CHANGES ONLY | Appends changes and new records, comparing against the latest known change so no duplicate change records are added | The usual choice for a changing source |
| LATEST VERSION | Applies changes and new records while preserving only one version per primary key | Master and reference data queried as it stands now |
| OVERWRITE | Incoming data overwrites the existing data | Small tables and full refreshes |
Target Formats
| Format | Description | Considerations |
|---|---|---|
| TABLE | The target databaseβs native table format | Best for most use cases, full SQL support |
| ICEBERG | Apache Iceberg table format | Better for multi-engine access, time travel, schema evolution |
| JSON | Semi-structured storage | When structure is unknown or highly variable |
SQL Generation
PDQ generates the loading SQL automatically based on your configuration:
- Sourcefile structure defines the columns and types
- Target method determines INSERT/MERGE/CREATE OR REPLACE behaviour
- Mappings control which fields map to which model attributes
- Settings provide schema, catalogue, and storage integration details
You can preview the generated SQL before execution to verify it matches your intent.
Publish considerationsβ
Data Normalisation
The targetNormalization setting controls how nested/hierarchical data is flattened:
| Mode | Behaviour | Example |
|---|---|---|
NONE | One table per sourcefile | Flat CSV β single target table |
LISTS | Split arrays into separate tables | JSON with nested arrays β parent + child tables |
LISTS AND OBJECTS | Split arrays and nested objects | Deeply nested JSON β fully normalised table set |
Path Token Support
Zone paths support date tokens for time-based partitioning:
| Token | Value |
|---|---|
[YYYY] | Four-digit year |
[MM] | Two-digit month |
[DD] | Two-digit day |
[HH] | Two-digit hour |
Example: trusted/SAP/customers/[YYYY]/[MM]/[DD]/
Landing Zone paths do not support date tokens β they use a flat path.
Compression & Encryption
| Setting | Options |
|---|---|
fileCompressionType | None, or gzip |
fileCompressionLevel | 1 (fastest) to 9 (smallest) |
enableEncryption | Enable/disable encryption at rest |
Supported file encodings: UTF-8 (default, recommended), UTF-16, Windows-1252
Best practicesβ
Zone Design
- Keep Raw Archive immutable β never modify or delete files in the Raw Zone. It's your recovery path.
- Use date partitioning β always include
[YYYY]/[MM]/[DD]in Raw and Trusted paths for efficient querying and lifecycle management. - Name zones clearly β use descriptive names that make it obvious where data is in its lifecycle.
Publishing Strategy
| Scenario | Recommended Approach |
|---|---|
| Small reference table (<100K rows) | OVERWRITE method, TABLE format |
| Large fact table (millions of rows) | APPEND method with CDC upstream |
| Slowly changing dimension | LATEST VERSION method with business keys |
| Multi-engine analytics (Spark + Snowflake) | ICEBERG format |
| Unknown/evolving schema | JSON format initially, migrate to TABLE once stable |
Performance Considerations
- Profiling adds compute cost β enable
enableProfilerfor critical data, consider disabling for high-volume, well-understood sources - Compression saves storage β use gzip level 6 as a good balance between speed and size
- Choose
LATEST VERSIONcarefully β it requires well-defined business keys and is slower thanAPPENDfor large volumes - Partition by date β it enables efficient pruning and makes cleanup/retention policies easier to implement
Data Governance
- Mark PII fields as
sensitive: 1in the sourcefile structure - Use
excludeFromProfiling: 1for fields that should not be analysed - Assign
fieldDomainclassifications for compliance tracking - Exclude internal/debug fields with
excludeField: 1to keep published data clean
Next stepsβ
- DWA layer β turning a published delivery into a model
- Data lake service β the field-by-field configuration guide
- Architecture β the model chain and its synonyms