Skip to main content

DLS layer

The DLS layer manages how data flows through the storage zones β€” from raw landing to standardised, query-ready tables. It handles archiving, processing, profiling and publishing.

Deviations against the data contract are recorded here, not blocked. See What the platform records, and what it stops.

For the screen-by-screen guide to configuring a sourcefile, see Data lake service.

The DLS layer manages how data flows through storage zones β€” from raw landing to validated, query-ready tables. It handles archiving, processing, validation, and publishing.

Zone architecture​

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Landing β”‚ ──►│ Raw Archive β”‚ ──►│ Trusted β”‚ ──►│ Published β”‚
β”‚ Zone β”‚ β”‚ Zone β”‚ β”‚ Zone β”‚ β”‚ Zone β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Landing Zone
  • Purpose: Temporary staging area where exported data first arrives
  • Lifecycle: Files are processed and moved to Raw Archive β€” not retained long-term
  • Path format: [landingZoneName]/[system]/[filename]/
  • No date tokens β€” landing is a transient buffer
Raw Archive Zone
  • Purpose: Immutable, permanent record of every file that was ingested
  • Lifecycle: Write-once, never modified β€” serves as the audit trail
  • Path format: [rawZoneName]/[system]/[filename]/[YYYY]/[MM]/[DD]/
  • Date-partitioned for efficient historical queries
Trusted Zone
  • Purpose: Standardised data ready for modelling
  • Lifecycle: Written by DLS Workers as each delivery is processed
  • Path format: [trustedZoneName]/[system]/[filename]/[YYYY]/[MM]/[DD]/
  • Deviations against the data contract are recorded here, not blocked β€” see Architecture
Published Zone
  • Purpose: SQL-optimised tables ready for direct querying and warehouse consumption
  • Lifecycle: Synced from Trusted via the Publish step
  • Controlled by: publishedZoneName and publishedModelSchema in Settings

Zone configuration​

All zone names are configurable in Settings β†’ Data Lake Settings:

SettingDescriptionExample
landingZoneNameLanding zone labellanding
rawZoneNameRaw archive labelraw
trustedZoneNameTrusted zone labeltrusted
publishedZoneNamePublished zone labelpublished
Storage Settings
SettingDescription
storageIntegrationSnowflake storage integration name (required for external stages)
externalVolumeExternal volume for data staging
formatFileFormat definition file (JSON/text)
tableCatalogSnowflake catalogue specification
publishedModelSchemaTarget schema for published models

The DLS pipeline​

Landing ──► Raw Archive ──► DLS Workers ──► Trusted + Profile
β”‚
β–Ό
Sync/Publish
β”‚
β–Ό
Published Zone

DLS Workers β€” parallel worker processes (minimum six containers) handle data transformation and validation:

  • Parse incoming file formats (CSV, JSON, XML)
  • Apply data type coercion and normalisation
  • Run profiling checks when enableProfiler is active
  • Write data to Trusted Zone, recording any deviation against the contract
  • Write profiling statistics to Profile Zone

Sync/Publish Step β€” bridges the gap between DLS (Trusted Zone) and DWA (Published Zone):

  • Generates loading SQL for the target format (TABLE or ICEBERG)
  • Applies the configured targetMethod (TRANSACTION, APPEND, CHANGES ONLY, LATEST VERSION or OVERWRITE)
  • Creates or updates tables in the Published schema

Publish patterns​

Target Methods

The targetMethod on a sourcefile controls how data is written to the trusted zone path, and from there into the published tables:

MethodBehaviourWhen to Use
TRANSACTIONAppends all incoming data to the target table without any controlsEvent and log data where every delivery is kept as delivered
APPENDAppends changes and new records without comparing against what is already there, so duplicate change records may occurAppend-only patterns where the source has already deduplicated
CHANGES ONLYAppends changes and new records, comparing against the latest known change so no duplicate change records are addedThe usual choice for a changing source
LATEST VERSIONApplies changes and new records while preserving only one version per primary keyMaster and reference data queried as it stands now
OVERWRITEIncoming data overwrites the existing dataSmall tables and full refreshes
Target Formats
FormatDescriptionConsiderations
TABLEThe target database’s native table formatBest for most use cases, full SQL support
ICEBERGApache Iceberg table formatBetter for multi-engine access, time travel, schema evolution
JSONSemi-structured storageWhen structure is unknown or highly variable
SQL Generation

PDQ generates the loading SQL automatically based on your configuration:

  1. Sourcefile structure defines the columns and types
  2. Target method determines INSERT/MERGE/CREATE OR REPLACE behaviour
  3. Mappings control which fields map to which model attributes
  4. Settings provide schema, catalogue, and storage integration details

You can preview the generated SQL before execution to verify it matches your intent.

Publish considerations​

Data Normalisation

The targetNormalization setting controls how nested/hierarchical data is flattened:

ModeBehaviourExample
NONEOne table per sourcefileFlat CSV β†’ single target table
LISTSSplit arrays into separate tablesJSON with nested arrays β†’ parent + child tables
LISTS AND OBJECTSSplit arrays and nested objectsDeeply nested JSON β†’ fully normalised table set
Path Token Support

Zone paths support date tokens for time-based partitioning:

TokenValue
[YYYY]Four-digit year
[MM]Two-digit month
[DD]Two-digit day
[HH]Two-digit hour

Example: trusted/SAP/customers/[YYYY]/[MM]/[DD]/

note

Landing Zone paths do not support date tokens β€” they use a flat path.

Compression & Encryption
SettingOptions
fileCompressionTypeNone, or gzip
fileCompressionLevel1 (fastest) to 9 (smallest)
enableEncryptionEnable/disable encryption at rest

Supported file encodings: UTF-8 (default, recommended), UTF-16, Windows-1252

Best practices​

Zone Design
  • Keep Raw Archive immutable β€” never modify or delete files in the Raw Zone. It's your recovery path.
  • Use date partitioning β€” always include [YYYY]/[MM]/[DD] in Raw and Trusted paths for efficient querying and lifecycle management.
  • Name zones clearly β€” use descriptive names that make it obvious where data is in its lifecycle.
Publishing Strategy
ScenarioRecommended Approach
Small reference table (<100K rows)OVERWRITE method, TABLE format
Large fact table (millions of rows)APPEND method with CDC upstream
Slowly changing dimensionLATEST VERSION method with business keys
Multi-engine analytics (Spark + Snowflake)ICEBERG format
Unknown/evolving schemaJSON format initially, migrate to TABLE once stable
Performance Considerations
  • Profiling adds compute cost β€” enable enableProfiler for critical data, consider disabling for high-volume, well-understood sources
  • Compression saves storage β€” use gzip level 6 as a good balance between speed and size
  • Choose LATEST VERSION carefully β€” it requires well-defined business keys and is slower than APPEND for large volumes
  • Partition by date β€” it enables efficient pruning and makes cleanup/retention policies easier to implement
Data Governance
  • Mark PII fields as sensitive: 1 in the sourcefile structure
  • Use excludeFromProfiling: 1 for fields that should not be analysed
  • Assign fieldDomain classifications for compliance tracking
  • Exclude internal/debug fields with excludeField: 1 to keep published data clean

Next steps​