Skip to main content

Installing and securing the platform

The PDQ platform is deployed as a modular set of Docker containers, with a strictly defined startup sequence and centralised metadata control through the Active Metadata Engine (AME). There is no PDQ-specific runtime underneath it: anything that can run containers can run the platform.

The core tools deployed are AME, INGEST, DLS and DWA, along with the Config UI and Caddy.


Where each component belongs​

Deployment is about efficiency and decentralisation where it pays:

  • INGEST agents are preferably deployed close to the data source — on-premises or in a separate cloud — to minimise data transmission.
  • DWA is ideally located in the same region as the target database.
  • AME and DLS run as close to each other as possible, in the same region as the repository database.

Deployment prerequisites and infrastructure setup​

System requirements​

MinimumRecommended
Docker Engine20.10+24.0+
RAM8 GB16 GB
CPU4 cores8 cores
Storage100 GB500 GB
DatabaseSharedDedicated instance (SQL Server, PostgreSQL or another supported engine)

Verify the host before you start:

docker --version
docker compose version

VM and core dependencies​

Deployment typically begins on a VM — an EC2 Ubuntu 22.04 instance is the reference platform — prepared as follows:

  1. Disk setup: Create and mount an extra disk (for example /datadrive), using mklabel gpt and mkfs.xfs. This disk holds the Docker data root.
  2. Docker installation: Install the Docker engine, CLI, containerd.io and the required plugins.
  3. Data root and logging: Edit or create /etc/docker/daemon.json to define the data root ("data-root": "/datadrive/docker") and configure rotating logs ("max-size": "10m", "max-file": "3").

Service orchestration and startup order​

The components are managed by systemd service files enforcing a strict, linear dependency chain through Requires= and After= directives, so a service starts only if the one before it succeeded.

The required startup order is:

  1. simplitics-ui.service (base service, requires only Docker)
  2. simplitics-ame.service (Active Metadata Engine — starts after UI)
  3. simplitics-dls.service (Data Lake Service — starts after AME)
  4. simplitics-dwa.service (Data Warehouse Automation — starts after DLS)
  5. simplitics-ingest.service (Data Importer — starts after DWA)

To start the whole stack, enable and start only the last service in the chain (simplitics-ingest.service).

1. Active Metadata Engine (AME)​

  • Security isolation: No component may publish a mapped port. Remove every port mapping from the Compose files — ame-compose.yml included — so that all services sit on the internal Docker network and only Caddy is exposed. See Network exposure below.
  • API communication: AME is the central API hub, serving all configuration and instructions to INGEST, DLS and DWA.

2. INGEST (Data Importer)​

  • Deployment model: INGEST runs as a container, one per source connection, and operates as an agent that can be deployed anywhere — it needs only an API connection back to AME.
  • Container start: Initiated with the source system name as a runtime parameter, for example command: -n <SOURCESYSTEMNAME> -f teams.yaml.
  • Teams/API integration: Requires an Azure AD app registration, client secrets, and the AME connection configured through /api/v3/ingest/connection/for/:sourceSystem with APIAuthMethod set to "OAuth 2.0".
SFTP host keys are not verified

An INGEST SFTP source auto-accepts the host key, so the server's identity is not checked. Restrict the network path to the SFTP host, and do not rely on the transport alone to tell you which machine answered. See Configuring ingestion.

3. DLS (Data Lake Service)​

  • Container scaling: DLS runs as several containers — a minimum of six — using workers that scale with available capacity.
  • Functionality deployment: Configuration (file structure, rawZonePath, targetMethod) is deployed through API calls to AME at /api/v3.2/sourcefiles/:sourceFilename.
  • Streaming configuration: High-performance streaming for large XML/JSON files is enabled by setting useForSplittingRecords to 1 on a hierarchy level, through a backend API update.

4. DWA (Data Warehouse Automation)​

  • Deployment location: DWA runs as several containers — a minimum of four — preferably in the same region as the target database.
  • Configuration: Relies on three metadata inputs, submitted through the API or the Config UI:
    1. Target model definitions (objects, keys, relations)
    2. Source description (from DLS)
    3. Source-to-target mappings (LOGIC), including attributes and filtering
  • Orchestration: DWA uses APIs to describe the source, create a runtime schedule, and log progress.

Post-installation configuration​

Once the stack is running, open Settings in the Config UI and set the installation-wide values before the first load:

  • Installation name and Data platform — the target engine the SQL generator writes for (Snowflake, Databricks, SQL Server, Synapse, Fabric or PostgreSQL).
  • Compute warehouse / workspace — the compute the platform runs against.
  • Global naming conventions — column prefix, table name casing, table name spacing and object key suffix. These are applied by the SQL generator to every generated object, so set them before generating target tables.
  • Verified file encodings — the encodings DLS accepts for incoming source files.

PDQ Settings, General: installation and platform, compute resources, global naming conventions and verified file encodings

The Dashboard confirms that the stack is up and reachable: it counts the configured systems, source files, fields, mappings and models, and draws the end-to-end flow across INGEST, DLS and DWA.

PDQ dashboard: summary tiles above the end-to-end platform flow


Security and access​

Caddy fronts the platform, handling TLS termination, API access and the login flow.

Network exposure​

Only port 443 is open, and only on Caddy. No other port is exposed, on any component.

  1. DNS: The PDQ machine must use a Fully Qualified Domain Name (FQDN).
  2. Firewall: Port 443 must be accessible on the server, but only from within the customer's environment — not publicly.
  3. Network: Create a new Docker network, for example xx_caddynet.
  4. Port mappings: Every mapped port must be removed from every component's Compose file. The components talk to each other over the internal Docker network; nothing but Caddy answers from outside.

Entra (Azure AD) authentication​

  1. App registration — create an app registration (for example sp-pdq-caddy-demo) with a redirect URI carrying the FQDN and HTTPS port: https://{dns_name}:{port}/auth/oauth2/azure/authorization-code-callback.
  2. Group claims — define three Entra groups (pdq_web_admins, pdq_web_developers, pdq_web_viewers) and update the token configuration to emit the user's assigned groups as a group claim.
  3. Permissions — grant admin consent for User.Read and GroupMember.Read.All.

Role mapping​

The installer maps roles at installation time, by setting the GUID of the Entra group as the value of each internal group — authp/admin and authp/user.

The roles span two interfaces, which the console's own role table does not make clear on its own:

RoleConfig UIDataOps Console
adminWriteEverything, including restart, skip, trigger and editing
developerWriteRead only — no admin actions
readerNo write accessRead only

A developer therefore has full write access where the platform is configured, and no ability to intervene in a running load. A reader cannot write anywhere.


Maintenance and upgrades​

  • General troubleshooting: The common resolution for infrastructure problems, such as a VM crash, is to stop and start the virtual machine from the cloud portal.
  • Container restart: If issues persist, navigate to /datadrive/configs and run docker compose down followed by docker compose up -d for all components (AME, UI, DLS, DWA, INGEST, TOOLS).
  • Upgrades: Stop all Docker services first. Update the component Compose files to point at the image tag for the release you are upgrading to — the product version, for example 3.2. Date-based numbers such as 25.6.1 are not product versions. Then perform a manual pull to fetch the new images.

Credentials after an upgrade​

After an upgrade, open each connection and save it again.

A stored credential is never reused, so passwords, keys and any client_secret have to be entered again on every save — see Credentials. There is no re-encryption routine, and there is nothing to recover: secrets are never stored unencrypted.

Do not GET a connection and POST it back

Retrieving a connection and posting it back does not re-encrypt it. It encrypts the already-encrypted value a second time and leaves a working connection unusable.

A downloaded Show JSON preview contains a live secret

The Show JSON preview renders the connection as it will be sent, secrets included. Treat a downloaded copy as a credential.