Skip to main content

Known issues

Last reviewed: 2026-09-21 — against product version 3.2.

This page documents known issues in the PDQ platform and their solutions or workarounds — behaviours that are known, expected and not yet fixed, with a status where a fix is planned.

Looking for what to do right now?

This page explains why a class of failure happens. For the step-by-step response when a daily check fails, use the operational runbook: Troubleshooting.

Infrastructure and system stability​

Docker containers stop unexpectedly​

Issue: Docker containers exit with code 0 or other error codes after VM restart.

Symptoms:

  • Containers show as "Exited" in docker ps -a
  • Services don't respond to network requests
  • DataOps Console unavailable

Solution:

cd /datadrive/configs
docker compose down
docker compose up -d

Status: Planned — automatic container restart after a VM restart.


Disk space reaches 100%​

Issue: /datadrive reaches 100% capacity, causing system failures.

Symptoms:

  • Docker containers fail to start
  • Logs show "No space left on device"
  • New data files cannot be processed

Solution:

  1. Check disk space: df -h
  2. Reclaim space from unused images, containers and build cache: docker system prune -a
  3. Check which container logs have grown: du -sh /datadrive/docker/containers/*
  4. Consider expanding the disk via the cloud portal

Workaround: Implement automatic log rotation in /etc/docker/daemon.json


Caddy and authentication​

FQDN resolution fails​

Issue: Caddy cannot obtain SSL certificates due to DNS problems.

Symptoms:

  • HTTPS connections fail
  • SSL certificate errors in browser
  • Caddy logs show ACME errors

Solution:

  1. Verify DNS configuration: nslookup <your-fqdn>
  2. Check that the Caddy configuration carries the correct FQDN
  3. Confirm the certificate issuance method matches your network exposure
Port 443 stays internal

Installing and securing the platform requires port 443 to be reachable inside the customer environment only, not publicly. A publicly-trusted ACME HTTP challenge cannot complete against a host the certificate authority cannot reach, so an internal-only deployment needs a DNS-based challenge or an internal CA instead. Do not open the port publicly to work around this.

Status: Under investigation — improved DNS validation.


Azure AD groups not syncing​

Issue: User groups from Azure AD don't appear correctly in PDQ.

Symptoms:

  • Users can log in but lack proper permissions
  • "Groups claim" missing in JWT token
  • Roles not mapped correctly

Solution:

  1. Check App Registration Token Configuration
  2. Ensure "Group Claims" configured for "Security Groups"
  3. Verify groups pdq_web_admins, pdq_web_developers, pdq_web_viewers exist
  4. Grant admin consent for User.Read and GroupMember.Read.All

Data ingestion (INGEST)​

OAuth 2.0 tokens expire prematurely​

Issue: Teams/SharePoint OAuth tokens expire more frequently than expected.

Symptoms:

  • INGEST tasks fail with "401 Unauthorized"
  • "Token expired" in INGEST logs
  • Manual re-authentication required daily

Solution:

  1. Check token refresh logic in INGEST configuration
  2. Ensure refresh_token is saved correctly
  3. Implement automatic token renewal via API

Workaround: Configure longer token lifetime in Azure AD App Registration

Status: Planned — improved token handling.


Large files cause memory errors​

Issue: Very large XML/JSON files (>2GB) exceed available system memory.

Symptoms:

  • INGEST process crashes with "OutOfMemoryError"
  • DLS Worker stops on large files
  • System becomes unresponsive during processing

Solution: Enable streaming for large files:

curl -X PUT "http://localhost:8080/api/v3.2/sourcefiles/{filename}/metadata" \
-H "Content-Type: application/json" \
-d '{"hierarchyLevel": "records", "useForSplittingRecords": 1}'

Status: By design — streaming is available but must be enabled manually.


Data Lake Service (DLS)​

DLS Worker stuck in "Raw" status​

Issue: Files remain in Raw Archive and don't process further to Trusted.

Symptoms:

  • DLS Trace shows data in Raw but not in Trusted
  • lastSeenOn = raw for affected files
  • No error messages in logs

Solution:

  1. Filter DLS Trace: lastSeenOn = raw
  2. Open Admin Console for affected files
  3. Press "Load Trusted" to force processing

Root Cause: a corrupt or unreadable file. Schema deviations do not stop a delivery — they are recorded, and the file carries on to Trusted.


Inconsistent data count between zones​

Issue: Significant differences in record counts between processing zones.

Symptoms:

  • Landing: 1000, Trusted: 850, Published: 850
  • A difference of more than 10–20 records between zones
  • Data missing in downstream systems

Solution:

  1. Check data quality in source files
  2. Review validation rules in Data Modifier
  3. Investigate filtering logic in DLS configuration

Workaround: A difference of up to 10–20 records between Published and the other zones is expected and can be accepted. Landing, Raw Archive, Trusted and Profile should match exactly — see The daily pass.


Data Warehouse Automation (DWA)​

SQL not generated for mappings​

Issue: DWA generates no SQL because a mapping group has no business key mapping. This is designed behaviour, not a defect — see Silent failure modes.

Symptoms:

  • "No SQL Generated" in DWA Loading Tasks
  • Mappings appear complete in UI
  • No error messages shown

Solution: Check mapping completeness:

  • Object Load: Requires complete key mapping
  • Attribute Load: Requires key mapping + at least one attribute
  • Relationship Load: Requires two complete keys from same source

Common Causes:

  • Case-sensitive field names don't match source data
  • Missing mandatory fields in mapping
  • Mappings spread across multiple mapping groups

Batch schedules run too early​

Issue: DWA schedules trigger before source data is available.

Symptoms:

  • "Expected" tasks > "Completed" tasks
  • Tasks stuck in "Scheduled" status >2 hours
  • Downstream system reports missing data

Solution:

  1. Identify affected tasks
  2. Mark all tasks for specific schedule and source file
  3. Use "Restart Tasks" from DataOps Console
  4. Adjust schedule timing for future runs

DataOps and monitoring​

DataOps Console won't load​

Issue: DataOps Console shows blank page or loads indefinitely.

Symptoms:

  • White screen when navigating to console
  • JavaScript errors in browser developer tools
  • Timeout on API calls

Solution:

  1. Check AME API status: curl http://localhost:8080/api/health
  2. Verify network connectivity between UI and AME
  3. Clear browser cache and cookies
  4. Check for port mapping issues in Docker

System health checks report false status​

Issue: System Health shows a degraded status for working components.

Symptoms:

  • Red status for active services
  • Conflicting status indicators
  • False alarms in monitoring systems

Solution:

  1. Manual verification of component status
  2. Restart health check services
  3. Check health check configuration in AME

Status: Under investigation — improved health check logic.


Performance and scaling​

Slow query response in Published data​

Issue: SQL queries against Published tables take long time to execute.

Symptoms:

  • Timeout in reporting tools
  • High CPU usage on database server
  • User complaints about slow performance

Solution:

  1. Analyse query patterns and add indexes
  2. Consider data partitioning for large tables
  3. Implement materialised views for common aggregations
  4. Optimise DWA-generated SQL

Workaround: Use batch processing for large analytical queries


Memory leak in long-running containers​

Issue: Docker containers consume increasingly more memory over time.

Symptoms:

  • Progressively increasing memory usage
  • System becomes slow after several days
  • OOM kills of containers

Solution: Schedule container restart:

# Add to crontab for weekly restart
0 2 * * 0 cd /datadrive/configs && docker compose restart

Status: Under investigation — memory optimisation.


Security​

Connections may need saving again after an upgrade​

Issue: After an upgrade a connection can stop authenticating, because the stored secret is not carried across. The platform never stores a secret unencrypted, so the older advice on this page — to GET the configuration and POST it back — described a problem that does not exist and a fix that breaks the connection.

Symptoms:

  • INGEST tasks fail to authenticate against a source that worked before the upgrade
  • The connection still looks complete in the UI

Solution: Open each affected connection and save it again, entering the password, keys and any client_secret. A stored credential is never reused, so every save requires them afresh. See Credentials after an upgrade.

Do not GET a connection and POST it back

There is nothing to re-encrypt. Retrieving a connection and posting it back encrypts the already-encrypted value a second time and leaves a working connection unusable.


Reporting new issues​

If you discover new issues not documented here:

  1. Gather information:

    • PDQ version and component versions
    • Detailed symptoms and error messages
    • Steps to reproduce the problem
    • System configuration and environment details
  2. Check logs:

    • Docker container logs: docker logs <container_name>
    • System logs: /var/log/
    • Application logs via DataOps Console
  3. Document workarounds: If you find temporary solutions, document them for the team

  4. Report via: Support channels according to organisational procedures