Known issues
Last reviewed: 2026-09-21 — against product version 3.2.
This page documents known issues in the PDQ platform and their solutions or workarounds — behaviours that are known, expected and not yet fixed, with a status where a fix is planned.
This page explains why a class of failure happens. For the step-by-step response when a daily check fails, use the operational runbook: Troubleshooting.
Infrastructure and system stability
Docker containers stop unexpectedly
Issue: Docker containers exit with code 0 or other error codes after VM restart.
Symptoms:
- Containers show as "Exited" in
docker ps -a - Services don't respond to network requests
- DataOps Console unavailable
Solution:
cd /datadrive/configs
docker compose down
docker compose up -d
Status: Planned — automatic container restart after a VM restart.
Disk space reaches 100%
Issue: /datadrive reaches 100% capacity, causing system failures.
Symptoms:
- Docker containers fail to start
- Logs show "No space left on device"
- New data files cannot be processed
Solution:
- Check disk space:
df -h - Reclaim space from unused images, containers and build cache:
docker system prune -a - Check which container logs have grown:
du -sh /datadrive/docker/containers/* - Consider expanding the disk via the cloud portal
Workaround: Implement automatic log rotation in /etc/docker/daemon.json
Caddy and authentication
FQDN resolution fails
Issue: Caddy cannot obtain SSL certificates due to DNS problems.
Symptoms:
- HTTPS connections fail
- SSL certificate errors in browser
- Caddy logs show ACME errors
Solution:
- Verify DNS configuration:
nslookup <your-fqdn> - Check that the Caddy configuration carries the correct FQDN
- Confirm the certificate issuance method matches your network exposure
Installing and securing the platform requires port 443 to be reachable inside the customer environment only, not publicly. A publicly-trusted ACME HTTP challenge cannot complete against a host the certificate authority cannot reach, so an internal-only deployment needs a DNS-based challenge or an internal CA instead. Do not open the port publicly to work around this.
Status: Under investigation — improved DNS validation.
Azure AD groups not syncing
Issue: User groups from Azure AD don't appear correctly in PDQ.
Symptoms:
- Users can log in but lack proper permissions
- "Groups claim" missing in JWT token
- Roles not mapped correctly
Solution:
- Check App Registration Token Configuration
- Ensure "Group Claims" configured for "Security Groups"
- Verify groups
pdq_web_admins,pdq_web_developers,pdq_web_viewersexist - Grant admin consent for
User.ReadandGroupMember.Read.All
Data ingestion (INGEST)
OAuth 2.0 tokens expire prematurely
Issue: Teams/SharePoint OAuth tokens expire more frequently than expected.
Symptoms:
- INGEST tasks fail with "401 Unauthorized"
- "Token expired" in INGEST logs
- Manual re-authentication required daily
Solution:
- Check token refresh logic in INGEST configuration
- Ensure
refresh_tokenis saved correctly - Implement automatic token renewal via API
Workaround: Configure longer token lifetime in Azure AD App Registration
Status: Planned — improved token handling.
Large files cause memory errors
Issue: Very large XML/JSON files (>2GB) exceed available system memory.
Symptoms:
- INGEST process crashes with "OutOfMemoryError"
- DLS Worker stops on large files
- System becomes unresponsive during processing
Solution: Enable streaming for large files:
curl -X PUT "http://localhost:8080/api/v3.2/sourcefiles/{filename}/metadata" \
-H "Content-Type: application/json" \
-d '{"hierarchyLevel": "records", "useForSplittingRecords": 1}'
Status: By design — streaming is available but must be enabled manually.
Data Lake Service (DLS)
DLS Worker stuck in "Raw" status
Issue: Files remain in Raw Archive and don't process further to Trusted.
Symptoms:
- DLS Trace shows data in Raw but not in Trusted
lastSeenOn = rawfor affected files- No error messages in logs
Solution:
- Filter DLS Trace:
lastSeenOn = raw - Open Admin Console for affected files
- Press "Load Trusted" to force processing
Root Cause: a corrupt or unreadable file. Schema deviations do not stop a delivery — they are recorded, and the file carries on to Trusted.
Inconsistent data count between zones
Issue: Significant differences in record counts between processing zones.
Symptoms:
- Landing: 1000, Trusted: 850, Published: 850
- A difference of more than 10–20 records between zones
- Data missing in downstream systems
Solution:
- Check data quality in source files
- Review validation rules in Data Modifier
- Investigate filtering logic in DLS configuration
Workaround: A difference of up to 10–20 records between Published and the other zones is expected and can be accepted. Landing, Raw Archive, Trusted and Profile should match exactly — see The daily pass.
Data Warehouse Automation (DWA)
SQL not generated for mappings
Issue: DWA generates no SQL because a mapping group has no business key mapping. This is designed behaviour, not a defect — see Silent failure modes.
Symptoms:
- "No SQL Generated" in DWA Loading Tasks
- Mappings appear complete in UI
- No error messages shown
Solution: Check mapping completeness:
- Object Load: Requires complete key mapping
- Attribute Load: Requires key mapping + at least one attribute
- Relationship Load: Requires two complete keys from same source
Common Causes:
- Case-sensitive field names don't match source data
- Missing mandatory fields in mapping
- Mappings spread across multiple mapping groups
Batch schedules run too early
Issue: DWA schedules trigger before source data is available.
Symptoms:
- "Expected" tasks > "Completed" tasks
- Tasks stuck in "Scheduled" status >2 hours
- Downstream system reports missing data
Solution:
- Identify affected tasks
- Mark all tasks for specific schedule and source file
- Use "Restart Tasks" from DataOps Console
- Adjust schedule timing for future runs
DataOps and monitoring
DataOps Console won't load
Issue: DataOps Console shows blank page or loads indefinitely.
Symptoms:
- White screen when navigating to console
- JavaScript errors in browser developer tools
- Timeout on API calls
Solution:
- Check AME API status:
curl http://localhost:8080/api/health - Verify network connectivity between UI and AME
- Clear browser cache and cookies
- Check for port mapping issues in Docker
System health checks report false status
Issue: System Health shows a degraded status for working components.
Symptoms:
- Red status for active services
- Conflicting status indicators
- False alarms in monitoring systems
Solution:
- Manual verification of component status
- Restart health check services
- Check health check configuration in AME
Status: Under investigation — improved health check logic.
Performance and scaling
Slow query response in Published data
Issue: SQL queries against Published tables take long time to execute.
Symptoms:
- Timeout in reporting tools
- High CPU usage on database server
- User complaints about slow performance
Solution:
- Analyse query patterns and add indexes
- Consider data partitioning for large tables
- Implement materialised views for common aggregations
- Optimise DWA-generated SQL
Workaround: Use batch processing for large analytical queries
Memory leak in long-running containers
Issue: Docker containers consume increasingly more memory over time.
Symptoms:
- Progressively increasing memory usage
- System becomes slow after several days
- OOM kills of containers
Solution: Schedule container restart:
# Add to crontab for weekly restart
0 2 * * 0 cd /datadrive/configs && docker compose restart
Status: Under investigation — memory optimisation.
Security
Connections may need saving again after an upgrade
Issue: After an upgrade a connection can stop authenticating, because the stored secret is not carried across. The platform never stores a secret unencrypted, so the older advice on this page — to GET the configuration and POST it back — described a problem that does not exist and a fix that breaks the connection.
Symptoms:
- INGEST tasks fail to authenticate against a source that worked before the upgrade
- The connection still looks complete in the UI
Solution:
Open each affected connection and save it again, entering the password, keys and any client_secret. A stored credential is never reused, so every save requires them afresh. See Credentials after an upgrade.
There is nothing to re-encrypt. Retrieving a connection and posting it back encrypts the already-encrypted value a second time and leaves a working connection unusable.
Reporting new issues
If you discover new issues not documented here:
-
Gather information:
- PDQ version and component versions
- Detailed symptoms and error messages
- Steps to reproduce the problem
- System configuration and environment details
-
Check logs:
- Docker container logs:
docker logs <container_name> - System logs:
/var/log/ - Application logs via DataOps Console
- Docker container logs:
-
Document workarounds: If you find temporary solutions, document them for the team
-
Report via: Support channels according to organisational procedures