Skip to main content

Fundamentals

The Simplitics way​

PDQ is an advanced framework that integrates core architectural layers, cutting-edge principles, and automation software to accelerate development and elevate quality. Our mission is to minimize code drift over time and optimise the total cost of ownership (TCO). We are committed to democratizing data access and empowering companies to unlock business value from their data, all while upholding key principles like searchability, interoperability, quality, and security. We call it: Simplifying analytics – The Simplitics way.

Simplitics Architecture Overview

Logical data architecture and flow of data

    1. Landing: The entry point, where delivered data first arrives. A transient buffer.

    2. Raw Archive: The immutable, permanent record of every file ingested. Write-once, and the recovery path.

    3. Trusted: The delivery normalised into a standardised, consumable format. Deviations against the data contract are recorded here, not blocked.

    4. Profile: Analysis of what actually arrived, written alongside the flow rather than in it.

    5. Published: The agreed table format for the delivery, open to SQL access. Also called Bronze.

    6. Base: The ensemble integration step, where sources are merged on business keys. Also called the Ensemble Model.

    7. Core: The unified model that integrates data from various sources into common core business concepts. Also called Integrated or Silver.

    8. Business: The final layer, where data is organised into specific data products for end-user consumption. Also called DM, Data Mart or Gold.

The full chain and its synonyms are in Architecture.

Key Features

    INGEST Data: Connect to a database source or API or Message Queue or file system and copy its data to the "Landing Zone".

    Data Lake Service Listener: Configure a listener that monitors file patterns in the "Landing Zone" and moves data to the "Raw Archive". The data is uniquely tagged, further processed into desired table format, and hence published in the "Published Data" layer.

    Core Business Concept Model: Model your core business concepts and define loading mappings from "Published" through "Base" into the "Core" model.

    Additional Model Dependencies: Set up dependencies to additional logic for building the "Business" data products.

    Data Contracts: Define data contracts for your published data.

    Deviation Discovery: Discover deviations from your data contracts in day-to-day operations.

    Data Re-modelling: Re-model data on demand.

    Design as Documentation: Your design serves as your documentation, which becomes your code β€” never outdated, never out of step with what runs.

Core Components

PDQ by Simplitic core component diagram

    The AME component is the core of all operations. It stores all logic and records all operations that have occurred, along with all planned operations.

    The INGEST component is an agent technology that can run anywhere, minimizing the volume of data transmitted over the network. It can be centrally controlled even if it is distributed.

    The DLS component triggers on new file events and loads data according to selected patterns into a selected published table format. It uses workers that run on virtual capacity, requiring an appropriate amount of CPU and memory aligned with current needs. It runs streaming data SDKs on Object Store technology and can scale to as many operators as needed. Ideally, it runs many small workers.

    The DLS component also helps us be more data-driven in our development. One component profiles content and structure, allowing us to publish the results as a data contract blueprint.

    The DWA component generates SQL code and loading orchestration by simply adding model descriptions and source-to-target mappings. It has an orchestration framework that allows for additional SQL-based operations to trigger after certain models and source files have completed their loading.

    The QPI component is the framework’s data quality engine. It allows us to schedule SQL-based controls on data in DWA β€” placed on Published, Core or Business. For example, domain values must match a certain list of values, or two tables combined must match specific criteria.

    The DataOps component allows us to interact with the metadata gathered in our Active Metadata Engine. It enables slicing and dicing of runtime operators and statistics, shows which logic is in use, and allows for simple adjustments and failure resolutions.

    The Config UI component consists of two parts: a CICD pipeline framework that allows us to save logic and definitions in a Git repository, and a web interface for direct operations with the AME.

    Data can be added or discontinued at any layer in the architecture, allowing us to bypass components and replace the functionality with favoured technology.

Interested in contributing? Please reach out to us and we will coordinate the next steps.