Skip to main content

INGEST Tasks

INGEST is the orchestrated extraction service that pulls data out of source databases and APIs and delivers it to the data lake. This page audits it: did every scheduled extraction run, did it finish, and did it bring back a plausible amount of data?

The page is scoped to the sidebar date range, with a Max records slider (default 20,000) capping how much history is fetched.

Incomplete work is never hidden

The record cap is applied to Completed rows only. Failed and still-running tasks are always kept in full, so they cannot fall off the end of the limit. The same applies to the system filter and the search box — the "Failed tasks" and "Tasks still in progress" sections are always drawn from the unfiltered snapshot and are labelled as such.

Export row/file count summary​

The first block on the page, and the one that catches silent failures. Row and file counts are parsed out of each task's message field (exported N rows from source connection, or Using binary copy, we Copied N files) and aggregated per export.

Four KPI cards, each with a percentage delta against the immediately preceding window of the same length:

CardMeaning
Total ExportsExports that reported a count
Total Rows/FilesSum across all exports
Zero Count ExportsExports that reported 0 — inverse colouring, up is bad
Avg Rows/FilesMean across non-zero exports

Below the cards, a warning lists any export that returned zero rows, and a horizontal bar chart shows volume per export (zero-count bars in red). With more than 20 exports the chart switches to top and bottom 10, and detail tabs give the full list.

A task that completes successfully but returns zero rows is the classic silent failure — the state is green everywhere else on the page. This is where it shows up.

State metrics​

A row of counters over the window: Expected, Completed, Queued, Running, Scheduled, Failed. Expected is calculated from each ingester's configured schedule (start time, cycle interval, cycle type) across the selected window, not from what actually ran.

Failed tasks​

Listed when the in-progress snapshot contains failures, with the reason for each. Two bulk actions:

  • Restart — re-queues the export for its scheduled datetime
  • Skip — records the task as completed without running it, with the trace Skipped from dataops

Choose the action, confirm, and the page clears its caches and reloads.

Long running tasks​

A task that is still running when its next scheduled execution is already due is flagged as overdue — a long runner. The section lists them with their next scheduled execution time, sorted by how overdue they are. This catches a hung extraction that would otherwise sit quietly in "Running" all day.

Tasks still in progress​

Every export with at least one non-completed task, shown as a percentage bar chart by state and an "Incomplete tasks" table. This view is deliberately unfilterable.

INGEST Task overview​

The main working area, with fast filters across the top:

  • Search — matches Export or System
  • Type, State, Scheduled At multiselects

Two tabs:

  • Overview — a Gantt chart of when each task ran and how long it took, with a 1d / 7d / All zoom control and a colour legend for the task states
  • Details — the same rows as a table

Admin console​

Admin only. Applies Restart or Skip to every row currently in the filtered view, behind a confirmation popover that names the row count. Both actions are audit-logged.

INGEST Task Checklist​

The configuration-versus-reality check, and the reason a green page is not enough. It lists every configured export from the scheduling schema and marks whether it has been seen executing in the window.

  • All executed → a green confirmation
  • Any missing → a red alert

Columns include the schedule and interval, last and next execution, CDC flag, included and excluded columns, source and filter. An export with schedule never is treated as satisfied.

An export that never started produces no task, no failure and no log line. This checklist is the only place it becomes visible.