Data lineage
Data lineage in Platform is in public preview. It is supported in AWS compute environments. It requires Nextflow v25.04 or later, AWS S3 object storage, and Amazon Simple Notification Service (SNS).
The feature is experimental and subject to change.
Data lineage tracks the full provenance of every pipeline run at both the task and workflow level, including what executed, what data it consumed, and what outputs it produced. Use it to audit results, verify reproducibility, and trace file provenance.
Why use data lineage
Production pipelines generate results that teams need to trust, audit, and reproduce. Data lineage provides a precise, immutable record of how each result was produced.
- Reproducibility: Every run, task, and output file receives a unique lineage ID (LID), a traversable URI that points to a structured record of what ran. Verify that two runs produced identical results, or identify where they diverged.
- Auditing and compliance: For teams in regulated industries such as pharma, clinical genomics, and contract research organizations (CROs), lineage provides the audit trail needed for regulatory compliance. Each record captures inputs, outputs, parameters, compute environment, and the user who launched the run.
- Debugging: When a cached task re-executes, or a pipeline produces an unexpected result, lineage traces backward from any output to all contributing tasks and parameters. Compare two task runs to isolate what changed.
- Broader team access: Exploring Nextflow lineage previously required CLI access and the ability to read raw JSON. Platform now shows lineage data on pipeline run detail pages and in Data Explorer.
- Pipeline output visibility: When lineage is enabled and a pipeline uses the Nextflow workflow output syntax (Nextflow 24.10.0 or later), all published output files appear in the Pipeline outputs sub-tab on the run details page. Each file entry includes its lineage ID, lineage labels, and a direct link to Data Explorer, so any team member can locate and open a result without navigating cloud storage.
- Discovery across runs: Workflow output labels make output files discoverable across runs. Navigate lineage records by label to find all matching outputs workspace-wide, without knowing which specific run produced a file.
Lineage records and event delivery
Nextflow creates a structured JSON record for each entity in your pipeline when lineage is enabled:
| Record type | Description |
|---|---|
| WorkflowRun | Full pipeline execution: repository, commit ID, parameters, compute environment, session ID, and Platform context (user, workspace, pipeline) |
| TaskRun | Individual task execution: script, code checksum, inputs, outputs, container, and dependencies |
| FileOutput | Output file: path, checksum, size, timestamp, and links back to the task and workflow that produced it |
Each record gets a lineage ID (LID), a lid:// URI that uniquely identifies the entity.
Functional flow
- Nextflow appends lineage record objects (
*.data.json) to the configured object storage bucket. - The bucket filters for objects matching
.data.jsonand publishess3:ObjectCreated:*events to an SNS topic. - The SNS topic pushes each event to a per-workspace Platform webhook over HTTPS.
- Platform verifies each delivery, buffers it, then reads the lineage object from the bucket and indexes it in the database.
- The index enriches the run details and the display of workflow-generated objects in Data Explorer.
Because delivery is a push over HTTPS, TOWER_SERVER_URL must resolve to an HTTPS endpoint that AWS can reach from the public internet. SNS refuses plain HTTP and cannot resolve a private address. An installation AWS cannot reach receives no lineage events. Nextflow still writes the records to your bucket, and Platform indexes them once delivery is established.
Lineage event ingestion moved from polling an Amazon SQS queue to SNS notifications pushed to Platform. If you are upgrading an installation that already has lineage configured, see Data lineage event ingestion moves from SQS to SNS for the permissions needed during the upgrade.
Enable data lineage
To start collecting data lineage for all pipeline runs in your workspace:
- Open Settings > Workspace settings.
- Select Lineage. If you don't see Lineage listed, contact your system administrator.
- Toggle Enable lineage by default on to collect data lineage for all pipeline runs in the workspace, or off to require per-pipeline launch configuration. Choose either a Manual or an Automatic configuration for lineage resources:
- Manual: Use your own pre-provisioned bucket and SNS topic. Define the credentials, region, bucket name, and SNS topic ARN. After saving, subscribe the webhook URL shown on the settings page to your topic. See Configure lineage manually.
- Automatic: Define the credentials and region. Platform creates the bucket, the SNS topic, the topic policies, the webhook subscription, and the bucket notification rule. This is the default setting.
- Once set and enabled, all pipeline runs in the workspace generate data lineage. See Lineage for more information about the settings.
Lineage credentials must be key-based or role-based AWS credentials. Lineage does not support workload identity federation credentials.
Updating the lineage settings after pipelines have generated lineage data results in historical data loss. The lineage index is tied to the lineage storage bucket and path. Changing it makes existing records inaccessible. To avoid data loss when updating the storage location, first copy all existing lineage data to the new bucket and path (for example, aws s3 cp --recursive s3://old-bucket/path s3://new-bucket/path), then update the workspace setting.
When launching a pipeline in a data-lineage enabled workspace, the Enable lineage toggle in the pipeline Run setup reflects the Enable lineage by default workspace setting. Turn it off to explicitly exclude data lineage for the pipeline run.
Users with the Maintain role or above can toggle lineage on or off when launching a specific pipeline run.
Additional IAM permissions required
If you use existing AWS Batch or AWS Cloud compute environments with custom IAM roles, the following service role policies are required:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListObjectsInBucket",
"Effect": "Allow",
"Action": [
"s3:ListBucket"
],
"Resource": "arn:aws:s3:::seqera-lineage-<workspace-id>"
},
{
"Sid": "AllObjectActions",
"Effect": "Allow",
"Action": "s3:*Object",
"Resource": "arn:aws:s3:::seqera-lineage-<workspace-id>/*"
},
{
"Sid": "AllowObjectTagging",
"Effect": "Allow",
"Action": [
"s3:PutObjectTagging",
"s3:GetObjectTagging"
],
"Resource": "arn:aws:s3:::seqera-lineage-<workspace-id>/*"
}
]
}
Platform integration credentials require the following additional permissions for Automatic provisioning, which creates the bucket, the notification topic, and the webhook subscription:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ManageNotificationTopics",
"Effect": "Allow",
"Action": [
"sns:CreateTopic",
"sns:SetTopicAttributes",
"sns:Subscribe",
"sns:ConfirmSubscription",
"sns:Unsubscribe",
"sns:DeleteTopic"
],
"Resource": "arn:aws:sns:*:*:seqera-lineage-*"
},
{
"Sid": "ManageLineageBuckets",
"Effect": "Allow",
"Action": [
"s3:CreateBucket",
"s3:GetBucketNotification",
"s3:PutBucketNotification",
"s3:GetObject",
"s3:ListBucket"
],
"Resource": [
"arn:aws:s3:::seqera-lineage-*",
"arn:aws:s3:::seqera-lineage-*/*"
]
}
]
}
No sqs:* permission is required. Platform holds no permission over messaging infrastructure in your account.
sns:ConfirmSubscription and s3:ListBucket are both required, and both fail quietly if omitted:
- Without
sns:ConfirmSubscription, provisioning completes and the workspace reports as configured, but Event delivery shows Failed and nothing is indexed. Platform completes the SNS handshake through the API. The permission is required even though you can also confirm a subscription by hand in a browser. - Without
s3:ListBucket, rebuilding a workspace's lineage index from its bucket fails withAccessDenied. Reindexing pages the store withListObjectsV2.
In Manual mode, Platform makes no control-plane calls other than confirming its own webhook subscription. The credentials need only:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadLineageBucket",
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:ListBucket"
],
"Resource": [
"arn:aws:s3:::<your-lineage-bucket>",
"arn:aws:s3:::<your-lineage-bucket>/*"
]
},
{
"Sid": "ConfirmLineageWebhook",
"Effect": "Allow",
"Action": [
"sns:ConfirmSubscription"
],
"Resource": "arn:aws:sns:<region>:<account>:<your-lineage-topic>"
}
]
}
Configure lineage manually
In Manual mode you own the bucket, the topic, and the subscription. Before saving the workspace settings:
-
Create the S3 bucket and the SNS topic.
-
Attach a topic access policy that allows the bucket to publish to the topic:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowBucketToPublishEvents",
"Effect": "Allow",
"Principal": { "Service": "s3.amazonaws.com" },
"Action": "sns:Publish",
"Resource": "arn:aws:sns:<region>:<account>:<your-lineage-topic>",
"Condition": {
"ArnEquals": {
"aws:SourceArn": "arn:aws:s3:::<your-lineage-bucket>"
}
}
}
]
} -
Configure a bucket notification rule that sends
s3:ObjectCreated:*events for the.data.jsonsuffix to the topic:{
"TopicConfigurations": [
{
"Id": "LineageRecordCreated",
"TopicArn": "arn:aws:sns:<region>:<account>:<your-lineage-topic>",
"Events": ["s3:ObjectCreated:*"],
"Filter": {
"Key": {
"FilterRules": [
{ "Name": "suffix", "Value": ".data.json" }
]
}
}
}
]
} -
Grant the compute environment's IAM role read/write access to the bucket, using the service role policy shown earlier.
Then save the workspace lineage settings, copy the Webhook URL shown on the settings page, and subscribe it to your topic:
aws sns subscribe \
--topic-arn arn:aws:sns:<region>:<account>:<your-lineage-topic> \
--protocol https \
--notification-endpoint '<webhook URL from the lineage settings page>'
SNS immediately posts a subscription confirmation to the endpoint. Platform verifies and confirms it with the workspace's lineage credentials. The Event delivery badge on the settings page moves from Awaiting confirmation to Active.
Set a delivery policy on your topic or subscription to widen the retry schedule. The AWS default of three attempts over roughly a minute drops events across an ordinary Platform restart. Automatically provisioned topics use a wider schedule for this reason. Tune it with TOWER_LINEAGE_SNS_MAX_RETRIES and TOWER_LINEAGE_SNS_MAX_DELAY_SECONDS.
The .data.json suffix filter is recommended to reduce cost and delivery volume, but it is not required. Platform discards any event whose object key does not end in .data.json.
Event delivery status
After you save the settings, the lineage settings page reports the Event delivery status (Active, Awaiting confirmation, Failed, or Not configured) alongside the workspace's Webhook URL. Delivery status is independent of the configuration status. A workspace can be configured and writable while Platform receives nothing.
If delivery does not become Active, confirm that AWS can reach the installation over public HTTPS and that the lineage credentials grant sns:ConfirmSubscription. Records already written to the bucket are intact, and Platform re-indexes them once delivery resumes.
Test lineage for a single pipeline or run
To test or troubleshoot data lineage for a specific pipeline, add the following to your Nextflow config file under Advanced options when adding a pipeline to the Launchpad.
lineage.enabled = true
lineage.store.location = '<PATH_TO_STORAGE>'
To test for a single pipeline run, add the same code to your Nextflow config file under Advanced options when launching the pipeline run.
If data lineage is defined for a workspace, only that data is displayed in Platform. Any unique specific pipeline or single pipeline run lineage data is only accessible via the AWS S3 console and other related services (such as Amazon Athena).
Lineage in the Platform UI
Platform shows lineage data on the run details page and in Data Explorer.
Workflow run details
For a run executed with lineage enabled, the run details page displays lineage data across the following tabs:
- Run Info: Shows the lineage ID, lineage labels, and the full Platform context captured at execution time, including user, workspace, compute environment, pipeline name, revision, and commit ID.
- Tasks: Displays the lineage ID and lineage labels for each
TaskRunalongside existing task data. You can trace any task back to its lineage record. The tab also displays all task file inputs and outputs, and the upstream and downstream tasks linked by lineage records. - Inputs: Lists all input datasets and parameters with file paths, types, and lineage IDs and lineage labels where available.
- Outputs: Lists all
FileOutputrecords linked to the workflow run, including output name, file path, type, lineage ID, and lineage labels. Files link directly to Data Explorer.
Data Explorer
Output objects from a lineage-enabled run display their LID and any lineage labels when you preview the object in Data Explorer. You can trace any file back to the pipeline run that produced it.
Search data lineage records
Global search is on by default. To hide the search bar, set TOWER_GLOBAL_SEARCH_ENABLED=false. See Configuration overview.
Use the search bar in the top navigation to find workflow runs, tasks, pipelines, and output files across every workspace you can access. To open it, select Search or press Cmd+K (macOS) or Ctrl+K (Windows and Linux). Search covers only workspaces that have data lineage enabled and in which you are a participant. Results include only records you have permission to view.
Results are ordered by most recently indexed and load as you scroll. An empty query returns the most recent records across all accessible workspaces. As you type, the field suggests keywords and, where supported, values.
Search syntax
A query is a series of space-separated tokens. Each token is either a qualifier:value pair or free text. Three rules apply to every qualifier:
- A space between tokens is AND:
type:file label:qcreturns output files that carry theqclabel. - A comma inside a value is OR:
type:workflow,taskreturns workflow runs and tasks. - Repeating a qualifier is AND:
label:qc label:validatedreturns records carrying both labels.
Qualifier names and free text are case-insensitive. Free text matches any substring of the record value. For example, salmon matches any record whose value contains salmon.
A record has exactly one type and lives in exactly one workspace. Repeating type: or workspace: returns an empty list because no record can match both values. For example, type:workflow type:file requires a record to be both a workflow run and a file. Use the comma form type:workflow,file to match either type.
Qualifiers
| Qualifier | Accepts | Description |
|---|---|---|
type: | workflow, task, file | Restrict results to a record type. Also accepts the internal names WorkflowRun, TaskRun, and FileOutput. |
label: | Any label | Records tagged with the label. Covers both Platform labels and Nextflow lineage labels. |
workspace: | organization/workspace | Scope the search to one or more workspaces by fully qualified name. |
workspaceId: | Numeric workspace ID | Numeric alias for workspace:. |
workflow: | A WorkflowRun LID | Scope the search to a single run. Results include the run itself, its tasks, and its published output files. |
pipeline: | A pipeline name | Scope the search to a pipeline name. Results include the pipeline itself and the output files in its work directory. |
task: | A TaskRun LID | Scope the search to a single task. Results include the task itself and the output files in its work directory. |
| Free text | Any string | Case-insensitive substring match on the record value. |
The field suggests workspace:, type:, and label: as you type. Enter the remaining qualifiers manually. Selecting a suggested type: or workspace: value replaces the current value for that qualifier. Selecting a suggested label: value adds another label: term to the query.
workspace: and workspaceId: set the scope of a search rather than filter its results. A query that contains only a workspace still returns that workspace's most recent records. Omit both to search every workspace available to you. Referencing a workspace you do not participate in returns an error rather than an empty list.
Renaming pipelines after execution can cause data lineage consistency issues. Pipeline names are mutable by design (can be edited). Data lineage records are immutable. If you run a pipeline, generate data lineage records, and then rename the pipeline, the indexed data lineage records are not associated with the new pipeline name.
Examples
| Query | Returns |
|---|---|
type:workflow,task | Workflow run or task records |
label:qc,validated | Records labeled qc or validated |
label:qc label:validated | Records labeled both qc and validated |
label:qc,draft label:validated | Records labeled validated and either qc or draft |
type:file salmon | Output files whose value contains salmon |
type:file multiqc pipeline:rnaseq | Output files whose value contains multiqc and associated with the rnaseq pipeline |
workspace:acme/dev label:qc | Records labeled qc in the acme/dev workspace |
workspace:acme/dev,acme/prod | Records in the acme/dev or acme/prod workspace |
workspace:acme/dev workspace:acme/prod | Nothing, because a record lives in one workspace. Use the comma form instead. |
workflow:lid://abc123 | The run lid://abc123, its tasks, and its published output files |
workflow:lid://abc123 type:task | The tasks of run lid://abc123 |
task:lid://abc123 type:file | The output files of task lid://abc123 |
Lineage search is also available through the Platform API. The GET /lineage/search endpoint accepts the same query syntax in its q parameter and returns paginated results. See the Platform API reference for the full set of lineage endpoints.
Lineage labels
Assign lineage labels to output files using the label directive in your Nextflow process definitions. Both Seqera Platform labels and Nextflow lineage labels propagate to lineage records. Seqera Platform excludes resource labels because they relate to underlying compute resources, not the data itself.
Nextflow sets lineage labels at execution time, and you cannot change them. Seqera Platform labels are mutable. Updating Platform labels after a run completes can produce a mismatch between Platform run labels and lineage labels. This is expected behavior.