Skip to main content

Asset-based vs task-based orchestration

You can describe a pipeline through the steps it runs or through the data it produces. These are the task-centric and asset-centric views. Both help answer "what should run next?", but neither automatically proves that the resulting data is complete or correct. This lesson compares the views and traces a delayed-input example.

Two ways to describe the same pipeline​

Picture ShopFlow's goal (ShopFlow — see Meet ShopFlow): the fact_sales table, built from raw.orders, which ingest_orders lands from the source database.

Task-centric (the classic model, e.g. Airflow). You describe the verbs — the steps to perform and their order. This is exactly the shopflow_daily DAG from the last lessons:

"Run ingest_orders, then dbt_run, then quality_check."

In this task-only definition, the orchestrator knows the steps but has not been told which tables they produce. To answer "what produces fact_sales?", it needs additional metadata or asset declarations. This is a limit of that definition, not a claim that Airflow cannot represent data assets.

Asset-centric (the newer model, pioneered by Dagster). You describe the nouns — the data objects you want to exist and what each depends on:

"fact_sales is an asset built from the stg_orders asset, built from the raw.orders asset, which is loaded from the source database."

A software-defined asset (SDA) is a declarative definition of a data object — a table, a file, an ML model — together with the code that produces it. You declare the asset and its upstream assets; the orchestrator derives the DAG from those dependencies.

Notice the inversion. In the task model you write the steps (ingest_orders → dbt_run → quality_check) and ShopFlow's tables are a byproduct. In the asset model you declare the tables — fact_sales and dim_customer as assets — and the steps (and their order) are derived. The DAG still exists — it's computed from asset dependencies instead of hand-wired.

Why the asset model caught on​

Declaring data objects instead of steps unlocks things that are awkward in the pure task world:

  • Visible declared lineage. The graph records the dependencies you declare: fact_sales comes from stg_orders, which comes from raw.orders. Undeclared external reads still need to be captured; a graph is only as complete as its definitions and integrations.
  • A named history for each data object. A recorded materialization answers "when did this asset's code run?" Whether its rows cover the expected interval is a separate freshness or completeness check.
  • A basis for data-aware scheduling. Declared dependencies let you configure work to react to upstream updates. You still choose the schedule, sensor, or automation rules that request runs.

Materialize is the asset-world verb for "run the code that produces this asset's data," the asset analog of "run this task."

Asset-based (data-aware) scheduling triggers work when an upstream data object updates, rather than purely on a time schedule. It's the model behind Dagster assets and, increasingly, Airflow 3's assets/datasets.

Time-based vs data-aware scheduling​

This is the practical heart of the shift, so make it concrete:

  • Time-based: "Run dbt_run at 2 a.m." But what if ingest_orders is late landing raw.orders? You either build fact_sales on stale data or pad with a guessed gap (the fragile cron pattern from lesson 8.1).
  • Data-aware: "Rebuild fact_sales whenever raw.orders has new data." No guessing about timing — the downstream reacts to the upstream actually being ready.

Data-aware scheduling replaces a guessed time gap with an explicit update signal. In Airflow, producer tasks declare asset outputs and consumers declare an asset schedule. A failed or skipped producer does not emit the successful update used by that schedule. The producer must still validate what "ready" means for its data.

Tool history, checked September 6, 2026

Airflow introduced dataset-based scheduling in 2.4; Airflow 3 uses the asset terminology. Data-aware scheduling is therefore not exclusive to Airflow 3 or to an asset-first framework. Dagster's software-defined assets combine the data object, its computation, and its dependencies. See the optional primary references below for the version-specific APIs.

Worked example: the producer is late​

ShopFlow's daily sales table must include the complete September 5 orders partition. A partition is the slice of data for that date. Assume the producer is responsible for confirming the partition has landed completely before publishing its update.

TimeWhat happenedClock-only consumerConsumer waiting for the validated asset update
01:55Ingestion starts, then slows downNothing yetNothing yet
02:00Only part of September 5 is loadedIts schedule fires; without a readiness check it can build an incomplete resultNo successful update exists, so it waits
02:12Ingestion finishes but validation finds missing order IDsA previously built result still needs correctionValidation fails; no ready event is published
02:20Retry lands the missing rows and validation passesRequires a retry or backfill to repair the earlier outputUpdate becomes eligible to trigger the consumer

The advantage comes from connecting the consumer to a validated completion signal, not merely renaming a task as an asset. If the producer reports success at 02:00 while data is partial, the event-driven consumer can make the same mistake. A clock-scheduled workflow can also be correct if it explicitly checks readiness before transforming data.

One more trap: an update to September 4 during a backfill is not proof that September 5 is ready. Record or resolve the relevant partition and make the consumer select it explicitly. Reprocessing the same partition must remain safe, as covered in Idempotency & backfills.

Try the trace: if the producer fails at 02:12 and never retries, should the downstream table be shown as fresh? No. An absent run is operationally understandable, but the data is still late and should trigger the freshness alert chosen for this dataset.

When each model fits​

Neither is universally "better" — they suit different shapes of work:

Lean task-centric when…Lean asset-centric when…
The pipeline is a sequence of operational actions (trigger a job, call an API, move a file) where the output isn't a neat tableThe pipeline's whole purpose is producing and keeping data assets fresh (tables, files, features, models)
You have a large existing Airflow estate and ecosystem of operatorsYou want built-in lineage, freshness, and data-aware scheduling from day one
Steps don't map cleanly to "one task = one data object"Work maps cleanly to "this code produces this dataset," and you value strong typing/testing of those datasets
The convergence

Airflow supports asset schedules alongside its task model, and Dagster assets are backed by operations. Think in both views: "what data must exist?" and "what steps produce it?" Choose the implementation after you can state those dependencies and the conditions that make an output ready.

Why it matters​

The task view makes the work explicit; the asset view makes the intended data and dependencies explicit. Neither removes the need to check completeness, handle partitions, or alert on late data. In ShopFlow's example, the useful change was waiting for a validated update, with a recovery path when that update never arrives.

Go deeper (optional): primary references

Next: Retries, SLAs, triggering & decoupling compute →