Asset-based vs task-based orchestration
You can describe a pipeline through the steps it runs or through the data it produces. These are the task-centric and asset-centric views. Both help answer "what should run next?", but neither automatically proves that the resulting data is complete or correct. This lesson compares the views and traces a delayed-input example.
Two ways to describe the same pipeline
Picture ShopFlow's goal (ShopFlow — see Meet ShopFlow): the fact_sales table, built from raw.orders, which ingest_orders lands from the source database.
Task-centric (the classic model, e.g. Airflow). You describe the verbs — the steps to perform and their order. This is exactly the shopflow_daily DAG from the last lessons:
"Run
ingest_orders, thendbt_run, thenquality_check."
In this task-only definition, the orchestrator knows the steps but has not been told which tables they produce. To answer "what produces fact_sales?", it needs additional metadata or asset declarations. This is a limit of that definition, not a claim that Airflow cannot represent data assets.
Asset-centric (the newer model, pioneered by Dagster). You describe the nouns — the data objects you want to exist and what each depends on:
"
fact_salesis an asset built from thestg_ordersasset, built from theraw.ordersasset, which is loaded from the source database."
A software-defined asset (SDA) is a declarative definition of a data object — a table, a file, an ML model — together with the code that produces it. You declare the asset and its upstream assets; the orchestrator derives the DAG from those dependencies.
Notice the inversion. In the task model you write the steps (ingest_orders → dbt_run → quality_check) and ShopFlow's tables are a byproduct. In the asset model you declare the tables — fact_sales and dim_customer as assets — and the steps (and their order) are derived. The DAG still exists — it's computed from asset dependencies instead of hand-wired.
Why the asset model caught on
Declaring data objects instead of steps unlocks things that are awkward in the pure task world:
- Visible declared lineage. The graph records the dependencies you declare:
fact_salescomes fromstg_orders, which comes fromraw.orders. Undeclared external reads still need to be captured; a graph is only as complete as its definitions and integrations. - A named history for each data object. A recorded materialization answers "when did this asset's code run?" Whether its rows cover the expected interval is a separate freshness or completeness check.
- A basis for data-aware scheduling. Declared dependencies let you configure work to react to upstream updates. You still choose the schedule, sensor, or automation rules that request runs.
Materialize is the asset-world verb for "run the code that produces this asset's data," the asset analog of "run this task."
Asset-based (data-aware) scheduling triggers work when an upstream data object updates, rather than purely on a time schedule. It's the model behind Dagster assets and, increasingly, Airflow 3's assets/datasets.
Time-based vs data-aware scheduling
This is the practical heart of the shift, so make it concrete:
- Time-based: "Run
dbt_runat 2 a.m." But what ifingest_ordersis late landingraw.orders? You either buildfact_saleson stale data or pad with a guessed gap (the fragile cron pattern from lesson 8.1). - Data-aware: "Rebuild
fact_saleswheneverraw.ordershas new data." No guessing about timing — the downstream reacts to the upstream actually being ready.
Data-aware scheduling replaces a guessed time gap with an explicit update signal. In Airflow, producer tasks declare asset outputs and consumers declare an asset schedule. A failed or skipped producer does not emit the successful update used by that schedule. The producer must still validate what "ready" means for its data.
Airflow introduced dataset-based scheduling in 2.4; Airflow 3 uses the asset terminology. Data-aware scheduling is therefore not exclusive to Airflow 3 or to an asset-first framework. Dagster's software-defined assets combine the data object, its computation, and its dependencies. See the optional primary references below for the version-specific APIs.
Worked example: the producer is late
ShopFlow's daily sales table must include the complete September 5 orders partition. A partition is the slice of data for that date. Assume the producer is responsible for confirming the partition has landed completely before publishing its update.
| Time | What happened | Clock-only consumer | Consumer waiting for the validated asset update |
|---|---|---|---|
| 01:55 | Ingestion starts, then slows down | Nothing yet | Nothing yet |
| 02:00 | Only part of September 5 is loaded | Its schedule fires; without a readiness check it can build an incomplete result | No successful update exists, so it waits |
| 02:12 | Ingestion finishes but validation finds missing order IDs | A previously built result still needs correction | Validation fails; no ready event is published |
| 02:20 | Retry lands the missing rows and validation passes | Requires a retry or backfill to repair the earlier output | Update becomes eligible to trigger the consumer |
The advantage comes from connecting the consumer to a validated completion signal, not merely renaming a task as an asset. If the producer reports success at 02:00 while data is partial, the event-driven consumer can make the same mistake. A clock-scheduled workflow can also be correct if it explicitly checks readiness before transforming data.
One more trap: an update to September 4 during a backfill is not proof that September 5 is ready. Record or resolve the relevant partition and make the consumer select it explicitly. Reprocessing the same partition must remain safe, as covered in Idempotency & backfills.
Try the trace: if the producer fails at 02:12 and never retries, should the downstream table be shown as fresh? No. An absent run is operationally understandable, but the data is still late and should trigger the freshness alert chosen for this dataset.
When each model fits
Neither is universally "better" — they suit different shapes of work:
| Lean task-centric when… | Lean asset-centric when… |
|---|---|
| The pipeline is a sequence of operational actions (trigger a job, call an API, move a file) where the output isn't a neat table | The pipeline's whole purpose is producing and keeping data assets fresh (tables, files, features, models) |
| You have a large existing Airflow estate and ecosystem of operators | You want built-in lineage, freshness, and data-aware scheduling from day one |
| Steps don't map cleanly to "one task = one data object" | Work maps cleanly to "this code produces this dataset," and you value strong typing/testing of those datasets |
Airflow supports asset schedules alongside its task model, and Dagster assets are backed by operations. Think in both views: "what data must exist?" and "what steps produce it?" Choose the implementation after you can state those dependencies and the conditions that make an output ready.
Why it matters
The task view makes the work explicit; the asset view makes the intended data and dependencies explicit. Neither removes the need to check completeness, handle partitions, or alert on late data. In ShopFlow's example, the useful change was waiting for a validated update, with a recovery path when that update never arrives.
- Airflow 2.4 data-aware scheduling documents the introduction of dataset schedules.
- Airflow asset-aware scheduling documents successful producer updates and consumer schedules.
- Dagster asset dependencies explains how declared upstream assets determine execution order within a run.