You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Proposal for Partial/Custom node ordering for SequentialRunner #3717
Support custom / partial node order where user desired.
Background
Kedro offers 3 Runners out of the box (SequentialRunner, ParallelRunner and ThreadRunner), there is another SoftFailRunner which I have implemented and can be installed in https://pypi.org/project/kedro-softfail-runner/.
In the past we have focus to make consistent support for runners and are reluctant to introduce feature parity. In fact, runners has been mostly unchanged for years. Improve resume pipeline suggestion for SequentialRunner introduce a concept similar to "Change Data Capture" (CDC), which fixed the broken suggestion and started to consider about persisted data and the closest checkpoint to recover a failed pipeline. This is the most obvious feature parity among runners as it only support SequentialRunner due to the non-deterministic nature of parallel computing.
Rethink how Kedro can play a role in multiprocessing / performance boost #3713 - I started to think more about runner recently, and my feeling is that SequentialRunner is the most important one, Kedro doesn't play an important role in terms of helping user to get code executed in a parallel fashion. This problem is usually solved by the 3rd party library, for example, polars support multi-core computing out of the box.
This proposal will only support SequentialRunner, which I am increasingly more comfortable with, details are discussed in #3713.
Design
Toposort will give ONE feasible solution, while there can be multiple possible solutions. In some case:
A -> B -> C -> D
A -> B -> D -> C
In terms of computation, both pipelines will have identical result if C & D doesn't depend on each other. However, for business logic or just ease of understanding, some ordering may be preferred. (Think about large pipeline like https://demo.kedro.org/?pipeline_id=__default__, how we perceive the execution order is usually from top-to-bottom, this is not necessary how nodes are executed)
Requirements:
Non-breaking
Validation of ordering, it cannot violate toposort result.
Nice to have:
Don't need to introduce a tons of new API
Don't need to change how Kedro resolve execution order fundamentally (keep toposort)
Feature:
Provide an argument to support custom ordering.
Next step is providing an user friendly API, as this is likely still too low-level for the end user.
High level Proposal
Add new constrains to "node_dependencies" addition to the existing inputs/outputs pair.
Possible Implementation
Can't think of anything, thus dummy outputs has been the workaround for years.
Possible Alternatives
Current workaround involves dummy inputs outputs which become tedious quickly.
Description
Support custom / partial node order where user desired.
Background
Kedro offers 3 Runners out of the box (
SequentialRunner,ParallelRunnerandThreadRunner), there is anotherSoftFailRunnerwhich I have implemented and can be installed in https://pypi.org/project/kedro-softfail-runner/.In the past we have focus to make consistent support for runners and are reluctant to introduce feature parity. In fact, runners has been mostly unchanged for years. Improve resume pipeline suggestion for SequentialRunner introduce a concept similar to "Change Data Capture" (CDC), which fixed the broken suggestion and started to consider about persisted data and the closest checkpoint to recover a failed pipeline. This is the most obvious feature parity among runners as it only support
SequentialRunnerdue to the non-deterministic nature of parallel computing.Context
TBD
SequentialRunneris the most important one, Kedro doesn't play an important role in terms of helping user to get code executed in a parallel fashion. This problem is usually solved by the 3rd party library, for example, polars support multi-core computing out of the box.This proposal will only support
SequentialRunner, which I am increasingly more comfortable with, details are discussed in #3713.Design
Toposort will give ONE feasible solution, while there can be multiple possible solutions. In some case:
In terms of computation, both pipelines will have identical result if C & D doesn't depend on each other. However, for business logic or just ease of understanding, some ordering may be preferred. (Think about large pipeline like https://demo.kedro.org/?pipeline_id=__default__, how we perceive the execution order is usually from top-to-bottom, this is not necessary how nodes are executed)
Requirements:
Nice to have:
Feature:
Next step is providing an user friendly API, as this is likely still too low-level for the end user.
High level Proposal
Add new constrains to "node_dependencies" addition to the existing inputs/outputs pair.
Possible Implementation
Can't think of anything, thus dummy outputs has been the workaround for years.
Possible Alternatives
Current workaround involves dummy inputs outputs which become tedious quickly.