Back to articles
Frameworks & Tools

RayOrch Brings Lineage-Aware Parallelism to Multimodal Data Pipelines

3 min read

Introduction

Preparing data for foundation models often involves several changes in granularity. A document may become a variable number of pages, while a video may expand into frames before each item is parsed, filtered, or converted into structured records. The number of children is not known in advance and can have a long-tail distribution. At the same time, results must remain attached to the right parent and preserve the original order. RayOrch is designed for this combination of dynamic expansion, GPU batching, and lineage-sensitive aggregation.

What RayOrch changes

  • Parent-child relations are part of the programming model. Applications declare ordered, variable-cardinality expansions and their corresponding gathers. A compiler validates the pairing, reducing the need for application code to manually regroup flat records.
  • Ready work is batched across parents. For each call, the runtime maintains a FIFO ready queue. Children that are ready at the same time can be combined across different documents or videos, improving GPU utilization without discarding their parent identity.
  • Results are reconstructed from lineage, not batch boundaries. RayOrch records child membership, immediate parents, immutable ordinals, and terminal states. The gather operation uses this metadata to rebuild outputs even when tasks finish in a different order from dispatch.
  • Progress and failures are scoped. A parent can move to the next stage once all of its required children become terminal. For typed failures, undispatched siblings belonging to the failed parent can be suppressed, while unrelated parents continue to make progress.

Reported results

On NVIDIA H20 GPUs, the paper reports a 15.14x speedup when scaling MinerU from four to 64 GPUs, and a 7.82x speedup for a video pipeline scaled from eight to 64 GPUs. In end-to-end MinerU measurements, RayOrch reduced runtime by 13.1% compared with Ray Data and 29.0% compared with Daft. On Docling, it reported a 16.0% improvement over Ray Data. These are results for the paper’s selected workloads; real-world gains will depend on operator costs, input skew, and cluster configuration.

Why it matters

The main contribution is an execution abstraction that treats lineage, dynamic parallelism, and gathering as one problem. Existing coarse-grained systems may hide useful parallelism, while flat-record interfaces often push ordering and regrouping into application code. RayOrch attempts to close that gap for document parsing, video understanding, and similar preparation pipelines.

The approach is not automatically a replacement for every batch-processing engine. Its benefits are most relevant when a workflow has clear hierarchical expansions and needs both fine-grained scheduling and deterministic reconstruction. The released implementation provides a basis for evaluating whether this model can extend to broader multimodal workloads and production fault-tolerance requirements.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Compile by Training Turns Natural-Language Specs into Local Neural Functions
Frameworks & Tools
cctest.ai

Compile by Training Turns Natural-Language Specs into Local Neural Functions

Compile by Training extends Program-as-Weights with a training-based compilation mode: teacher models synthesize examples, then a lightweight adapter specializes a compact local interpreter. The approach improves semantic accuracy on a hard benchmark subset, while increasing compile-time cost and raising questions about coverage.

Read more