Skip to content

Support for dynamic parameteric inputs in inference on heterogeneous architectures - #45

Merged
sanjibansg merged 32 commits into
ML4EP:gpu/alpakafrom
harz05:feat/dynamic-input-alpaka
Sep 15, 2026
Merged

sanjibansg merged 32 commits into
ML4EP:gpu/alpakafrom
harz05:feat/dynamic-input-alpaka

Conversation

@harz05

@harz05 harz05 commented Jun 30, 2026 •

Copy link
Copy Markdown
Member

Implements #43

Adds dynamic shape support to the GPU (alpaka) code generator: models whose input dimensions are symbolic (N, n_pf, n_sv) now generate, compile and run, instead ofhaving their shapes baked in at codegen time.

The operator set, tests and results are driven by particle-net.onnx, the target dynamic model for this work.

The dynamic buffer fix

Generating a dynamic model on the GPU path failed to compile. The dynamic intermediate tensor buffers were emitted with a type alias built on the runtime device object devAcc where a type is needed, declared under the name bufDev_while every operator referencesdeviceBuf_(so the buffer was undefined at use) and allocated into a discardedautolocal in the constructor with a hardcoded float type, so the member was never set. After fixing those, a second error showed up becausealpaka::Buf` has no default constructor, so the bare member declaration would not compile.

The fix declares the dynamic buffers as Session members with the correct Buf type and deviceBuf_ name according to the tensor dtype, default initializes each member to a one element placeholder buffer to satisfy the no default constructor requirement and allocates the real buffer in the constructor sized by the runtime length once N is known.

Operator changes:

Extended for dynamic shapes: Tile, Transpose, Concat, Gather, Slice, Reduce, Conv, BatchNormalization, BasicBinary, Comparision, Range. Each keeps a dual size_t/Dim representation mirroring the ROOT's SOFIE cpu operator with a dynamic Initialize branch that registers the output as a dynamic tensor without materializing it and index math driven by runtime dimensions.

BatchNormalization: the scale/variance fusion produces a per-channel [C] array rather than materializing weights to the full [N,C,...] tensor, which would hardcode in the batch size and block any dynamic shape. Batch and spatial dims are handled by the kernel's index math instead.

New GPU kernels for two operators that previously had none:

  • TopK": one thread per slice with a K-sized insertion-sorted register buffer; k stays a compile-time constant so the buffer is a fixed register array, while the axis length is a runtime argument.
  • Softmax: block-per-row online softmax (a single fused max/sum pass, then a shared-memory tree reduction). Threads per row is the axis length rounded up to a power of two, clamped to [32,1024], falling back to 256 for a dynamic axis.

Other operators:

  • Range: output length derived symbolically from a shape tensor instead of an opaque range_size variable that collided across ops.
  • Gather: a Gather indexing a shape tensor now produces a shape tensor whose value is written host-side and copied to the device buffer, so a shape tensor consumed by a compute kernel has its value available on device.
  • BasicNary (Max/Min/Sum/Mean): had no GPU codegen at all; the base returned "" silently, so the output was never computed on device. Added a generic elementwise kernel with per-input template types (inputs can have mixed dtypes) and per-input index decomposition for multidirectional broadcast, so differently-shaped inputs are read through their own strides instead of a shared flat index.
  • RModel_ALPAKA shape tensor declarations, and an _infer_impl argument order fix. The definition interleaves dynamic params with inputs, but the two call sites passed all params first. Those orders coincide when a single input introduces every symbol, so the bug only appears on multi-symbol models.

Kernel argument convention

Index math kernels that reference a runtime dimension receive the model shape parameters (for example N) as size_t kernel arguments, supplied at the createTaskKernel callsite. GetGPUDynParams() computes that parameter list once per operator and is called by both the kernel signature generation and the launch, so the two cannot change

Known gaps

Skipped deliberately with a comment rather than emitting wrong code:

  • Range with a fully run time size: the length expression dereferences the scalar inputs host-side, which on GPU would need a device to host read.

Tests

30 new gtests, each constructed and run at two different sizes and compared against an independent host reference:

area tests
shape / indexing Transpose, Concat, Tile, Gather, Slice, Range, RangeMul
reduce ReduceSumLast, ReduceMeanMid, ReduceMaxFirst, ReduceSumMulti
conv Conv1D, Conv1DNoBias, Conv2DNoBias
normalization BatchNorm4D, BatchNormDynSpatialRelu, BatchNorm2D
elementwise AddBroadcast, Equal, NegRelu
linear Linear (Gemm+Relu fusion)
new operators TopK, Softmax 1D/2D/3D/4D
BasicNary broadcast Max/Min/Mean/SumMultidirectionalBroadcast

Results

140/140 alpaka gtests pass on an NVIDIA H100
image

Verified end to end on particle-net.onnx, the dynamic target model: generates with 0 failures, compiles under nvcc and runs.


Memory reported by the built-in profiler:

image

Per-operator timings, N=1, n_pf=100, n_sv=10, averaged over 100 runs:

operator class count total
Conv 15 ~594 us
BatchNormalization 17 ~182 us
TopK 3 ~153 us
ReduceMean 3 ~149 us
ReduceSum 6 ~65 us
MatMul 3 ~53 us
Gemm 2 ~26 us

Conv dominates. The profiler synchronizes after every operator, so its overall figure is not a throughput measurement; the per-op breakdown is the meaningful part.

Note on multi-size sessions

A Session registers its cuBLASLt layouts at construction size, so each test here constructs a fresh Session per size. With ML4EP/sofieBLAS#11 a single Session serves multiple sizes: verified by running one ParticleNet Session at n_pf/n_sv of 100/10, 50/5 and 128/16, each returning a valid softmax.

@harz05
harz05 force-pushed the feat/dynamic-input-alpaka branch from 2946253 to 78eb28e Compare July 5, 2026 17:38
@harz05
harz05 force-pushed the feat/dynamic-input-alpaka branch from 01ddd4c to e8623a1 Compare July 8, 2026 16:15
@harz05
harz05 marked this pull request as draft July 10, 2026 18:10
@harz05
harz05 marked this pull request as ready for review July 18, 2026 09:30
@sanjibansg

Copy link
Copy Markdown
Member

/runtest h100

@github-actions

Copy link
Copy Markdown

/runtest (h100): triggered - view run

@sanjibansg

Copy link
Copy Markdown
Member

/runtest h100-47gb

@github-actions

Copy link
Copy Markdown

/runtest (h100-47gb): triggered - view run

@github-actions

Copy link
Copy Markdown

/runtest (h100-47gb): GPU Unit Tests ✅ passed - view run

@sanjibansg sanjibansg left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this very useful implementation, couple of initial comments.

Comment thread core/inc/SOFIE/ROperator_BatchNormalization.hxx
Comment thread core/inc/SOFIE/ROperator_BasicBinary.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_Reduce.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_Softmax.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_Softmax.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_Softmax.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_Softmax.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_TopK.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_TopK.hxx Outdated
Comment thread core/inc/SOFIE/ROperator_TopK.hxx Outdated
Comment thread core/src/RModel_ALPAKA.cxx Outdated
Comment thread core/src/RModel_ALPAKA.cxx Outdated
Comment thread core/src/SOFIE_common.cxx Outdated
Comment thread core/src/SOFIE_common.cxx
Comment thread core/inc/SOFIE/ROperator_Conv.hxx
Comment thread core/inc/SOFIE/ROperator_Conv.hxx
Comment thread core/inc/SOFIE/ROperator_Reduce.hxx
Comment thread core/src/RModel_ALPAKA.cxx
Comment thread core/src/RModel_ALPAKA.cxx
Comment thread core/inc/SOFIE/RModel.hxx Outdated
@sanjibansg sanjibansg changed the title Initial dynamic input support for gpu codegen Support for dynamic parameteric inputs in inference on heterogeneous architectures Sep 15, 2026
@sanjibansg
sanjibansg force-pushed the feat/dynamic-input-alpaka branch from 22e639d to 21dac6b Compare September 15, 2026 09:46
@sanjibansg
sanjibansg force-pushed the feat/dynamic-input-alpaka branch from 21dac6b to 1f8472c Compare September 15, 2026 11:54

@sanjibansg sanjibansg left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The PR has now reached the state for the first addition of dynamic parameteric tensor shapes, so its its ready to be merged, upon passing of the unit tests in CI.

@sanjibansg

Copy link
Copy Markdown
Member

/runtest h100-47gb

@github-actions

Copy link
Copy Markdown

/runtest (h100-47gb): triggered - view run

@sanjibansg

Copy link
Copy Markdown
Member

/runtest l40s

@github-actions

Copy link
Copy Markdown

/runtest (l40s): triggered - view run

@sanjibansg

Copy link
Copy Markdown
Member

/runtest h100

@github-actions

Copy link
Copy Markdown

/runtest (h100): triggered - view run

@sanjibansg

Copy link
Copy Markdown
Member

/runtest l40s

@github-actions

Copy link
Copy Markdown

/runtest (l40s): triggered - view run

@sanjibansg

Copy link
Copy Markdown
Member

/runtest h100-47gb

@github-actions

Copy link
Copy Markdown

/runtest (h100-47gb): triggered - view run

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants