Repository navigation
Support for dynamic parameteric inputs in inference on heterogeneous architectures - #45
Merged
Merged
Conversation
harz05
force-pushed
the
feat/dynamic-input-alpaka
branch
from
July 5, 2026 17:38
2946253 to
78eb28e
Compare
harz05
force-pushed
the
feat/dynamic-input-alpaka
branch
from
July 8, 2026 16:15
01ddd4c to
e8623a1
Compare
harz05
marked this pull request as draft
July 10, 2026 18:10
harz05
marked this pull request as ready for review
July 18, 2026 09:30
Member
|
/runtest h100 |
|
|
Member
|
/runtest h100-47gb |
|
|
|
|
sanjibansg
requested changes
Aug 19, 2026
sanjibansg
left a comment
Member
There was a problem hiding this comment.
Thanks for this very useful implementation, couple of initial comments.
harz05
commented
Aug 25, 2026
harz05
commented
Aug 25, 2026
harz05
commented
Aug 25, 2026
harz05
commented
Aug 25, 2026
harz05
commented
Aug 25, 2026
sanjibansg
force-pushed
the
feat/dynamic-input-alpaka
branch
from
September 15, 2026 09:46
22e639d to
21dac6b
Compare
…xing codegen bugs
… math define directives
…ND, Split, Trilu, Expand and Pool
…eration space range
sanjibansg
force-pushed
the
feat/dynamic-input-alpaka
branch
from
September 15, 2026 11:54
21dac6b to
1f8472c
Compare
sanjibansg
approved these changes
Sep 15, 2026
sanjibansg
left a comment
Member
There was a problem hiding this comment.
The PR has now reached the state for the first addition of dynamic parameteric tensor shapes, so its its ready to be merged, upon passing of the unit tests in CI.
Member
|
/runtest h100-47gb |
|
|
Member
|
/runtest l40s |
|
|
Member
|
/runtest h100 |
|
|
Member
|
/runtest l40s |
|
|
Member
|
/runtest h100-47gb |
|
|
This was referenced Sep 15, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #43
Adds dynamic shape support to the GPU (alpaka) code generator: models whose input dimensions are symbolic (
N,n_pf,n_sv) now generate, compile and run, instead ofhaving their shapes baked in at codegen time.The operator set, tests and results are driven by particle-net.onnx, the target dynamic model for this work.
The dynamic buffer fix
Generating a dynamic model on the GPU path failed to compile. The dynamic intermediate tensor buffers were emitted with a type alias built on the runtime device object
devAccwhere a type is needed, declared under the name bufDev_while every operator referencesdeviceBuf_(so the buffer was undefined at use) and allocated into a discardedautolocal in the constructor with a hardcoded float type, so the member was never set. After fixing those, a second error showed up becausealpaka::Buf` has no default constructor, so the bare member declaration would not compile.The fix declares the dynamic buffers as Session members with the correct Buf type and
deviceBuf_name according to the tensor dtype, default initializes each member to a one element placeholder buffer to satisfy the no default constructor requirement and allocates the real buffer in the constructor sized by the runtime length onceNis known.Operator changes:
Extended for dynamic shapes: Tile, Transpose, Concat, Gather, Slice, Reduce, Conv, BatchNormalization, BasicBinary, Comparision, Range. Each keeps a dual
size_t/Dimrepresentation mirroring the ROOT's SOFIE cpu operator with a dynamicInitializebranch that registers the output as a dynamic tensor without materializing it and index math driven by runtime dimensions.BatchNormalization: the scale/variance fusion produces a per-channel
[C]array rather than materializing weights to the full[N,C,...]tensor, which would hardcode in the batch size and block any dynamic shape. Batch and spatial dims are handled by the kernel's index math instead.New GPU kernels for two operators that previously had none:
kstays a compile-time constant so the buffer is a fixed register array, while the axis length is a runtime argument.[32,1024], falling back to 256 for a dynamic axis.Other operators:
range_sizevariable that collided across ops.Max/Min/Sum/Mean): had no GPU codegen at all; the base returned "" silently, so the output was never computed on device. Added a generic elementwise kernel with per-input template types (inputs can have mixed dtypes) and per-input index decomposition for multidirectional broadcast, so differently-shaped inputs are read through their own strides instead of a shared flat index._infer_implargument order fix. The definition interleaves dynamic params with inputs, but the two call sites passed all params first. Those orders coincide when a single input introduces every symbol, so the bug only appears on multi-symbol models.Kernel argument convention
Index math kernels that reference a runtime dimension receive the model shape parameters (for example
N) assize_tkernel arguments, supplied at thecreateTaskKernelcallsite.GetGPUDynParams()computes that parameter list once per operator and is called by both the kernel signature generation and the launch, so the two cannot changeKnown gaps
Skipped deliberately with a comment rather than emitting wrong code:
Rangewith a fully run time size: the length expression dereferences the scalar inputs host-side, which on GPU would need a device to host read.Tests
30 new gtests, each constructed and run at two different sizes and compared against an independent host reference:
Results
140/140 alpaka gtests pass on an NVIDIA H100

Verified end to end on
particle-net.onnx, the dynamic target model: generates with 0 failures, compiles under nvcc and runs.Memory reported by the built-in profiler:
Per-operator timings,
N=1, n_pf=100, n_sv=10, averaged over 100 runs:Conv dominates. The profiler synchronizes after every operator, so its overall figure is not a throughput measurement; the per-op breakdown is the meaningful part.
Note on multi-size sessions
A Session registers its cuBLASLt layouts at construction size, so each test here constructs a fresh Session per size. With ML4EP/sofieBLAS#11 a single Session serves multiple sizes: verified by running one ParticleNet Session at
n_pf/n_svof 100/10, 50/5 and 128/16, each returning a valid softmax.