Repository navigation
[gpu][alpaka] Implementation of inference on heterogeneous architectures using alpaka - #2
Open
sanjibansg wants to merge 99 commits into
Open
sanjibansg wants to merge 99 commits into
sanjibansg wants to merge 99 commits into
Conversation
sanjibansg
force-pushed
the
gpu/alpaka
branch
from
November 23, 2025 13:42
b55f173 to
f7e44ad
Compare
…output errors for now
Co-authored-by: Saransh Chopra <saransh0701@gmail.com> Co-authored-by: Francesco Derme <francesco.derme02@gmail.com> Co-authored-by: PietroFumagalli <pfuma02@gmail.com>
Co-authored-by: Saransh Chopra <saransh0701@gmail.com> Co-authored-by: Francesco Derme <francesco.derme02@gmail.com> Co-authored-by: PietroFumagalli <pfuma02@gmail.com>
Member
Author
|
/runtest h100-47gb |
|
|
|
|
Member
Author
|
/runbenchmark h100-47gb |
|
|
|
|
Member
Author
|
/runbenchmark h100 |
|
|
|
|
…ectures (#39) * extended gpu support for selu * add gtest for selu gpu support * feat: make Selu alpha and gamma configurable via ONNX attributes * test: rename Selu gtest and model to SeluNonDefaultCoeffs
shaunlee8
added a commit
that referenced
this pull request
Aug 17, 2026
…s on heterogeneous architectures
…U operators on heterogeneous architectures
Member
Author
|
/runtest h100-47gb |
|
|
|
|
Member
Author
|
/runbenchmark h100 |
|
|
…rchitectures (#45) * dynamic tensor allocation * testing: dynamic support for Tile * test added for dynamic input * dynamic shape support for comparision op * onnx models added and regression tests for dynamic Tile and Equal * dynamic input for batchnorm op and test cases * dynamic input for conv op and tests * emit dynamic session ctor params in declaration order to match infer * dynamic input for binary, gather, transpose, concat and slice ops with tests * dynamic support for fused kernels and tests for gemm relu fusion and no bias conv * dynamic input for reduce op and test cases * add diagnostic logging * chore: remove init-fail diagnostic logging * feat: dynamic-input GPU codegen for Shape/Range/BasicNary and the particle-net op set * test: add dynamic-input GPU operator tests * chore: in line comments cleanup in the code * fix: add missing reference headers for softmax and TopK tests * feat: broadcast support in BasicNary GPU codegen * refactor: address review comments on naming, launch style and comments * refactor: modifying dynamic input codegen, share shape helpers and fixing codegen bugs * fix: guard dynamic sizes at infer and finish the run-time Range path on GPU * style: drop alpaka::math from exp and log in the Softmax kernel * refactor: pass dyn params to the gpu emitters and rename the Dim shape helpers * feat: pass the operation epilogue when registering blas configs * feat: special case of reduction operator for the last and first axis * feat: dynamic configuration of resource allocation for reduce operator * feat: avoid D-H movements by using session initialized values in range operator * feat: grid-strides to better parallelize the layernorm kernel, remove math define directives * feat: use shared memory and registers for hierarchial search of top-K * feat: support for inference on dynmainc tensors for GatherND, ScatterND, Split, Trilu, Expand and Pool * feat: minor improvments in naming conventions, and redundant code removal * feat: rename lengthExpr for RModel to iterationLengthExpr to imply iteration space range --------- Co-authored-by: Sanjiban Sengupta <sanjiban.sengupta@cern.ch>
…#40) * pad gpu kernel and test added * fix test ref values; skip redundant pad lower-bound check * feat: launch kernel as a task in PAD operator * fix: incorrect formulation for RGLRU * feat: support for inferring dynamic parameteric tensors for the pad operator on heterogeneous operators --------- Co-authored-by: Sanjiban Sengupta <sanjiban.sengupta@cern.ch>
…itectures (#52) * feat: add Alpaka support for ONNX NonZero * feat: add Alpaka support for ONNX ScatterND * test: add choice for CUDA architectures * test: integrate NonZero and ScatterND Alpaka tests * fix: address Alpaka operator review comments * feat: support for dynamic parametric inputs, remove post-op host-device trip --------- Co-authored-by: Sanjiban Sengupta <sanjiban.sengupta@cern.ch>
…neous architectures (#54) * dynamic tensor allocation * testing: dynamic support for Tile * test added for dynamic input * dynamic shape support for comparision op * onnx models added and regression tests for dynamic Tile and Equal * feat: add GPU memory optimization * feat: adjust memory expansion * chore: adjust gitignore * feat: add memory tracking in benchmark * dynamic input for batchnorm op and test cases * dynamic input for conv op and tests * emit dynamic session ctor params in declaration order to match infer * feat: add benchmark results to a file * dynamic input for binary, gather, transpose, concat and slice ops with tests * fix for conv output shape expression for stride > 1 * feat: add support for benchmarking fragmentation and peak gpu usage * dynamic support for fused kernels and tests for gemm relu fusion and no bias conv * feat: add ncu for memory tracking * Fix: Delete extract_ncu.sh * dynamic input for reduce op and test cases * add diagnostic logging * feat: add support for peak gpu tracking for ONNX and Tensorrt * fix: adjust directory creation * fix: make tensrorrt able to run with profiling * chore: ignore cached data * fix: fix cache path * feat: add memory reporting * Revert "Merge branch 'pr-45' into gpu/memory" This reverts commit b5655f1, reversing changes made to 864bfef. * fix: allow tests to take different architectures * fix: remove test log * fix: allow tests to take different architectures * fix: fix kernel launches tracking * chore: update gitignore * chore: remove build-profile files from tracking * chore: update README * fix: stop tracking .idea directory * feat: implement generic ALPAKA operator fusion * feat: support typed shuffle fusion * feat: add reorganize support to fusion planner * feat: add Where support to fusion * feat: support nonconsecutive kernel fusion * fix: fix memory release for non consecutive kernels * refactor: separate ALPAKA fusion planning * refactor: extract fusion planning helpers * refactor: extract fused ALPAKA kernel generation * feat: fuse transpose inputs into MatMul * feat: add fusion support for Clip * feat: fuse L2 normalization pattern * feat: remove redundant softmax stabilization * feat: absorb transparent casts into L2 normalization * feat: eliminate transparent tensor adapters * fix: optimize ALPAKA Concat indexing * feat: improve ALPAKA fusion planning and DAG support * feat: integrate special fusion groups into global planner * feat: add INT64 support for Reduce operators * feat: add Alpaka support for ONNX NonZero * feat: add Alpaka support for ONNX ScatterND * feat: add Alpaka support for ONNX NonZero * feat: add Alpaka support for ONNX ScatterND * test: add choice for CUDA architectures * test: integrate NonZero and ScatterND Alpaka tests * feat: support dynamic CMS models in Alpaka benchmarks * feat: add PyTorch AOTInductor benchmark backend * feat: add generic horizontal GPU kernel fusion * feat: extend Alpaka fusion to indexed tensor mappings * feat: extend Alpaka fusion to concat and slice * feat: add persistent GPU state support and MambaV2 benchmarks * chore: consolidate benchmark models and AOT exporters * fix: correct Griffin RGLRU and RWKV WKV6 recurrences * fix: remove temporary Where broadcast constants * fix: fix Mamba GPU inference correctness and performance * feat: add PyTorch AOT support for Mamba v2 * fix: handle constant integer Pow on GPU * feat: parallelize MambaScan across sequence * fix: remove LayerNormalization debug output * feat: extend generic GPU kernel fusion * fix: improve benchmark configuration and documentation * fix: address merge conflicts, dynamic parameters for sdpa * fix: add benchamrking of dynamic models with parameters --------- Co-authored-by: Harsh Chauhan <harsh2005.hc@gmail.com> Co-authored-by: Sanjiban Sengupta <sanjiban.sengupta@cern.ch>
This reorganises the repository layout and build interface of the project. It also picks up the latest ROOT upstream to sync the SOFIE changes since the last update. * feat: update standalone sofie with root upstream * fix: shape op data move from host to device using memcpy * fix: remove createBasicBinary from ROperator_BasicBinary since it is not used anywhere * feat: structure change to move test files to modular directories --------- Co-authored-by: Jonas Rembser <jonas.rembser@cern.ch> Co-authored-by: Giacomo De Pietro <giacomo.pietro@kit.edu> Co-authored-by: Lorenzo Moneta <lorenzo.moneta@cern.ch> Co-authored-by: Sainava <sainava.modak@gmail.com> Co-authored-by: Maher Ahmed Maher Shola <maher.shola.officer@gmail.com> Co-authored-by: Sargun Singh <gluonparticle@gmail.com> Co-authored-by: Aditya <adityarathore7067@gmail.com> Co-authored-by: Harsh Chauhan <harsh2005.hc@gmail.com> Co-authored-by: Ramjan Khandelwal <ramjankhandelwal7@gmail.com> Co-authored-by: oliasiri <131645734+olia110@users.noreply.github.com> Co-authored-by: markknoffler <samreedh.bhuyan@gmail.com> Co-authored-by: Neeraj <krishnaneeraj773@gmail.com> Co-authored-by: ferdymercury <ferdymercury@users.noreply.github.com> Co-authored-by: Yongyan Liu <liut19707@gmail.com> Co-authored-by: Mattias Ellert <mattias.ellert@physics.uu.se> Co-authored-by: MrViiV <mrviiv9@gmail.com>
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds initial implementation to generate GPU inference code with ALPAKA from SOFIE.