Skip to content

feat(bazel): run the DPC++ device code on NVIDIA GPUs through oneMath - #3834

Open
Alexandr-Solovev wants to merge 9 commits into
uxlfoundation:mainfrom
Alexandr-Solovev:dev/asolovev_onemath_nvidia
Open

Alexandr-Solovev wants to merge 9 commits into
uxlfoundation:mainfrom
Alexandr-Solovev:dev/asolovev_onemath_nvidia

Conversation

@Alexandr-Solovev

@Alexandr-Solovev Alexandr-Solovev commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Description

Lets oneDAL's DPC++ device code run on NVIDIA GPUs by swapping oneMKL for oneMath, which implements the same specification on top of cuBLAS, cuSOLVER, cuSPARSE and cuRAND. Motivation: today a GPU evaluation of oneDAL requires Intel hardware, which is a hard stop for anyone who wants to try the library before committing to it.

Nothing changes for the default build. oneMKL stays the default and the kernels are untouched.

Scope, up front: the library builds and links as a whole against a real oneMath install, and 44 of the 45 DPC++ examples run on it. But oneMath does not cover every domain oneDAL uses: sparse BLAS, and three of the five device RNG engines, are absent for reasons not fixable inside this PR. Those are compiled out behind ONEDAL_MATH_BACKEND_ONEMATH and throw unimplemented at the point of use rather than failing the build, so what is missing is missing at run time and says so, instead of blocking the whole configuration. Those paths are deselected in the tests and examples rather than left to fail at run time. Details in Validation below and in the README. Labelled experimental. Both build systems have the switch: --dpc_math_backend=onemath under Bazel, DPC_MATH_BACKEND=onemath under Make.

How

  • New header cpp/oneapi/dal/backend/math_backend.hpp is the single place the SYCL math library is chosen. It includes <oneapi/mkl.hpp> or <oneapi/math.hpp> depending on ONEDAL_MATH_BACKEND_ONEMATH and exposes the result as oneapi::dal::backend::math.
  • The nine existing namespace mkl = oneapi::mkl; aliases now point at that alias, so all ~150 mkl:: call sites and every kernel stay byte-identical. No rename sweep — and notably cpp/oneapi/dal/test/engine/mkl/blas.hpp defines an unrelated oneapi::dal::test::engine::mkl, which a blanket rename would have corrupted.
  • New @onemath prebuilt-libs repository (dev/bazel/deps/onemath.bzl), located through ONEMATHROOT. Like the other math backends there is nothing to download: which CUDA backends a libonemath.so contains is fixed when oneMath itself is configured, so there is no redistributable to pin.
  • New flag --dpc_math_backend={mkl,onemath} (default mkl). It is orthogonal to --backend_config, which selects the host math library. cpp/oneapi/dal/BUILD resolves it through an alias, because dal_module walks its dependency lists at macro-evaluation time where a select is still opaque.
  • ONEDAL_MATH_BACKEND_ONEMATH is carried in the defines of the @onemath cc_library rather than set in the .bazelrc, so the define cannot drift from the library actually being linked.
  • Retargeting device code is purely a toolchain concern — oneDAL's device sources use no Intel SYCL extensions and query sub-group sizes at run time — so -fsycl-targets is read from ONEDAL_SYCL_TARGETS and appended to both the DPC++ compile and link flags. It has to be an env var rather than a build flag because the DPC++ flag sets are baked in when the toolchain repository is configured, before any build flag is visible.
  • --config=nvidia-gpu bundles the flag, plus --test_env=LD_LIBRARY_PATH on the test lane; dev/bazel/README.md documents the oneMath cmake recipe and the limitations below.

Three Bazel-only fixes were needed to make the configuration runnable and not just linkable — the Make side already records an rpath:

  • The SONAME had to be staged. oneMath installs libonemath.so as a symlink to libonemath.so.<abi>. Staging only the symlink links fine, since ld follows it, but records a DT_NEEDED on the SONAME that is then absent from the runfiles, so every test binary died at startup with libonemath.so.0: cannot open shared object file. The library patterns are globbed now, the way mkl.bzl globs its shared objects; Bazel treats the versioned file as a runtime-only input, so naming both is exactly right.
  • Globs were silently dropped from optional_libs. _create_optional_symlinks asked repo_ctx.path(...).exists about the pattern itself, which is never true, so a globbed optional entry was skipped even when the package shipped matching files. A globbed entry is now resolved to concrete names before the presence test. Pre-existing latent bug; it only bites a caller that globs optional libraries, which @onemath is the first to do.
  • LD_LIBRARY_PATH has to reach the test. oneMath's dispatcher dlopens its per-domain backends through a $ORIGIN RUNPATH, and under Bazel $ORIGIN is the _solib directory the dispatcher is staged into, not the oneMath install — so the backends are not beside it and every test aborted with oneapi::math::backend_not_found. test:nvidia-gpu forwards the variable; the README says to put $ONEMATHROOT/lib on it.

Deselection. Rather than letting the uncovered paths build and then throw, they are deselected the way the reference host backend deselects what it does not implement:

  • te::device_sparse_blas_supported() (in test/engine/fixtures.hpp) is false under ONEDAL_MATH_BACKEND_ONEMATH; the CSR-only cases in kmeans and the logloss primitive check it with SKIP_IF. The define is set only for device translation units, so the host versions of the same tests keep running.
  • Type lists that mix a dense and a sparse method — kmeans_types, log_reg_types — drop the sparse method under that define, and the device RNG test's engine lists drop the three engines oneMath does not declare. A COMBINE_TYPES list cannot be empty, which is why the wholly-sparse cases use SKIP_IF instead.
  • //cpp/oneapi/dal/backend/primitives/sparse_blas:tests is target_compatible_with @platforms//:incompatible under --dpc_math_backend=onemath, so Bazel skips it instead of building a test whose every assertion would throw. This needed target_compatible_with threaded through dal_test, where **kwargs goes to dal_module rather than to the test rule.
  • The DPC++ examples take -DONEMATH_BACKEND=ON, which excludes kmeans_lloyd_csr_batch, and -DCUDA_BACKEND=ON, which additionally excludes pca_svd_dense_batch. There is no Bazel equivalent of the second: which backends a libonemath.so dispatches to is fixed when oneMath is configured and is not visible to the build graph.

The same configuration through Make:

export ONEMATHROOT=/path/to/onemath
export ONEDAL_SYCL_TARGETS=nvptx64-nvidia-cuda
make -f makefile oneapi_dpc PLAT=lnx32e DPC_MATH_BACKEND=onemath
  • DPC_MATH_BACKEND mirrors --dpc_math_backend and is likewise independent of BACKEND_CONFIG, which names the host math library. mkl needs no file of its own because deps.mkl.mk already carries the oneMKL device link line, so only dev/make/deps.dpc.onemath.mk is new.
  • That file replaces only the device-side pieces. The oneMath include directory is appended rather than assigned, because the oneAPI sources still include DAAL headers that need whatever the host backend's include path is; <oneapi/math.hpp> and <oneapi/mkl.hpp> do not collide and math_backend.hpp includes exactly one. Only the dispatching libonemath.so is a link input; the per-domain backends are dlopened. $ONEMATHROOT/lib is also recorded as an rpath on libonedal_dpc.so, which is what makes the result usable rather than merely linkable: ld does not consult LD_LIBRARY_PATH when resolving a dependency of a shared library, so without it every consumer's link fails with libonemath.so.0 ... not found followed by every oneapi::math symbol reported undefined, even though oneDAL itself linked fine. An absolute rpath is warranted here in a way it would not be for a redistributable — oneMath has no redistributable and no standard prefix, so it cannot be reached through the release tree's $ORIGIN-relative rpath the way oneMKL is.
  • ONEDAL_SYCL_TARGETS becomes -fsycl-targets on both the DPC++ compile and the link in compiler_definitions/dpcpp.mk. Make takes it as a make variable as well as an env var; under Bazel it has to be an env var for the toolchain reason above.
  • DPC_MATH_BACKEND=onemath builds into __work_onemath instead of __work. The incremental build only re-examines a command line when a makefile is newer than the object, so one shared directory would mix objects compiled against two different math libraries and then fail at link with undefined oneapi::mkl:: or oneapi::math:: symbols. The default path keeps the directory name it has always had.
  • Refused rather than guessed: a missing or non-oneMath ONEMATHROOT, and any PLAT other than lnx32e, where the import-library naming and the NVPTX toolchain setup are unverified.

Three source changes were needed beyond the aliasing, each found only by building against the real library:

  • mkl::blas::gemm → mkl::blas::column_major::gemm (5 call sites in primitives/blas/{gemm,gemv,syrk}_dpc.cpp). oneMKL declares column_major as an inline namespace, oneMath declares it plainly, so the unqualified spelling resolves only against oneMKL. The qualified spelling is correct in both and is what the specification actually names.
  • oneapi::mkl::rng:: → mkl::rng:: (16 sites in primitives/rng/device_engine{.hpp,_dpc.cpp}). These were fully qualified and so escaped the alias sweep entirely.
  • allow_empty = True on the header glob in onemath.tpl.BUILD. oneMath ships .hpp and .hxx but no .h, and Bazel's glob fails an unmatched pattern by default.

Also drops the unused #include <mkl_version.h> from sparse_blas/misc.hpp. It is the only reference to that header anywhere in cpp/, the file contains no version guards, and it is one of the things a genuine oneMath install cannot satisfy.

Validation

oneMath was built from source (commit 3273ca2, mklcpu + mklgpu backends) and oneDAL built against that install on Intel GPU hardware, through both build systems. Per domain, identical either way:

domain result
BLAS compiles and links after the column_major qualification
LAPACK compiles and links unchanged
RNG philox4x32x10 and mrg32k3a work; mt2203, mt19937 and mcg59 throw unimplemented — absent from oneMath
sparse BLAS throws unimplemented — disjoint API generation

The two gaps, since they bound what this PR delivers:

  • RNG. oneapi/math/rng/engines.hpp declares philox4x32x10 and mrg32k3a and nothing else. primitives/rng/device_engine.hpp also instantiates mt2203, mt19937 and mcg59. No aliasing reaches types that do not exist; this needs the engines contributed upstream, or a fallback for those three — and mt2203 has no GPU skip_ahead, so it cannot simply be remapped onto a counter-based engine.
  • Sparse BLAS. The two libraries expose different generations of the interface. oneDAL uses oneMKL's handle API (init_matrix_handle / set_csr_data / gemv / gemm / optimize_gemv); oneMath ships the newer specification (init_csr_matrix / release_sparse_matrix, and spmv / spmm driven by descriptors with separate buffer-size, optimize and execute stages). oneMKL 2026 does not offer the newer form, so no single source satisfies both — the sparse_blas primitives need a per-backend implementation, which is a separate change.

Three further translation units fail in this configuration (reduction_rm_cw_perf_dpc.cpp, reduction_rm_rw_perf_dpc.cpp, stat/test/perf_cov_dpc.cpp) — verified to fail identically on the default oneMKL backend at this base, so they are pre-existing and not attributed here.

Through Make specifically: all eight blas and lapack device objects compile, and their undefined symbols are oneapi::math::blas / oneapi::math::lapack with no oneapi::mkl left — every one of which resolves in libonemath.so. rng and sparse_blas were where the two blockers showed up first, with exactly the diagnostics you would expect (no type named 'mt2203' in namespace 'oneapi::math::rng', no member named 'gemv' in namespace 'oneapi::math::sparse'); see Whole-library build below for how they are handled now.

Default (oneMKL) path re-verified after the source edits above: all 7 blas / rng / sparse_blas device test targets pass, so the column_major and alias changes are behaviour-neutral where they matter today. On the Make side the default is byte-identical — same WORKDIR, same command line — and the same objects rebuild to oneapi::mkl::blas.

Whole-library build. The measurements above cover the primitives; a full make onedal_dpc PLAT=lnx32e DPC_MATH_BACKEND=onemath was then run to see what the library as a whole does. It surfaced one gap that had nothing to do with the two documented blockers: pca/backend/gpu/misc.hpp spelled the syevd template arguments mkl::job::vec / mkl::uplo::upper unqualified, which from inside oneapi::dal::pca::backend finds oneapi::mkl by enclosing-namespace lookup and so bypassed the backend alias entirely. Qualified as pr::mkl:: — the spelling the linear regression GPU kernels already use — five PCA GPU kernels now compile. A sweep for other such sites found none; everything else outside primitives either qualifies with pr:: or is a comment.

After that fix, eight translation units still failed, all downstream of the two gaps: primitives/rng/device_engine_dpc, primitives/sparse_blas/{gemm,gemv,set_csr_data}_dpc, detail/sparse_matrix_handle_impl, and the three decision forest GPU training kernels, which instantiate the missing engines through device_engine.hpp. Those eight are now compiled out behind ONEDAL_MATH_BACKEND_ONEMATH and throw unimplemented at the point of use, so the library builds and links as a whole:

  • make onedal_dpc PLAT=lnx32e DPC_MATH_BACKEND=onemath — 0 errors. libonedal_dpc.so has no undefined oneapi::mkl symbols and 16 undefined oneapi::math ones, all resolving in libonemath.so.0 through the new rpath.
  • All 45 DPC++ examples build against that install and 44 run correctly on an Intel Data Center GPU Max 1100 (level_zero), with PCA output identical to the oneMKL build. The passing set includes all six decision forest examples, kmeans_init_dense, kmeans_lloyd_dense_batch, dbscan, three hdbscan, three knn, svm, logistic regression, linear regression, basic statistics, covariance and the table examples.
  • The one failure is kmeans_lloyd_csr_batch, which throws exactly oneapi::dal::v1::unimplemented: Sparse BLAS is not provided by the math backend this build of oneDAL was linked against....

Bazel test suites, run rather than built. After the three runtime fixes above, bazel test --config=nvidia-gpu over primitives/{blas,lapack,rng,objective_function} and algo/{pca,kmeans,decision_forest,covariance,linear_regression,basic_statistics,logistic_regression} gives 67 passing targets, the three sparse_blas targets SKIPPED as incompatible, and decision_forest:test_spmd_dpc failing — the one failure is the pre-existing GPU segfault that this machine also hits on the default backend at main. This is the first run of the oneMath configuration against the test suites rather than the examples, and it is what the deselection above was written against. (A same-set control run on the default backend is not quotable: the card wedged partway through and every _dpc target timed out, including blas:test_gemm_dpc, which had passed in 5.8 s minutes earlier. sycl-ls hanging is the tell.)

Throwing rather than substituting is deliberate for the RNG: the absent engines have no numerically equivalent stand-in, so remapping them onto a counter-based engine would return different numbers under the name the caller asked for. The wrappers for the three stay constructible rather than being removed, because device_engine always builds a DAAL host engine and a device engine as a pair and the host-side entry points — shuffle, uniform_without_replacement, the Fisher-Yates helper — never touch the device one; only a device-side generate on an absent engine throws.

That is also why the example run is less restricted than the domain table suggests. decision_forest is the only oneAPI descriptor exposing engine_type; its default is philox4x32x10, which oneMath provides, so its one device-side draw runs and only an explicit request for mt2203/mt19937/mcg59 throws. kmeans_init draws through the Fisher-Yates helper, i.e. the DAAL host engine, so it is unaffected by the engine set entirely. Sparse BLAS is entered only for a csr_table on the device, so it bites exactly the CSR-input algorithms.

Default path, end to end. A full make onedal_dpc PLAT=lnx32e at this branch builds clean (0 errors, libonedal_dpc.so staged), and the DPC++ PCA examples built against that install with CMake and run to correct output — pca_cov_dense_batch (which goes through the syevd call site touched above), pca_svd_dense_batch, pca_cor_dense_online. So neither the -fsycl-targets plumbing, the WORKDIR keying, nor the PCA requalification regresses the path that ships.

What the CUDA backends cover

The two gaps above are interface gaps and show up on any oneMath build. On top of them, oneMath's CUDA backends do not implement everything oneMath declares. The table below is derived from the unimplemented throw sites under src/{blas,lapack,rng,sparse_blas}/backends/cu* at oneMath 3273ca2, for exactly the entry points oneDAL's device code calls. It is read from the source, not measured — there is no NVIDIA hardware here, which is why the feature stays experimental.

oneDAL uses CUDA backend covered?
blas::column_major::{gemm,gemv,syrk}, float and double cuBLAS yes — only bfloat16 gemm and the row_major forms are unimplemented
lapack::{potrf,potrs,syevd} and their scratchpad sizes cuSOLVER yes, with no restriction
lapack::gesvd cuSOLVER only for m ≥ n
rng::uniform over philox4x32x10 / mrg32k3a, USM API, float / double / std::int32_t cuRAND yes, including scalar skip_ahead — but see below
sparse::* cuSPARSE n/a — blocked by the interface-generation gap, not by the backend

Two consequences, neither visible from the domain table:

  • PCA method::svd cannot run on NVIDIA. train_kernel_svd_impl_dpc.cpp calls pr::gesvd<somevec, novec>(queue, column_count, row_count, ...) — m = columns, n = rows — so m < n for any dataset with more rows than columns, i.e. the normal case, and cuSOLVER throws unimplemented("lapack", "gesvd", "cusolver gesvd does not support m < n"). PCA method::cov goes through syevd and is unaffected, so PCA as a whole is usable and only the SVD method is not. This is the -DCUDA_BACKEND=ON example exclusion.
  • Random forest works, on its default engine. decision_forest is the only descriptor exposing engine_type, its default is philox4x32x10, and cuRAND implements everything its one device-side draw needs; mrg32k3a is the other valid explicit choice. One caveat on repeated draws: oneDAL advances an engine with skip_ahead_gpu(count) after each device draw, expecting oneMKL's relative advance, whereas cuRAND's backend implements skip_ahead(n) as curandSetGeneratorOffset(engine_, n), an absolute offset. Single-draw use is unaffected, but it is a real semantic divergence.

These, plus a cuRAND buffer-API uniform<float> overload that appears to call curandGenerateNormalDouble (oneDAL uses the USM API, so it is not on our path), are worth reporting upstream to oneMath rather than worked around here.

Known limitations (also in the README)

  1. oneapi::math is reached only for BLAS and LAPACK. The sparse_blas primitives and three of the five device RNG engines throw unimplemented under onemath, per the table above — so the library builds, links and runs, but CSR input on the device and an explicitly requested mt2203/mt19937/mcg59 do not. Both gaps live behind ONEDAL_MATH_BACKEND_ONEMATH, so the default oneMKL build compiles exactly the code it did before.
  2. oneMath's own kernels were exercised through its mklgpu backend on an Intel GPU. The NVPTX target and the CUDA backends are not validated — that needs NVIDIA hardware, which we have none of. This is the main reason the feature is experimental.
  3. Make covers PLAT=lnx32e only. Windows and the non-x86 platforms error out rather than link against a guessed oneMath library name.
  4. No CI coverage, for the same hardware reason.
  5. Not every NVIDIA GPU reports aspect::fp64; on those, add --test_disable_fp64=yes. That covers the tests, not the algorithms.

Checklist:

Completeness and readability

  • I have commented my code, particularly in hard-to-understand areas.
  • I have updated the documentation to reflect the changes or created a separate PR with updates and provided its number in the description, if necessary.
  • Git commit message contains an appropriate signed-off-by string (see CONTRIBUTING.md for details).
  • I have resolved any merge conflicts that might occur with the base branch.

Testing

  • I have run it locally and tested the changes extensively.
  • All CI jobs are green or I have provided justification why they aren't.
  • I have extended testing suite if new functionality was introduced in this PR. — no new tests: the existing device test suites are the coverage, they are simply run against a second math library. New tests would need NVIDIA hardware in CI.

Performance

  • I have measured performance for affected algorithms using scikit-learn_bench and provided at least a summary table with measured data, if performance change is expected. — no performance change is expected or possible on the default path: it compiles to the same code.
  • I have provided justification why performance and/or quality metrics have changed or why changes are not expected.
  • I have extended the benchmarking suite and provided a corresponding scikit-learn_bench PR if new measurable functionality was introduced in this PR. — n/a.

🤖 Generated with Claude Code

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Current oneMath APIs are incompatible with several call sites, and required runtime backend libraries are not propagated into Bazel runfiles.

Review effort: Balanced
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

Adds experimental NVIDIA GPU support for Bazel DPC++ builds by selecting oneMath instead of oneMKL.

Changes:

  • Adds a configurable SYCL math-backend abstraction.
  • Adds oneMath dependency and NVIDIA/NVPTX toolchain configuration.
  • Updates device primitives to use the selected backend namespace.
File Description
MODULE.bazel Registers the local oneMath repository.
.bazelrc Adds the backend flag and NVIDIA configuration.
dev/​bazel/​README.md Documents setup and limitations.
dev/​bazel/​config/​config.tpl.BUILD Defines the math-backend setting.
dev/​bazel/​deps/​onemath.bzl Configures the oneMath installation.
dev/​bazel/​deps/​onemath.tpl.BUILD Declares oneMath libraries and headers.
dev/​bazel/​toolchains/​cc_toolchain.bzl Tracks the SYCL target environment variable.
dev/​bazel/​toolchains/​cc_toolchain_lnx.bzl Applies SYCL target flags.
cpp/​oneapi/​dal/​BUILD Selects the DPC++ math dependency.
cpp/​oneapi/​dal/​backend/​math_backend.hpp Introduces the backend namespace abstraction.
cpp/​oneapi/​dal/​detail/​sparse_matrix_handle_impl.hpp Uses the selected sparse backend.
cpp/​oneapi/​dal/​backend/​primitives/​sparse_blas/​misc.hpp Redirects sparse-BLAS types.
cpp/​oneapi/​dal/​backend/​primitives/​rng/​device_engine.hpp Redirects RNG declarations.
cpp/​oneapi/​dal/​backend/​primitives/​rng/​device_engine_dpc.cpp Uses the backend header for RNG.
cpp/​oneapi/​dal/​backend/​primitives/​lapack/​syevd.hpp Redirects SYEVD declarations.
cpp/​oneapi/​dal/​backend/​primitives/​lapack/​syevd_dpc.cpp Uses the backend header for SYEVD.
cpp/​oneapi/​dal/​backend/​primitives/​lapack/​misc.hpp Redirects LAPACK utility types.
cpp/​oneapi/​dal/​backend/​primitives/​lapack/​gesvd.hpp Redirects GESVD declarations.
cpp/​oneapi/​dal/​backend/​primitives/​lapack/​gesvd_dpc.cpp Uses the backend header for GESVD.
cpp/​oneapi/​dal/​backend/​primitives/​blas/​syrk.hpp Redirects SYRK declarations.
cpp/​oneapi/​dal/​backend/​primitives/​blas/​syrk_dpc.cpp Uses the backend header for SYRK.
cpp/​oneapi/​dal/​backend/​primitives/​blas/​misc.hpp Redirects BLAS utility types.
cpp/​oneapi/​dal/​backend/​primitives/​blas/​gemv_dpc.cpp Uses the backend header for GEMV.
cpp/​oneapi/​dal/​backend/​primitives/​blas/​gemm_dpc.cpp Uses the backend header for GEMM.
cpp/​oneapi/​dal/​algo/​pca/​backend/​gpu/​train_kernel_svd_impl_dpc.cpp Redirects PCA SVD math calls.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +44 to +45
#ifdef ONEDAL_MATH_BACKEND_ONEMATH
#include <oneapi/math.hpp>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed by the later commits on this branch (0d1edce … 7fcdc29), which validate against a real oneMath build (commit 3273ca2, mklcpu/mklgpu backends) rather than a shim. Under ONEDAL_MATH_BACKEND_ONEMATH, device_engine exposes only philox4x32x10 and mrg32k3a. The sparse BLAS primitives and sparse_matrix_handle_impl are guarded and throw unimplemented, and the tests that reach them are deselected via te::device_sparse_blas_supported(). No fully qualified oneapi::mkl call is left outside math_backend.hpp. dev/bazel/README.md records the supported surface (BLAS, LAPACK, 2 of 5 RNG engines, no sparse BLAS).

Comment on lines +28 to +31
deps = [
":headers",
"@opencl//:opencl_binary",
],

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, it was unreferenced. But wiring it as runfiles would not help: the dispatcher dlopens its backends through a $ORIGIN RUNPATH, and under Bazel $ORIGIN is the _solib directory, not the place runfiles would put them. So the backends are found through LD_LIBRARY_PATH (--config=nvidia-gpu passes it through, as the README documents). I removed the dead filegroup and the optional_libs staging that only fed it (e2c92df). The oneMath blas:tests still pass on an Intel GPU.

Comment thread dev/bazel/deps/onemath.bzl Outdated
Comment on lines +47 to +49
"lib/libonemath_blas_mklgpu.so",
"lib/libonemath_lapack_mklgpu.so",
"lib/libonemath_rng_mklgpu.so",

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moot after e2c92df: nothing is staged from the backend list anymore, because the backends are found through LD_LIBRARY_PATH (see the reply above). So libonemath_sparse_blas_* is no more missing than any other backend. Sparse BLAS is unsupported under oneMath anyway, for the interface reasons documented in the README.

Alexandr-Solovev and others added 8 commits October 1, 2026 11:13
oneDAL's device primitives call oneMKL directly, which pins the GPU
backend to Intel hardware. oneMath implements the same specification on
top of cuBLAS, cuSOLVER, cuSPARSE and cuRAND, so swapping the library is
enough to let people evaluate oneDAL without an Intel GPU.

Route the nine existing `namespace mkl = ...` aliases through a single
new header, `dal/backend/math_backend.hpp`, which picks
`<oneapi/mkl.hpp>` or `<oneapi/math.hpp>` on `ONEDAL_MATH_BACKEND_ONEMATH`
and exposes the choice as `oneapi::dal::backend::math`. The ~150 `mkl::`
call sites and every kernel stay byte-identical; the decision is made
once, at build time.

The library comes from a new `@onemath` prebuilt-libs repository
(`ONEMATHROOT`), selected by `--dpc_math_backend=onemath`, which is
orthogonal to `--backend_config` (the host math library). The define
travels with the cc_library so it cannot drift from the library actually
linked. Retargeting device code is a toolchain concern, so
`-fsycl-targets` is read from `ONEDAL_SYCL_TARGETS` and added to both the
DPC++ compile and link flags. `--config=nvidia-gpu` bundles the flag, and
dev/bazel/README.md documents the oneMath configure recipe and the known
limitations.

Also drop the unused `<mkl_version.h>` include from sparse_blas/misc.hpp:
it is the only reference to that header in cpp/, the file has no version
guards, and it is the one thing left that a real oneMath install cannot
satisfy.

Validation on Intel GPU hardware: the default oneMKL path is unchanged
and its BLAS/LAPACK/RNG/sparse-BLAS device tests pass. The oneMath path
was exercised with oneMKL re-exported as `oneapi::math`, which compiles
every DPC++ source, links, and passes the same six test suites. oneMath's
own kernels and the NVPTX target still need CUDA hardware.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Built oneMath from source (mklcpu + mklgpu backends) and pointed
ONEMATHROOT at it. Three things only a real install could show:

The header glob required `include/**/*.h` to match. oneMath ships no `.h`
at all -- 122 `.hpp` and 25 `.hxx` -- so the glob failed outright and the
`.hxx` files the headers include were not staged.

oneMKL declares `inline namespace column_major`, oneMath declares it
plainly, so `mkl::blas::{gemm,gemv,syrk}` resolves only against oneMKL.
Naming `column_major` explicitly is what both libraries accept, and for
oneMKL it names the same functions it already resolved to.

The RNG sources reach for `oneapi::mkl::rng::` directly rather than
through the alias, so they never followed the backend switch. Routed
through it.

BLAS and LAPACK now compile and link against real oneMath. RNG and sparse
BLAS do not, for reasons no amount of aliasing fixes; both are recorded in
the README and the PR description.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Bazel could already put oneMath behind the DPC++ primitives; Make could
not, which meant the NVIDIA configuration was unavailable to anyone
building the way the product is actually released. Add the same switch:

  export ONEMATHROOT=/path/to/onemath
  export ONEDAL_SYCL_TARGETS=nvptx64-nvidia-cuda
  make -f makefile oneapi_dpc PLAT=lnx32e DPC_MATH_BACKEND=onemath

DPC_MATH_BACKEND mirrors Bazel's --dpc_math_backend and is likewise
independent of BACKEND_CONFIG, which names the host math library; `mkl`
needs no file of its own because deps.mkl.mk already carries the oneMKL
device link line. dev/make/deps.dpc.onemath.mk replaces only the
device-side pieces: the oneMath include directory is appended rather than
assigned, since the oneAPI sources still include DAAL headers that need
the host backend's include path, and only the run-time dispatching
libonemath.so is linked because the per-domain backends are dlopened.

ONEDAL_SYCL_TARGETS becomes -fsycl-targets on both the DPC++ compile and
the link, since device code is produced at both.

The non-default backend gets its own __work_onemath directory. The
incremental build only re-examines a command line when a makefile is
newer than the object, so a shared directory would silently mix objects
compiled against two different math libraries and then fail at link.
The default path keeps the directory name it has always had.

Refused rather than guessed: a missing or non-oneMath ONEMATHROOT, and
any PLAT other than lnx32e, where the import-library naming and the NVPTX
toolchain setup are unverified.

Measured against the oneMath install the Bazel side was validated with
(commit 3273ca2, mklcpu + mklgpu). Through Make, all eight blas and
lapack device objects compile and their undefined symbols are
oneapi::math::blas / oneapi::math::lapack with no oneapi::mkl left, and
every one of them resolves in libonemath.so. rng and sparse_blas fail
with exactly the two blockers already documented for Bazel -- engines
absent from oneapi::math::rng, and a different sparse generation. The
default backend is byte-identical: same WORKDIR, same command line, and
the same objects rebuild to oneapi::mkl::blas.

Signed-off-by: Alexandr-Solovev <alexandr.solovev@intel.com>

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A whole-library build -- the first one attempted, the earlier measurements
covered the primitives only -- showed five PCA GPU kernels failing that
have nothing to do with the two documented gaps.

`pca/backend/gpu/misc.hpp` spells the syevd template arguments `mkl::job::vec`
and `mkl::uplo::upper` unqualified. From inside `oneapi::dal::pca::backend`
that finds `oneapi::mkl` by enclosing-namespace lookup, so it silently bypassed
the backend alias and there is no `oneapi::mkl` in a oneMath build. Qualifying
it as `pr::mkl::` matches what the linear regression GPU kernels already do and
resolves to the same oneMKL entities as before by default.

The sweep for other such sites found only this one: everything else outside
`primitives` either qualifies with `pr::` already or is a comment.

The README's per-domain table was read as "only those primitives fail". Records
what a full `make onedal_dpc DPC_MATH_BACKEND=onemath` actually leaves broken --
eight translation units, including the three decision forest GPU kernels that
instantiate the missing RNG engines -- and that `libonedal_dpc.so` consequently
does not link, so no DPC++ example or test runs against oneMath yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Before this, `DPC_MATH_BACKEND=onemath` compiled most of the library but eight
translation units failed outright, so `libonedal_dpc.so` never linked and the
configuration could only be checked domain by domain. The failures were all in
two places -- the device RNG engines oneMath does not declare, and the sparse
BLAS interface generation it does not ship -- and none of them were
device-specific, so they blocked an Intel GPU build just as much as an NVIDIA
one.

Both gaps are now compiled out behind ONEDAL_MATH_BACKEND_ONEMATH and throw
`unimplemented` at the point of use. Throwing rather than substituting matters
for the RNG: the missing engines have no numerically equivalent stand-in, so
remapping them onto a counter-based engine would return different numbers under
the name the caller asked for.

The RNG wrappers stay constructible instead of being removed, because
`device_engine` always builds a DAAL host engine and a device engine as a pair,
and the host-side entry points -- `shuffle`, `uniform_without_replacement` and
the Fisher-Yates helper -- never touch the device one. Keeping the object lets
those keep working with any engine type; only a device-side `generate` on an
absent engine throws.

Linking also needed an rpath on the oneMath install: `ld` does not consult
LD_LIBRARY_PATH when resolving a dependency of a shared library, so without it
every consumer's link failed with `libonemath.so.0 ... not found` followed by
every `oneapi::math` symbol reported undefined, even though oneDAL itself had
linked.

Measured with oneMath built from source with the mklcpu and mklgpu backends, on
an Intel Data Center GPU Max 1100: the library builds with no errors, all 45
DPC++ examples build, and 44 run with the same output as the oneMKL build. The
remaining one is `kmeans_lloyd_csr_batch`, which throws out of sparse BLAS.

The default oneMKL build compiles exactly the code it did before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…annot do

Makes a --dpc_math_backend=onemath build testable rather than merely
buildable, and deselects the paths the backend cannot run instead of
letting them build and then throw.

Two runtime problems had to be fixed first, both Bazel-only -- the Make
build already records an rpath:

* oneMath installs libonemath.so as a symlink to the SONAME'd
  libonemath.so.<abi>. Staging only the symlink links fine, but records a
  DT_NEEDED on the SONAME that is absent from the runfiles, so every
  binary died at startup with `libonemath.so.0: cannot open shared object
  file`. The library patterns are globbed now, the way mkl.bzl globs its
  shared objects, which stages the versioned file as a runtime input.

* Globs were silently dropped from `optional_libs`, because the presence
  test asked `exists` about the pattern itself. A globbed entry is now
  resolved to the concrete names before that test.

* oneMath's dispatcher dlopens its per-domain backends through a $ORIGIN
  RUNPATH, and under Bazel $ORIGIN is the _solib directory the dispatcher
  is staged into rather than the install, so tests aborted with
  `backend_not_found`. `test:nvidia-gpu` forwards LD_LIBRARY_PATH.

With that, the affected suites run and pass against oneMath: BLAS,
LAPACK, RNG, objective function and logistic regression.

Deselection, following the reference host backend's pattern:

* te::device_sparse_blas_supported() is false under
  ONEDAL_MATH_BACKEND_ONEMATH; the CSR-only cases in kmeans and the
  logloss primitive check it with SKIP_IF. The define is set only for
  device translation units, so the host versions keep running.
* Type lists that mix a dense and a sparse method -- kmeans_types,
  log_reg_types -- drop the sparse method under that define, and the
  device RNG test's engine lists drop the three engines oneMath does not
  declare.
* The sparse_blas test suite is target_compatible_with
  @platforms//:incompatible under --dpc_math_backend=onemath, so Bazel
  skips it. This needed target_compatible_with threaded through
  dal_test, where **kwargs goes to dal_module rather than to the test.
* The DPC++ examples take -DONEMATH_BACKEND=ON, excluding
  kmeans_lloyd_csr_batch, and -DCUDA_BACKEND=ON, additionally excluding
  pca_svd_dense_batch.

That last one comes from an audit of what oneMath's CUDA backends
actually implement of the surface oneDAL uses, now in the README:
cuBLAS covers float and double gemm/gemv/syrk; cuSOLVER covers potrf,
potrs and syevd but rejects gesvd for m < n, which is exactly how the
PCA SVD kernel calls it; cuRAND covers philox4x32x10 and mrg32k3a
including the USM uniform overloads, so decision forest works on its
default engine, but implements skip_ahead as an absolute offset rather
than a relative advance. Read from the oneMath sources, not measured --
there is no NVIDIA hardware here.

Signed-off-by: Alexandr-Solovev <aleksandr.solovev@intel.com>
…e alias rename

Routing the device BLAS through the `mkl::` backend alias changed the length of
the callee, so clang-format wants the continuation lines realigned. Pure
whitespace; the CI check added in uxlfoundation#3813 flags the three files otherwise.

Signed-off-by: Alexandr-Solovev <aleksandr.solovev@intel.com>
…sults

The oneMath section claimed the implemented surface was BLAS and LAPACK only,
which undersold it: two of the five device RNG engines work too, and sparse BLAS
is the single domain that does not. Say that in the lead paragraph and in the
backend table, and add what the Bazel suites do under `--config=nvidia-gpu` --
67 targets pass, the sparse_blas targets are skipped as incompatible, and the one
failure is the decision forest SPMD test, which fails the same way on the default
backend.

Signed-off-by: Alexandr-Solovev <aleksandr.solovev@intel.com>
@Alexandr-Solovev
Alexandr-Solovev force-pushed the dev/asolovev_onemath_nvidia branch from bbf34ee to 7fcdc29 Compare October 1, 2026 18:24
@Alexandr-Solovev

Copy link
Copy Markdown
Contributor Author

Rebased onto main to clear the conflicts (7fcdc2987).

The conflict was with #3812, which consolidated the .github/instructions/*.instructions.md files. This branch had added four lines about the device math backend to .github/instructions/build-systems.instructions.md, and upstream deleted that file — a modify/delete conflict. Resolved by accepting the deletion and porting the four lines to where that content now lives, dev/AGENTS.md:

- `BACKEND_CONFIG`: `mkl` (default on x86-64) or `ref` (OpenBLAS; default on ARM and RISC-V). Names the host math library only.
- `DPC_MATH_BACKEND`: `mkl` (default) or `onemath`, the math library behind the DPC++ device code. `onemath` additionally reaches NVIDIA GPUs and needs `ONEMATHROOT`; it mirrors Bazel's `--dpc_math_backend` and is independent of `BACKEND_CONFIG`.

Two commits on top of the rebase:

  • f387f6d38 — pure whitespace. clang-format realigns the continuation lines in blas/{gemm,gemv,syrk}_dpc.cpp after the mkl:: alias rename; verified against the upstream versions of those files that the three violations come from this branch and not from a formatter version difference.
  • 7fcdc2987 — dev/bazel/README.md: the oneMath surface is BLAS, LAPACK and two of the five device RNG engines, but not sparse BLAS, and the suite results are recorded (the three sparse_blas targets are skipped as incompatible, and decision_forest:test_spmd_dpc fails the same way it does on the default backend on that GPU).

Rebuilt after the rebase on both backends: //cpp/oneapi/dal/backend/primitives/blas:tests builds clean under --config=dpc and under --config=nvidia-gpu.

…ries

The onemath_runtime filegroup had no users. Staging the backends in the repo
cannot help either: the dispatcher dlopens them through a $ORIGIN RUNPATH,
which under Bazel is the _solib directory, so they are found through
LD_LIBRARY_PATH as dev/bazel/README.md documents.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants