A unified C++20 inference library that compiles ONNX models and runs them across heterogeneous compute backends through a clean, backend-agnostic API.
Yukti splits the inference lifecycle into two independent stages:
Compile — An ONNX graph is compiled to a backend-specific plan by the Parser using a registered IBackend. The resulting ModelDesc (I/O metadata + plan path) is serialised to JSON and can be shared, versioned, or cached independently of the runtime.
Run — An EngineHandle loads a JSON config, dlopen-s the matching backend plugin (libyukti_engine_<backend>.so), and exposes a zero-copy inference interface. Backends can be hot-swapped at runtime without rebuilding the handle.
ONNX file
│
▼
Parser ──► IBackend::compile() ──► plan file + ModelDesc ──► JSON config
│
▼
EngineHandle (dlopen)
│
alpaka buffers ─┤
raw pointers ─┤
▼
IEngine::infer()
- Backend-agnostic API — compile and run code is identical regardless of which backend is active
- Plugin architecture — inference backends ship as separate shared libraries; the core library has no hard dependency on TensorRT or any other framework at link time
- Zero-copy inference —
BufferViewandAlpakaAdapterforward raw device pointers directly to the engine; host memory is never touched during inference - Alpaka interop — first-class support for alpaka buffers (CUDA, ROCm, CPU Serial) via
EngineHandle::inferAlpaka() - Hot-swap — call
EngineHandle::load()to swap backends or reload a plan without reconstructing the handle - Registry-driven compilation — backends register themselves via
registerBuiltinBackends(); no concrete backend type needs to appear in user code
| Backend | Status | Notes |
|---|---|---|
| TensorRT | Supported | Requires TensorRT ≥ 8, CUDA Toolkit |
| SOFIE | Supported | TMVA SOFIE header-only inference |
| Plugin | Status | Notes |
|---|---|---|
libyukti_engine_tensorrt.so |
Supported | FP32, FP16, INT8; explicit-batch networks |
libyukti_engine_sofie.so |
Supported |
| Device | Alpaka Backend | Status |
|---|---|---|
| NVIDIA GPU | alpaka::AccGpuCudaRt |
Supported |
| AMD GPU | alpaka::AccGpuHipRt |
Planned |
| CPU | alpaka::AccCpuSerial |
Supported |
| Dependency | Version | When required |
|---|---|---|
| CMake | ≥ 3.18 | Always |
| C++ compiler | C++20 | Always |
| nlohmann/json | 3.11.3 | Always (fetched automatically) |
| CUDA Toolkit | ≥ 11.0 | YUKTI_ENABLE_TENSORRT=ON |
| TensorRT | ≥ 8.0 | YUKTI_ENABLE_TENSORRT=ON |
| alpaka | commit 2fa91a3 |
YUKTI_ENABLE_ALPAKA=ON (fetched automatically) |
| Catch2 | 3.5.4 | YUKTI_BUILD_TESTS=ON (fetched automatically) |
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build buildcmake -B build \
-DCMAKE_BUILD_TYPE=Release \
-DYUKTI_ENABLE_TENSORRT=ON
cmake --build buildIf TensorRT headers are not on the default search path, point CMake at them:
cmake -B build \
-DYUKTI_ENABLE_TENSORRT=ON \
-DTENSORRT_INCLUDE_DIR=/path/to/tensorrt/includecmake -B build \
-DCMAKE_BUILD_TYPE=Release \
-DYUKTI_ENABLE_TENSORRT=ON \
-DYUKTI_ENABLE_ALPAKA=ON
cmake --build build| CMake option | Default | Effect |
|---|---|---|
YUKTI_ENABLE_TENSORRT |
OFF |
Build TensorRT parser backend + engine plugin |
YUKTI_ENABLE_SOFIE |
OFF |
Build SOFIE parser backend + inference plugin |
YUKTI_ENABLE_ALPAKA |
OFF |
Enable alpaka buffer adapter and inferAlpaka() |
YUKTI_BUILD_TESTS |
OFF |
Build the test suite |
TENSORRT_INCLUDE_DIR |
auto | Override TensorRT header search path |
CMAKE_EXPORT_COMPILE_COMMANDS |
OFF |
Generate compile_commands.json for IDEs |
After building, copy the headers and libraries manually or install via CMake:
cmake --install build --prefix /usr/localThe engine plugin (libyukti_engine_tensorrt.so) must be reachable at runtime. Either:
- place it in a directory on
LD_LIBRARY_PATH, or - set
YUKTI_PLUGIN_DIRto the directory containing the.so.
export YUKTI_PLUGIN_DIR=/path/to/plugins#include "yukti/parser/parser.hxx"
#include "yukti/parser/backend_registry.hxx"
#include "yukti/common/model_desc.hxx"
using namespace yukti;
registerBuiltinBackends();
Parser parser;
auto graph = parser.load("resnet50.onnx");
CompileOptions opts;
opts.outputPath = "resnet50.trt";
opts.fp16 = true;
opts.workspaceMB = 1024;
auto backend = BackendRegistry::instance().create("tensorrt");
ModelDesc desc = parser.compile(graph, *backend, opts);
desc.name = "resnet50";
saveModelDesc(desc, "resnet50.json");#include "yukti/engine/engine_handle.hxx"
using namespace yukti;
EngineHandle handle("resnet50.json");
// Raw pointer path (any device allocation)
std::vector<NamedBuffer> inputs = {{ "input", BufferView{...} }};
std::vector<NamedBuffer> outputs = {{ "output", BufferView{...} }};
handle.infer(inputs, outputs);#include "yukti/engine/engine_handle.hxx"
#include <alpaka/alpaka.hpp>
using Dim = alpaka::DimInt<1>;
using Idx = std::size_t;
alpaka::PlatformCudaRt platform;
auto device = alpaka::getDevByIdx(platform, 0u);
auto in_buf = alpaka::allocBuf<float, Idx>(device, 1600u);
auto out_buf = alpaka::allocBuf<float, Idx>(device, 160u);
// Fill in_buf ...
decltype(in_buf)* p_in = &in_buf;
decltype(out_buf)* p_out = &out_buf;
EngineHandle handle("resnet50.json");
handle.inferAlpaka(std::span{&p_in, 1}, std::span{&p_out, 1});handle.load("resnet50_int8.json"); // unloads old plugin, loads new oneBuild with tests enabled:
cmake -B build \
-DYUKTI_BUILD_TESTS=ON \
-DYUKTI_ENABLE_TENSORRT=ON \
-DYUKTI_ENABLE_ALPAKA=ON
cmake --build buildRun all tests:
ctest --test-dir build --output-on-failure| Test binary | What it covers | Requirements |
|---|---|---|
test_adapters |
View, BufferView, AlpakaAdapter, RawAdapter, DataType helpers |
None |
test_engine |
ModelDesc JSON round-trip, tensor metadata, NamedBuffer |
None |
test_e2e_trt |
ONNX → TRT plan → JSON → alpaka CUDA inference, hot-swap | CUDA GPU + TensorRT |
The end-to-end test uses Linear_16.onnx (a 100→10 linear layer, batch 16) bundled in the tests/ directory. Individual test cases are skipped automatically if their prerequisites (GPU, plan file) are not present.
Run only the non-GPU tests:
ctest --test-dir build -E e2e_trt --output-on-failureinclude/yukti/
├── common/
│ ├── types.hxx # DataType, DeviceType, exception hierarchy
│ ├── model_desc.hxx # TensorDesc, ModelDesc (I/O metadata + plan path)
│ ├── view.hxx # View<T> — non-owning typed span
│ ├── buffer_view.hxx # BufferView — device-aware, untyped buffer descriptor
│ └── adapters/
│ ├── alpaka_adapter.hxx # Zero-copy alpaka ↔ BufferView interop
│ └── raw_adapter.hxx # Raw pointer → BufferView helpers
├── parser/
│ ├── IBackend.hxx # Compilation backend interface + CompileOptions
│ ├── backend_registry.hxx # Singleton registry; registerBuiltinBackends()
│ ├── parser.hxx # Parser: load ONNX, compile with a backend
│ └── backends/
│ └── tensorrt.hxx # TensorRTBackend (requires YUKTI_ENABLE_TENSORRT)
└── engine/
├── IEngine.hxx # Runtime engine interface
└── engine_handle.hxx # Plugin loader; infer() and inferAlpaka()
- Implement
IBackendininclude/yukti/parser/backends/mybackend.hxxandsrc/parser/backends/mybackend.cxx. - Register it in
src/parser/parser_registry.cxxinsideregisterBuiltinBackends(). - Implement
IEngineinsrc/engine/mybackend/mybackend_engine.cxx, exportyukti_create_engine/yukti_destroy_engineasextern "C", and build it as a shared library namedlibyukti_engine_mybackend.so. - Add the corresponding CMake targets guarded by a
YUKTI_ENABLE_MYBACKENDoption.
- AMD / ROCm support — HIP-based engine plugin, alpaka
AccGpuHipRtintegration - Dynamic shape support — runtime shape profiles and optimisation ranges in
CompileOptions - Multi-stream inference — expose CUDA stream handles through
IEnginefor concurrent execution