Skip to content
ML4EPPublic

About

A Unified Interface for ML Inference on heterogeneous architectures

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Yukti - Your Unified KiT for Inference

A unified C++20 inference library that compiles ONNX models and runs them across heterogeneous compute backends through a clean, backend-agnostic API.


Overview

Yukti splits the inference lifecycle into two independent stages:

Compile — An ONNX graph is compiled to a backend-specific plan by the Parser using a registered IBackend. The resulting ModelDesc (I/O metadata + plan path) is serialised to JSON and can be shared, versioned, or cached independently of the runtime.

Run — An EngineHandle loads a JSON config, dlopen-s the matching backend plugin (libyukti_engine_<backend>.so), and exposes a zero-copy inference interface. Backends can be hot-swapped at runtime without rebuilding the handle.

ONNX file
    │
    ▼
Parser ──► IBackend::compile() ──► plan file + ModelDesc ──► JSON config
                                                                    │
                                                                    ▼
                                                           EngineHandle (dlopen)
                                                                    │
                                                    alpaka buffers ─┤
                                                    raw pointers   ─┤
                                                                    ▼
                                                               IEngine::infer()

Features

  • Backend-agnostic API — compile and run code is identical regardless of which backend is active
  • Plugin architecture — inference backends ship as separate shared libraries; the core library has no hard dependency on TensorRT or any other framework at link time
  • Zero-copy inference — BufferView and AlpakaAdapter forward raw device pointers directly to the engine; host memory is never touched during inference
  • Alpaka interop — first-class support for alpaka buffers (CUDA, ROCm, CPU Serial) via EngineHandle::inferAlpaka()
  • Hot-swap — call EngineHandle::load() to swap backends or reload a plan without reconstructing the handle
  • Registry-driven compilation — backends register themselves via registerBuiltinBackends(); no concrete backend type needs to appear in user code

Current Support

Compilation Backends (IBackend)

Backend Status Notes
TensorRT Supported Requires TensorRT ≥ 8, CUDA Toolkit
SOFIE Supported TMVA SOFIE header-only inference

Runtime Engine Plugins (IEngine)

Plugin Status Notes
libyukti_engine_tensorrt.so Supported FP32, FP16, INT8; explicit-batch networks
libyukti_engine_sofie.so Supported

Device / Accelerator Support

Device Alpaka Backend Status
NVIDIA GPU alpaka::AccGpuCudaRt Supported
AMD GPU alpaka::AccGpuHipRt Planned
CPU alpaka::AccCpuSerial Supported

Requirements

Dependency Version When required
CMake ≥ 3.18 Always
C++ compiler C++20 Always
nlohmann/json 3.11.3 Always (fetched automatically)
CUDA Toolkit ≥ 11.0 YUKTI_ENABLE_TENSORRT=ON
TensorRT ≥ 8.0 YUKTI_ENABLE_TENSORRT=ON
alpaka commit 2fa91a3 YUKTI_ENABLE_ALPAKA=ON (fetched automatically)
Catch2 3.5.4 YUKTI_BUILD_TESTS=ON (fetched automatically)

Build

Minimal (no backends, headers + model serialisation only)

cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build

With TensorRT backend

cmake -B build \
    -DCMAKE_BUILD_TYPE=Release \
    -DYUKTI_ENABLE_TENSORRT=ON

cmake --build build

If TensorRT headers are not on the default search path, point CMake at them:

cmake -B build \
    -DYUKTI_ENABLE_TENSORRT=ON \
    -DTENSORRT_INCLUDE_DIR=/path/to/tensorrt/include

With TensorRT + Alpaka (CUDA)

cmake -B build \
    -DCMAKE_BUILD_TYPE=Release \
    -DYUKTI_ENABLE_TENSORRT=ON \
    -DYUKTI_ENABLE_ALPAKA=ON

cmake --build build

All options

CMake option Default Effect
YUKTI_ENABLE_TENSORRT OFF Build TensorRT parser backend + engine plugin
YUKTI_ENABLE_SOFIE OFF Build SOFIE parser backend + inference plugin
YUKTI_ENABLE_ALPAKA OFF Enable alpaka buffer adapter and inferAlpaka()
YUKTI_BUILD_TESTS OFF Build the test suite
TENSORRT_INCLUDE_DIR auto Override TensorRT header search path
CMAKE_EXPORT_COMPILE_COMMANDS OFF Generate compile_commands.json for IDEs

Installation

After building, copy the headers and libraries manually or install via CMake:

cmake --install build --prefix /usr/local

The engine plugin (libyukti_engine_tensorrt.so) must be reachable at runtime. Either:

  • place it in a directory on LD_LIBRARY_PATH, or
  • set YUKTI_PLUGIN_DIR to the directory containing the .so.
export YUKTI_PLUGIN_DIR=/path/to/plugins

Usage

1. Compile an ONNX model

#include "yukti/parser/parser.hxx"
#include "yukti/parser/backend_registry.hxx"
#include "yukti/common/model_desc.hxx"

using namespace yukti;

registerBuiltinBackends();

Parser parser;
auto graph = parser.load("resnet50.onnx");

CompileOptions opts;
opts.outputPath  = "resnet50.trt";
opts.fp16        = true;
opts.workspaceMB = 1024;

auto backend  = BackendRegistry::instance().create("tensorrt");
ModelDesc desc = parser.compile(graph, *backend, opts);

desc.name = "resnet50";
saveModelDesc(desc, "resnet50.json");

2. Run inference

#include "yukti/engine/engine_handle.hxx"

using namespace yukti;

EngineHandle handle("resnet50.json");

// Raw pointer path (any device allocation)
std::vector<NamedBuffer> inputs  = {{ "input",  BufferView{...} }};
std::vector<NamedBuffer> outputs = {{ "output", BufferView{...} }};
handle.infer(inputs, outputs);

3. Run inference with alpaka buffers (zero-copy)

#include "yukti/engine/engine_handle.hxx"
#include <alpaka/alpaka.hpp>

using Dim = alpaka::DimInt<1>;
using Idx = std::size_t;

alpaka::PlatformCudaRt platform;
auto device = alpaka::getDevByIdx(platform, 0u);

auto in_buf  = alpaka::allocBuf<float, Idx>(device, 1600u);
auto out_buf = alpaka::allocBuf<float, Idx>(device, 160u);

// Fill in_buf ...

decltype(in_buf)*  p_in  = &in_buf;
decltype(out_buf)* p_out = &out_buf;

EngineHandle handle("resnet50.json");
handle.inferAlpaka(std::span{&p_in, 1}, std::span{&p_out, 1});

4. Hot-swap the backend

handle.load("resnet50_int8.json");   // unloads old plugin, loads new one

Tests

Build with tests enabled:

cmake -B build \
    -DYUKTI_BUILD_TESTS=ON \
    -DYUKTI_ENABLE_TENSORRT=ON \
    -DYUKTI_ENABLE_ALPAKA=ON

cmake --build build

Run all tests:

ctest --test-dir build --output-on-failure

Test suite

Test binary What it covers Requirements
test_adapters View, BufferView, AlpakaAdapter, RawAdapter, DataType helpers None
test_engine ModelDesc JSON round-trip, tensor metadata, NamedBuffer None
test_e2e_trt ONNX → TRT plan → JSON → alpaka CUDA inference, hot-swap CUDA GPU + TensorRT

The end-to-end test uses Linear_16.onnx (a 100→10 linear layer, batch 16) bundled in the tests/ directory. Individual test cases are skipped automatically if their prerequisites (GPU, plan file) are not present.

Run only the non-GPU tests:

ctest --test-dir build -E e2e_trt --output-on-failure

Architecture

include/yukti/
├── common/
│   ├── types.hxx            # DataType, DeviceType, exception hierarchy
│   ├── model_desc.hxx       # TensorDesc, ModelDesc (I/O metadata + plan path)
│   ├── view.hxx             # View<T>  — non-owning typed span
│   ├── buffer_view.hxx      # BufferView — device-aware, untyped buffer descriptor
│   └── adapters/
│       ├── alpaka_adapter.hxx   # Zero-copy alpaka ↔ BufferView interop
│       └── raw_adapter.hxx      # Raw pointer → BufferView helpers
├── parser/
│   ├── IBackend.hxx         # Compilation backend interface + CompileOptions
│   ├── backend_registry.hxx # Singleton registry; registerBuiltinBackends()
│   ├── parser.hxx           # Parser: load ONNX, compile with a backend
│   └── backends/
│       └── tensorrt.hxx     # TensorRTBackend (requires YUKTI_ENABLE_TENSORRT)
└── engine/
    ├── IEngine.hxx          # Runtime engine interface
    └── engine_handle.hxx    # Plugin loader; infer() and inferAlpaka()

Adding a new backend

  1. Implement IBackend in include/yukti/parser/backends/mybackend.hxx and src/parser/backends/mybackend.cxx.
  2. Register it in src/parser/parser_registry.cxx inside registerBuiltinBackends().
  3. Implement IEngine in src/engine/mybackend/mybackend_engine.cxx, export yukti_create_engine / yukti_destroy_engine as extern "C", and build it as a shared library named libyukti_engine_mybackend.so.
  4. Add the corresponding CMake targets guarded by a YUKTI_ENABLE_MYBACKEND option.

Future Plans

  • AMD / ROCm support — HIP-based engine plugin, alpaka AccGpuHipRt integration
  • Dynamic shape support — runtime shape profiles and optimisation ranges in CompileOptions
  • Multi-stream inference — expose CUDA stream handles through IEngine for concurrent execution

License

GNU General Public License v3.0

About

A Unified Interface for ML Inference on heterogeneous architectures

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors