Performance

Benchmarks

Cross-vendor, cross-stack profiling benchmarks. Measure what matters on the hardware you have.

Matrix Multiplication Benchmark

Tutorial: Matrix MultiplicationOaFnMatrix::MatMul through OaContext. All numbers are wall-clock including full public API path (submit + sync overhead). Methodology: 50 iterations after 5 warmup runs, RTX 5090 Laptop GPU.

Single-Dispatch Throughput

ShapeMNKWall msOA BF16OA Fp32Ref TF32OA / Ref
square-5125125125120.0823,2802,58232,02710.2%
square-10241024102410240.13815,5486,21542,14236.9%
square-20482048204820480.89719,15910,47038,65049.6%
tall-skinny409612810240.1537,0044,78637,37718.7%
short-wide128409610240.1357,9585,01639,72120.0%
gemv-decode1409640960.90337494208.8%

Device theoretical peak: 64 TFLOPS (BF16 tensor cores). Throughput climbs with problem size as fixed per-dispatch overhead amortizes, peaking near 19.2 TFLOP/s at 2048³.

Batch-Dispatch Throughput (4 ops per submit)

ShapeMNKOA BF16OA Fp32Ref TF32 (single)OA Batch / Ref Single
square-51251251251234,03126,88835,22597%
square-102410241024102477,60343,30881,39995%
tall-skinny2048128102434,97828,16574,59947%

Batching amortizes CPU→GPU submission overhead. For small ops where GPU time ≈ submit latency, batching provides 4–8× speedup vs single-dispatch.

Cross-Device Portability

RTX 5090 LaptopIntel ARL iGPU
Precision pathBF16 CoopMat (tensor cores)FP32 tiled/naive
norm_err (2048³)~5e-30.00 (bit-exact)
Throughput (2048³)19.2 TFLOP/s259 GFLOP/s
Theoretical peak64 TFLOPS BF162.5 TFLOPS FP32
% of peak (2048³)30.0%10.4%

Same binary, same source — runtime device selection via OA_DEVICE=integrated. Intel iGPU runs the FP32 fallback path at 259 GFLOP/s (10.4% of 2.5 TFLOP theoretical peak).

Run.sh

cmake --build Build/Release --target TutorialCoreMatMulIntro -j
./Bin/Release/Tutorial/Core/TutorialCoreMatMulIntro