Sitelet https://github.com/uxlfoundation/oneDNN/releases
Skip to content

Releases: uxlfoundation/oneDNN

v3.13.4

Choose a tag to compare

@tprimak tprimak released this 06 Oct 03:16

This is a patch release containing the following changes to v3.13.3:

  • Fixed crash in grouped matmul with large shapes on Intel GPUs based on Xe architecture (bb0e26c, 93e82fc, e81aa24, 1411796, c899370, b933459, a93026c)
  • Fixed correctness issue in f16, bf16, and s8/u8 convolutions with small number of input channels on x64 processors with Intel AMX instruction set support (fed7e41)
  • Fixed correctness issue in f16, bf16, and s8/u8 convolutions with trivial output width on x64 processors with Intel AMX instruction set support (8b5f0ab, e5cb5ff)
  • Fixed correctness issue in f32 convolution weight gradienton x64 processors with Intel AVX-512 instruction set support (4256033)

v3.14-rc

v3.14-rc Pre-release
Pre-release

Choose a tag to compare

@vpirogov vpirogov released this 02 Oct 23:06

Performance Optimizations

Intel 64/AMD64 Processors

  • Introduced initial support for AI Compute Extensions (ACE) instructions. This functionality is not dispatched by default and requires opt-in with environment variable ONEDNN_MAX_CPU_ISA=AVX10_2_ACE.
  • Improved performance of future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
  • Improved performance of fp8 matmul with block-wise weights on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
  • Improved performance of bf16 and f16 matmul with relaxed accumulation mode on Intel Xeon processors with Intel AMX instruction set support.
  • Improved performance of strided deconvolution on Intel Xeon Scalable processors (formerly Sapphire Rapids).
  • Improved performance of floating-point Scaled Dot Product Attention (SDPA) training forward propagation subgraph with Graph API.

Intel Graphics

  • Improved performance of future discrete GPUs based on Xe3p-XPC architecture (codename Crescent Island).
  • Improved performance of future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
  • Improved performance of Scaled Dot Product Attention (SDPA) training forward and backpropagation subgraphs.
  • Improved performance of Scaled Dot Product Attention (SDPA) subgraphs with asymmetric head sizes.
  • Improved performance of grouped matmul with small group sizes.
  • Improved performance of f8 matmul with source and weights scales mask 3.
  • Improved performance of f16 and s8 matmul with u4 weights and N=1.
  • Reduced convolution and deconvolution primitives creation time.

AArch64 Processors

  • Improved the performance of u8 and s8 matmul on platforms with SVE support.
  • Improved the performance of f32 matmul on platforms with 128-bit SVE vector lengths.
  • Improved the performance of f32 depthwise convolution on platforms with ASIMD support.
  • Improved the performance of convolutions.
  • Improved the performance of eltwise_log on platforms with ASIMD support.
  • Improved the performance of bf16 eltwise for the gelu_erf, swish, gelu_tanh, exp, log, and sqrt algorithms.
  • Improved the performance of f32, and f16 PReLU.
  • Improved the performance of eltwise post-ops.
  • Improved the performance of the logsoftmax algorithm for the softmax primitive on platforms with ASIMD support.
  • Reduced penalties on small utility functions on clang builds by changing the default stack-protection level from all to strong.

RISC-V Processors

  • Improved performance of f32 binary, eltwise, pooling, softmax, and logsoftmax on processors with V extension support.
  • Improved performance of f16 matmul, eltwise, and softmax on processors with Zvfh extension support, including softmax with non-contiguous axes.
  • Extended RVV-optimized implementations to bf16 binary, eltwise, pooling, softmax, batch normalization, and resampling, and improved bf16 matmul performance on processors with Zvfbfwma extension support.
  • Introduced RVV-optimized forward resampling with nearest-neighbor and linear interpolation for f32 and f16 data types.
  • Introduced RVV-optimized shuffle for f32, s32, f16, and bf16 data types.
  • Extended the RVV matmul implementation to support unsigned 8-bit source and weights.

Functionality

Functional API

  • Introduced the binary_mul_inplace algorithm for binary post-ops. Unlike binary_mul, the new algorithm allows the matmul destination tensor to be used as one of its inputs. An optimized implementation is available for matmul on Intel GPUs.
  • [experimental] Extended eltwise post-ops support in grouped matmul with all supported algorithms. Optimized implementation is available on Intel GPUs.
  • [experimental] Extended grouped matmul with support for backpropagation cases (2D grouped by 3D dense and 2D grouped by 2D grouped) covering f32, f16 and bf16 data types. Optimized implementation is available for Intel GPUs. This is an experimental feature that requires opt-in with ONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ON build option.

Graph API

Usability

Common

  • Updated mxfp8 downconversion implementations to saturate instead of overflowing. New behavior is consistent with OCP MX specification and aligned with preferred behavior in PyTorch.
  • Version number can now be used as a passable approximation of pi.

Intel 64/AMD64 processors

  • Cleaned up implicit narrowing conversions and removed suppression of MSVC compiler warning C4244.

Intel Graphics

  • Refactored verbose profiling implementation for Level Zero runtime to avoid spurious synchronizations.
  • Introduced support for concurrent primitive execution with the Level Zero runtime on Intel GPUs.
  • [experimental] Introduced support for verbose profiling based on sycl_ext_oneapi_profiling_tag SYCL extension. This is an experimental feature that requires opt-in with ONEDNN_EXPERIMENTAL_ENABLE_SYCL_PROFILING_TAG=ON build option.

AArch64 Processors

  • Introduced initial asynchronous runtime support to AArch64 platforms for the matmul, convolution, eltwise, binary, lnorm, and reorder primitives.
  • Fixed a memory leak in convolutions on platforms with SVE support.

Validation

  • Updated benchdnn smoke and CI test sets for matmul using parameter space sampling approach.
  • [experimental] Extended benchdnn --grouped knob with balanced, hot, and decode strategies for offset generation to generate MoE-style group distributions in grouped matmul validation.
  • Extended benchdnn graph driver: operation attribute removal via --op-attrs knob, scalar tensor support via --in-shapes knob, tensor property rewriting via --tensor-property knob, operation removal via --op-kind.

Deprecated Functionality

  • BLAS-like API including dnnl::sgemm, dnnl::gemm_u8s8s32, and dnnl::gemm_s8s8s32 functions is deprecated and will be removed in future releases. If you are using this API consider switching to matmul primitive.

Breaking changes

  • Removed optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th generation Intel Core processors. oneDNN remains functional on these platforms and dispatches a generic OpenCL implementation.
  • Removed optimizations for processors with Intel SSE4.1 and Intel AVX instruction sets. oneDNN remains functional on these platforms and dispatches a generic C++ implementation.
  • Removed optimizations for tf32 fpmath_mode in matmul on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).

Thanks to our Contributors

This release contains contributions from the project core team as well as Abhishek Kumar @abhishek-iitmadras, Aditya Singh @adityasingh2400, Akihiro Tabuchi @Akihiro-Tabuchi, AragornOfKebroyd @AragornOfKebroyd, Aron Xu @happyaron, @AyushSinghBaiswar, Codrut Irimie @CodrutIrimieARM, Crefeda Rodrigues @cfRod, elimor01 @MorelElian, Emilio Cota @cota, Ishita Shreya @ishita-shreya, Kamil Jackiewicz @kjackiew, Kamil Wieloch @kwieloch-intel, Keerthana KT @Keerthana-64, Léandre LE DUC @leduclean, Leon Kennedy @leoken01, Megha Sangtani @megha-sangtani, Mohammed Bilgrami @mohbil01, Nikhil Gupta @nikhil-arm, PiotrReiterIntel @PiotrReiterIntel, Puneet Matharu @puneetmatharu, @rinatrap, Thiago Macieira @thiagomacieira, @Tiwari-Avanish, Udit Kumar Agarwal @uditagarwal97, @velonica0, Wang hongyan @ww8191201-coder, and @xinghai-zh.

Each of them contributed an invaluable slice of the pi.

v3.13.3

Choose a tag to compare

@vpirogov vpirogov released this 25 Sep 21:24

This is a patch release containing the following changes to v3.13.2:

  • Improved performance of Scaled Dot Product Attention (SDPA) training forward subgraph
    with Graph API on x64 CPUs (e971e88, b5206d3, bf8ffaa)
  • Fixed build errors with -Wunused-template diagnostic promoted to an error (75c3dfa, 6e7312c)
  • Improved performance of SDPA backpropagation on Intel GPUs (fc9fed9, 695317e)
  • Introduced dnnl::verbose_profiling_enabled function to allow applications to check
    whether verbose profiling mode requires queue profiling to be enabled (e2a772f, d005652, 1716509)
  • Fixed crash in convolution with large padding on CPUs with Intel AVX-512 and
    Intel AVX2 instruction set support (3c1407a)
  • Fixed crash and hang in verbose mode with Graph API and
    ONEDNN_CPU_RUNTIME=THREADPOOL when using an asynchronous threadpool
    (a35fd4f, 25b6be6, a1c07f2, b7dc0d5)
  • Fixed sporadic correctness issue in SDPA subgraph on Intel GPUs (b38b47c)
  • Fixed correctness issue in RNN primitive backpropagation with GRU cell type and dhc == 1 on Intel GPUs (0e486f6)
  • Fixed a crash in matmul with transposed tensor B and non-trivial strides on x64 CPUs (c0abb1f)
  • [experimental] Introduced support for verbose profiling for SYCL runtime based on sycl_ext_oneapi_profiling_tag extension (6a45b07, 303e817, e26fd5a, 3fc4f4e, 8f8e9ec, bdf5118, bdf5118, 4f1cd92, 13d58e9)
  • Fixed crash in matmul with binary post-ops and bf16 or f16 broadcasted tensors on x64 CPUs with Intel AVX-512 and Intel DL Boost support (933b30a, b29c933)
  • Fixed correctness issues in f32 3D matmul with fp16 or bf16 weights on x64 CPUs (68f6a26, 171872e)
  • Fixed f32 SDPA subgraph performance regression on Intel GPUs (1a38c2d)
  • Fixed performance regression in f32 convolutions with 1x1 kernel on AArch64 CPUs with SVE support (bb92917)
  • Fixed performance regression in batch normalization on AArch64 CPUs with SVE support (a72c4a7)
  • Changed Clang compiler flag from -fstack-protector-all to -fstack-protector-strong for builds on AArch64 CPUs (27fc2c1)
  • Fixed correctness issue in f32 matmul with binary add post-op preceding sum post-op on Intel GPUs (f1c9e43)

v3.13.2

Choose a tag to compare

@tprimak tprimak released this 26 Aug 22:53

This is a patch release containing the following changes to v3.13.1:

  • Removed check for AMX_TF32 CPUID bits on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (code name Diamond Rapids) (9b2cdda)

v3.12.5

Choose a tag to compare

@tprimak tprimak released this 26 Aug 22:54

This is a patch release containing the following changes to v3.12.4:

  • Removed check for AMX_TF32 CPUID bits on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (code name Diamond Rapids) (ea02464)

v3.13.1

Choose a tag to compare

@vpirogov vpirogov released this 19 Aug 23:01

This is a patch release containing the following changes to v3.13:

v3.12.4

Choose a tag to compare

@vpirogov vpirogov released this 19 Aug 21:09

This is a patch release containing the following changes to v3.12.3:

  • Fixed performance regression in f16 matmul on processors with Intel AMX (bf16, u8/s8) instruction set support (e9c54e5)

v3.13

Choose a tag to compare

@vpirogov vpirogov released this 17 Jul 21:55

Performance Optimizations

Intel 64/AMD64 Processors

  • Improved performance on future Intel Core Ultra processors with Intel AVX10.2 instruction set support (codename Nova Lake).
  • Improved performance of matmul on processors with Intel AMX instruction set support.
  • Improved performance of bf16, f16, and f32 matmul with unit M, N or K dimentions (GEMV-like) on processors with Intel AVX2 instruction set support.
  • Improved performance of u8/s8 matmul with u4/s4 weights and grouped scales.
  • Improved peformance of bf16 and f16 matmul with f8 weights.
  • Improved performance of f8 quantized Scaled Dot Product Attention (SDPA) subgraph with Graph API.

Intel Graphics

  • Improved performance for future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
  • Improved u8/s8 convolution performance on Intel Arc B-series graphics.
  • Improved f16 and u8/s8 matmul performance with u8/s8 and u4/s4 weights in non-transposed layout.

AArch64 Processors

  • Improved performance of u8/s8 matmuls with u8/s8, f16, or s32 outputs.
  • Improved performance of u8/s8 convolutions with bf16 or f32 outputs.
  • Improved u8/s8 layer normalization performance.
  • Improved performance of convolution backpropagation and pooling on platforms with 128-bit SVE.
  • Improved performance of bf16 inner-product.
  • Improved multi-threaded bnorm performance.
  • Improved binary primitive and post-op performance.
  • Improved performance of eltwise primitive with gelu_erf algorithm and post-op.

RISC-V Processors

  • Improved f32 convolution, matmul, inner product, binary, eltwise, pooling, batch normalization, and group normalization primitive performance on processors with V extension support.
  • Improved f16 matmul, binary, eltwise, pooling, softmax, and layer normalization primitive performance on processors with Zvfh extension support.
  • Improved bf16 matmul primitive performance on processors with Zvfbfwma extension support.

Functionality

Functional API

  • [experimental] Introduced support for eltwise and binary post-ops in matmul with grouped memory. Optimized implementation is available on Intel GPUs.
  • [experimental] Extended grouped matmul with NVFP4 quantization scheme, including support for f4_e2m1 tensors with f8_e4m3 grouped scales and per-group binary post-op to implement global fp32 scale. This is an experimental feature that requires opt-in with ONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ON build option.

Graph API

  • Introduced support for device-side seed, offset, and probability arguments for Dropout operation.

Usability

Common

Intel Graphics

  • Refactored verbose profiling on Intel GPUs to avoid spurious synchronizations with SYCL or OpenCL runtimes. The new implementation reports device time instead of host time and is compatible with SYCL Graph record/replay mode.
  • Reduced memory consumption of Gated MLP subgraph with Graph API.
  • Enabled interoperability with SYCL Graph native recording mode for Intel GPUs.
  • Introduced ONEDNN_ZE_INCLUDE_DIR and ONEDNN_OCL_INCLUDE_DIR build knobs to use Level Zero or OpenCL headers from a user-defined location instead of the vendored headers.
  • [experimental] Introduced support for persistent cache with Level Zero runtime on GPU. Level Zero support is experimental.

AArch64 Processors

  • Update convolutions accumulation data type in certain cases to conform with library numerical behavior requirements.
  • Reduced stack-space usage across all primitives.
  • Fixed a correctness issue with leaky ReLU with alpha > 1.

Validation

  • Extended SYCL Graph validation mode in benchdnn with support for native recording mode. This mode is enabled using --execution-mode=native_graph knob.
  • Enabled SYCL recording mode validation --execution-mode=graph for benchdnn --graph driver.
  • Introduced benchdnn knob --mode=S to improve performance validation speed in simulation or emulation environments.
  • Improved GPU performance reporting for --mode=F by stabilizing measurement methodology and reducing inaccuracies caused by cache effects and run-to-run variability.

Deprecated Functionality

  • The BLAS-like API, including dnnl::sgemm, dnnl::gemm_u8s8s32, and dnnl::gemm_s8s8s32, is deprecated
    and will be removed in future releases. If you are using this API, consider switching to the matmul primitive.
  • f4_e3m0 data type is deprecated and will be removed in future releases.
  • Optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th Generation Intel Core Processors are deprecated and will be removed in future releases.
  • Optimizations for processors with Intel SSE4.1 support and Intel AVX support are deprecated and will be removed in the future releases.

Breaking Changes

  • Updated minimal supported Arm Compute Library version to v53.1.0 (was v52.7.0).

Thanks to our Contributors

This release contains contributions from the project core team as well as Alexandre de Limas Santana @alexandrelimassantana, Andrei Hutu @Anndrey24, Anna Sztukowska @asztukow, @bhanuprasad14, Fadi Arafeh @fadara01, George Nash @georgen117, Georgii Zagoruiko @AstonMartin-one-77, Henry Gardiner @henry-gar, Kamil Wieloch @kwieloch-intel, Keanu Czirjak @keanucz, Michał Patronik @mikita12, Qize Li @Ga1axy0, Rohan @Rohanjames1997, @velonica0, and Xiuchuan Zhai @azhai219.

v3.12.3

Choose a tag to compare

@vpirogov vpirogov released this 15 Jul 23:08

This is a patch release containing the following changes to v3.12.2:

  • Fixed potential memory corruption in Graph API for logical tensors with number of dimensions exceeding 12 (1470adb)
  • Enabled matmul and inner product primitives compatibility with SYCL Graph native recording mode on Intel GPUs (817bf1f, a16c22f, ad55577)
  • Fixed performance regression in SDPA subgraph with head size 64 on Intel GPUs based on Xe2 architecture (486c7f7, ee77a16)

v3.13-rc

v3.13-rc Pre-release
Pre-release

Choose a tag to compare

@vgvozdeva vgvozdeva released this 09 Jul 11:33

Performance Optimizations

Intel 64/AMD64 Processors

  • Improved performance on future Intel Core Ultra processors with Intel AVX10.2 instruction set support (codename Nova Lake).
  • Improved performance of matmul on processors with Intel AMX instruction set support.
  • Improved performance of bf16/f16/f32 matmul with unit M/N or K dimentions (GEMV-like) on processors with Intel AVX2 instruction set support.
  • Improved performance of u8/s8 matmul with u4/s4 weights and grouped scales.
  • Improved peformance of bf16/f16 matmul with f8 weights.
  • Improved performance of f8 quantized Scaled Dot Product Attention (SDPA) subgraph with Graph API.

Intel Graphics

  • Improved performance for future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
  • Improved u8/s8 convolution performance on Intel Arc B-series graphics.
  • Improved f16 and u8/s8 matmul performance with u8/s8 and u4/s4 weights in non-transposed layout.

AArch64 Processors

  • Improved u8/s8 matmuls with u8/s8/f16/s32 outputs
  • Improved u8/s8 convolutions with bf16/f32 outputs
  • Improved u8/s8 lnorm performance
  • Improved performance of convolution training on platforms with 128-bit SVE
  • Improved performance of pooling on platforms with 128-bit SVE
  • Improved performance of bf16 inner-product
  • Improved multi-threaded bnorm performance
  • Improved binary operator, and post-op performance
  • Improved performance of gelu_erf activations

RISC-V ProcessorsExpand commentComment on line R28Resolved

  • Improved f32 convolution, matmul, inner product, binary, eltwise, pooling, batch normalization, and group normalization primitive performance on processors with V extension support.
  • Improved f16 matmul, binary, eltwise, pooling, softmax, and layer normalization primitive performance on processors with Zvfh extension support.
  • Improved bf16 matmul primitive performance on processors with Zvfbfwma extension support.

Functionality

Functional API

  • [experimental] Introduced support for eltwise and binary post-ops in matmul with grouped memory. Optimized implementation is available on Intel GPUs.
  • [experimental] Extended grouped matmul with NVFP4 quantization scheme, including support for f4_e2m1 tensors with f8_e4m3 grouped scales and per-group binary post-op to implement global fp32 scale. This is an experimental feature that requires opt-in with ONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ON build option.

Graph API

  • Introduced support for device-side seed, offset, and probability arguments for Dropout operation.

Usability

Common

Intel Graphics

  • Refactored verbose profiling on Intel GPUs to avoid spurious synchronizations with SYCL or OpenCL runtimes. The new implementation reports device time instead of host time and is compatible with SYCL Graph record/replay mode.
  • Reduced memory consumption of Gated MLP subgraph with Graph API.
  • Enabled interoperability with SYCL Graph native recording mode for Intel GPUs.
  • Introduced ONEDNN_ZE_INCLUDE_DIR and ONEDNN_OCL_INCLUDE_DIR build knobs to use Level Zero or OpenCL headers from a user-defined location instead of the vendored headers.
  • [experimental] Introduced support for persistent cache with Level Zero runtime on GPU. Level Zero support is experimental.

AArch64 Processors

  • Fixed a correctness issue with leaky ReLU with alpha > 1
  • Fixed an issue where convolutions could be accumulated in a lower precision than intended
  • Reduced baseline stack-space usage across all operators

Validation

  • Extended SYCL Graph validation mode in benchdnn with support for native recording mode. This mode is enabled using --execution-mode=native_graph knob.
  • Enabled SYCL recording mode validation --execution-mode=graph for benchdnn --graph driver.
  • Introduced benchdnn knob --mode=S to improve performance validation speed in simulation or emulation environments.
  • Improved GPU performance reporting for --mode=F by stabilizing measurement methodology and reducing inaccuracies caused by cache effects and run-to-run variability.

Deprecated Functionality

  • The BLAS-like API, including dnnl::sgemm, dnnl::gemm_u8s8s32, and dnnl::gemm_s8s8s32, is deprecated
    and will be removed in future releases. If you are using this API, consider switching to the matmul primitive.
  • f4_e3m0 data type is deprecated and will be removed in future releases.
  • Optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th Generation Intel Core Processors are deprecated and will be removed in future releases.

Breaking Changes

  • The minimum version of Arm® Compute Library is now v53.1.0

Thanks to our Contributors

This release contains contributions from the project core team as well as Alexandre de Limas Santana @alexandrelimassantana, Andrei Hutu @Anndrey24, Anna Sztukowska @asztukow, @bhanuprasad14, Fadi Arafeh @fadara01, George Nash @georgen117, Georgii Zagoruiko @AstonMartin-one-77, Henry Gardiner @henry-gar, Kamil Wieloch @kwieloch-intel, Keanu Czirjak @keanucz, Michał Patronik @mikita12, Qize Li @Ga1axy0, Rohan @Rohanjames1997, velonica0 @velonica0 and Xiuchuan Zhai @azhai219.