Repository navigation
Releases: uxlfoundation/oneDNN
Release list
v3.13.4
This is a patch release containing the following changes to v3.13.3:
- Fixed crash in grouped matmul with large shapes on Intel GPUs based on Xe architecture (bb0e26c, 93e82fc, e81aa24, 1411796, c899370, b933459, a93026c)
- Fixed correctness issue in
f16,bf16, ands8/u8convolutions with small number of input channels on x64 processors with Intel AMX instruction set support (fed7e41) - Fixed correctness issue in
f16,bf16, ands8/u8convolutions with trivial output width on x64 processors with Intel AMX instruction set support (8b5f0ab, e5cb5ff) - Fixed correctness issue in
f32convolution weight gradienton x64 processors with Intel AVX-512 instruction set support (4256033)
v3.14-rc
Performance Optimizations
Intel 64/AMD64 Processors
- Introduced initial support for AI Compute Extensions (ACE) instructions. This functionality is not dispatched by default and requires opt-in with environment variable
ONEDNN_MAX_CPU_ISA=AVX10_2_ACE. - Improved performance of future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
- Improved performance of
fp8matmul with block-wise weights on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids). - Improved performance of
bf16andf16matmul with relaxed accumulation mode on Intel Xeon processors with Intel AMX instruction set support. - Improved performance of strided deconvolution on Intel Xeon Scalable processors (formerly Sapphire Rapids).
- Improved performance of floating-point Scaled Dot Product Attention (SDPA) training forward propagation subgraph with Graph API.
Intel Graphics
- Improved performance of future discrete GPUs based on Xe3p-XPC architecture (codename Crescent Island).
- Improved performance of future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
- Improved performance of Scaled Dot Product Attention (SDPA) training forward and backpropagation subgraphs.
- Improved performance of Scaled Dot Product Attention (SDPA) subgraphs with asymmetric head sizes.
- Improved performance of grouped matmul with small group sizes.
- Improved performance of
f8matmul with source and weights scales mask3. - Improved performance of
f16ands8matmul withu4weights andN=1. - Reduced convolution and deconvolution primitives creation time.
AArch64 Processors
- Improved the performance of
u8ands8matmul on platforms with SVE support. - Improved the performance of
f32matmul on platforms with 128-bit SVE vector lengths. - Improved the performance of
f32depthwise convolution on platforms with ASIMD support. - Improved the performance of convolutions.
- Improved the performance of
eltwise_logon platforms with ASIMD support. - Improved the performance of
bf16eltwise for thegelu_erf,swish,gelu_tanh,exp,log, andsqrtalgorithms. - Improved the performance of
f32, andf16PReLU. - Improved the performance of eltwise post-ops.
- Improved the performance of the
logsoftmaxalgorithm for the softmax primitive on platforms with ASIMD support. - Reduced penalties on small utility functions on clang builds by changing the default stack-protection level from
alltostrong.
RISC-V Processors
- Improved performance of
f32binary, eltwise, pooling, softmax, and logsoftmax on processors withVextension support. - Improved performance of
f16matmul, eltwise, and softmax on processors withZvfhextension support, including softmax with non-contiguous axes. - Extended RVV-optimized implementations to
bf16binary, eltwise, pooling, softmax, batch normalization, and resampling, and improvedbf16matmul performance on processors withZvfbfwmaextension support. - Introduced RVV-optimized forward resampling with nearest-neighbor and linear interpolation for
f32andf16data types. - Introduced RVV-optimized shuffle for
f32,s32,f16, andbf16data types. - Extended the RVV matmul implementation to support unsigned 8-bit source and weights.
Functionality
Functional API
- Introduced the
binary_mul_inplacealgorithm for binary post-ops. Unlikebinary_mul, the new algorithm allows the matmul destination tensor to be used as one of its inputs. An optimized implementation is available for matmul on Intel GPUs. - [experimental] Extended eltwise post-ops support in grouped matmul with all supported algorithms. Optimized implementation is available on Intel GPUs.
- [experimental] Extended grouped matmul with support for backpropagation cases (2D grouped by 3D dense and 2D grouped by 2D grouped) covering
f32,f16andbf16data types. Optimized implementation is available for Intel GPUs. This is an experimental feature that requires opt-in withONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ONbuild option.
Graph API
- Extended DynamicQuantize and DynamicDequantize operation to support the new
maskattribute.
Usability
Common
- Updated
mxfp8downconversion implementations to saturate instead of overflowing. New behavior is consistent with OCP MX specification and aligned with preferred behavior in PyTorch. - Version number can now be used as a passable approximation of pi.
Intel 64/AMD64 processors
- Cleaned up implicit narrowing conversions and removed suppression of MSVC compiler warning C4244.
Intel Graphics
- Refactored verbose profiling implementation for Level Zero runtime to avoid spurious synchronizations.
- Introduced support for concurrent primitive execution with the Level Zero runtime on Intel GPUs.
- [experimental] Introduced support for verbose profiling based on sycl_ext_oneapi_profiling_tag SYCL extension. This is an experimental feature that requires opt-in with
ONEDNN_EXPERIMENTAL_ENABLE_SYCL_PROFILING_TAG=ONbuild option.
AArch64 Processors
- Introduced initial asynchronous runtime support to AArch64 platforms for the matmul, convolution, eltwise, binary, lnorm, and reorder primitives.
- Fixed a memory leak in convolutions on platforms with SVE support.
Validation
- Updated benchdnn
smokeandCItest sets for matmul using parameter space sampling approach. - [experimental] Extended benchdnn
--groupedknob withbalanced,hot, anddecodestrategies for offset generation to generate MoE-style group distributions in grouped matmul validation. - Extended benchdnn graph driver: operation attribute removal via
--op-attrsknob, scalar tensor support via--in-shapesknob, tensor property rewriting via--tensor-propertyknob, operation removal via--op-kind.
Deprecated Functionality
- BLAS-like API including
dnnl::sgemm,dnnl::gemm_u8s8s32, anddnnl::gemm_s8s8s32functions is deprecated and will be removed in future releases. If you are using this API consider switching to matmul primitive.
Breaking changes
- Removed optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th generation Intel Core processors. oneDNN remains functional on these platforms and dispatches a generic OpenCL implementation.
- Removed optimizations for processors with Intel SSE4.1 and Intel AVX instruction sets. oneDNN remains functional on these platforms and dispatches a generic C++ implementation.
- Removed optimizations for
tf32fpmath_modein matmul on future Intel Xeon processors with Intel AVX10.2 and Intel AMX instruction set support (codename Diamond Rapids).
Thanks to our Contributors
This release contains contributions from the project core team as well as Abhishek Kumar @abhishek-iitmadras, Aditya Singh @adityasingh2400, Akihiro Tabuchi @Akihiro-Tabuchi, AragornOfKebroyd @AragornOfKebroyd, Aron Xu @happyaron, @AyushSinghBaiswar, Codrut Irimie @CodrutIrimieARM, Crefeda Rodrigues @cfRod, elimor01 @MorelElian, Emilio Cota @cota, Ishita Shreya @ishita-shreya, Kamil Jackiewicz @kjackiew, Kamil Wieloch @kwieloch-intel, Keerthana KT @Keerthana-64, Léandre LE DUC @leduclean, Leon Kennedy @leoken01, Megha Sangtani @megha-sangtani, Mohammed Bilgrami @mohbil01, Nikhil Gupta @nikhil-arm, PiotrReiterIntel @PiotrReiterIntel, Puneet Matharu @puneetmatharu, @rinatrap, Thiago Macieira @thiagomacieira, @Tiwari-Avanish, Udit Kumar Agarwal @uditagarwal97, @velonica0, Wang hongyan @ww8191201-coder, and @xinghai-zh.
Each of them contributed an invaluable slice of the pi.
v3.13.3
This is a patch release containing the following changes to v3.13.2:
- Improved performance of Scaled Dot Product Attention (SDPA) training forward subgraph
with Graph API on x64 CPUs (e971e88, b5206d3, bf8ffaa) - Fixed build errors with
-Wunused-templatediagnostic promoted to an error (75c3dfa, 6e7312c) - Improved performance of SDPA backpropagation on Intel GPUs (fc9fed9, 695317e)
- Introduced
dnnl::verbose_profiling_enabledfunction to allow applications to check
whether verbose profiling mode requires queue profiling to be enabled (e2a772f, d005652, 1716509) - Fixed crash in convolution with large padding on CPUs with Intel AVX-512 and
Intel AVX2 instruction set support (3c1407a) - Fixed crash and hang in verbose mode with Graph API and
ONEDNN_CPU_RUNTIME=THREADPOOLwhen using an asynchronous threadpool
(a35fd4f, 25b6be6, a1c07f2, b7dc0d5) - Fixed sporadic correctness issue in SDPA subgraph on Intel GPUs (b38b47c)
- Fixed correctness issue in RNN primitive backpropagation with GRU cell type and
dhc == 1on Intel GPUs (0e486f6) - Fixed a crash in matmul with transposed tensor
Band non-trivial strides on x64 CPUs (c0abb1f) - [experimental] Introduced support for verbose profiling for SYCL runtime based on
sycl_ext_oneapi_profiling_tagextension (6a45b07, 303e817, e26fd5a, 3fc4f4e, 8f8e9ec, bdf5118, bdf5118, 4f1cd92, 13d58e9) - Fixed crash in matmul with binary post-ops and
bf16orf16broadcasted tensors on x64 CPUs with Intel AVX-512 and Intel DL Boost support (933b30a, b29c933) - Fixed correctness issues in
f323D matmul withfp16orbf16weights on x64 CPUs (68f6a26, 171872e) - Fixed
f32SDPA subgraph performance regression on Intel GPUs (1a38c2d) - Fixed performance regression in
f32convolutions with 1x1 kernel on AArch64 CPUs with SVE support (bb92917) - Fixed performance regression in batch normalization on AArch64 CPUs with SVE support (a72c4a7)
- Changed Clang compiler flag from
-fstack-protector-allto-fstack-protector-strongfor builds on AArch64 CPUs (27fc2c1) - Fixed correctness issue in
f32matmul with binary add post-op preceding sum post-op on Intel GPUs (f1c9e43)
v3.13.2
v3.12.5
v3.13.1
This is a patch release containing the following changes to v3.13:
- Fixed correctness issue in matmul with non-
f32common scales on x64 CPUs (8649862) - Fixed an
unimplementederror in grouped matmul with post-ops on Intel GPUs (168ccaf) - Fixed correctness issue in
f8grouped matmul on Intel GPUs based on Xe-LPG architecture (0a2644a) - Fixed matmul correctness issue for non-trivial source strides on x64 CPUs (6e5689d, 21ed470)
- Extended grouped matmul post-ops to support all eltwise algorithms on Intel GPUs (1b5607c, b24a1db, 3829e47, 879d412, 678f419)
- Fixed performance regression in
f16matmul withN = 1and binary post-op on x64 CPUs (0b9ccf7) - Fixed a performance regression in
f64matmul with large K on Intel GPUs (4a44f5f) - Fixed an
unimplementederror in grouped matmul withu8weights and zero points on Intel GPUs based on Xe-HPG architecture (f9d64d0, 347e3f4) - Cleaned up implicit narrowing conversions and removed suppression of MSVC compiler warning C4244 on x64 CPUs (af6fc5e, 2a1e6af, d7b4acd, 428ab19, 21ebf76, 9392d8b, f0172dc, 4345414, 6203120, 29c3bb3, ca64e87, 554c09f, 514dc23, 4afe4b2, c92aa9d, 800de48, 4f681cb, 82192d6, f28a821, 764729f, 71e5cc3, 48e7204, cdedd44, 23b98ba, 6375c24, 0929f31, d36e03a, 39dbfc5, 1dc80b5, 3508778)
- Reduced convolution and deconvolution primitives creation time on Intel GPUs (9be3cfe, 521463a, 251a8d6, dcc10e6, 4c2a7d7, 3a78962, da8af99, 268bafe)
- Improved performance of grouped grouped matmul primitive with eltwise post-ops on Intel GPUs (99b824e, b6fa787, b0c8687, d46a6cb, 2e3e918, 226a0da, dd0f425, 44391fc, e7b8ab8, 281167b, e09f474, 2be907b, a209014, 82200ce)
- Fixed crash during convolution primitive creation with large shapes on x64 CPUs (a3d4597)
v3.12.4
v3.13
Performance Optimizations
Intel 64/AMD64 Processors
- Improved performance on future Intel Core Ultra processors with Intel AVX10.2 instruction set support (codename Nova Lake).
- Improved performance of matmul on processors with Intel AMX instruction set support.
- Improved performance of
bf16,f16, andf32matmul with unit M, N or K dimentions (GEMV-like) on processors with Intel AVX2 instruction set support. - Improved performance of
u8/s8matmul withu4/s4weights and grouped scales. - Improved peformance of
bf16andf16matmul withf8weights. - Improved performance of
f8quantized Scaled Dot Product Attention (SDPA) subgraph with Graph API.
Intel Graphics
- Improved performance for future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
- Improved
u8/s8convolution performance on Intel Arc B-series graphics. - Improved
f16andu8/s8matmul performance withu8/s8andu4/s4weights in non-transposed layout.
AArch64 Processors
- Improved performance of
u8/s8matmuls withu8/s8,f16, ors32outputs. - Improved performance of
u8/s8convolutions withbf16orf32outputs. - Improved
u8/s8layer normalization performance. - Improved performance of convolution backpropagation and pooling on platforms with 128-bit SVE.
- Improved performance of
bf16inner-product. - Improved multi-threaded bnorm performance.
- Improved binary primitive and post-op performance.
- Improved performance of eltwise primitive with
gelu_erfalgorithm and post-op.
RISC-V Processors
- Improved
f32convolution, matmul, inner product, binary, eltwise, pooling, batch normalization, and group normalization primitive performance on processors withVextension support. - Improved
f16matmul, binary, eltwise, pooling, softmax, and layer normalization primitive performance on processors withZvfhextension support. - Improved
bf16matmul primitive performance on processors withZvfbfwmaextension support.
Functionality
Functional API
- [experimental] Introduced support for eltwise and binary post-ops in matmul with grouped memory. Optimized implementation is available on Intel GPUs.
- [experimental] Extended grouped matmul with NVFP4 quantization scheme, including support for
f4_e2m1tensors withf8_e4m3grouped scales and per-group binary post-op to implement globalfp32scale. This is an experimental feature that requires opt-in withONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ONbuild option.
Graph API
- Introduced support for device-side seed, offset, and probability arguments for
Dropoutoperation.
Usability
Common
- Introduced user-managed scratchpad support in Graph API.
Intel Graphics
- Refactored verbose profiling on Intel GPUs to avoid spurious synchronizations with SYCL or OpenCL runtimes. The new implementation reports device time instead of host time and is compatible with SYCL Graph record/replay mode.
- Reduced memory consumption of Gated MLP subgraph with Graph API.
- Enabled interoperability with SYCL Graph native recording mode for Intel GPUs.
- Introduced
ONEDNN_ZE_INCLUDE_DIRandONEDNN_OCL_INCLUDE_DIRbuild knobs to use Level Zero or OpenCL headers from a user-defined location instead of the vendored headers. - [experimental] Introduced support for persistent cache with Level Zero runtime on GPU. Level Zero support is experimental.
AArch64 Processors
- Update convolutions accumulation data type in certain cases to conform with library numerical behavior requirements.
- Reduced stack-space usage across all primitives.
- Fixed a correctness issue with leaky ReLU with alpha > 1.
Validation
- Extended SYCL Graph validation mode in benchdnn with support for native recording mode. This mode is enabled using
--execution-mode=native_graphknob. - Enabled SYCL recording mode validation
--execution-mode=graphfor benchdnn--graphdriver. - Introduced benchdnn knob
--mode=Sto improve performance validation speed in simulation or emulation environments. - Improved GPU performance reporting for
--mode=Fby stabilizing measurement methodology and reducing inaccuracies caused by cache effects and run-to-run variability.
Deprecated Functionality
- The BLAS-like API, including
dnnl::sgemm,dnnl::gemm_u8s8s32, anddnnl::gemm_s8s8s32, is deprecated
and will be removed in future releases. If you are using this API, consider switching to the matmul primitive. f4_e3m0data type is deprecated and will be removed in future releases.- Optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th Generation Intel Core Processors are deprecated and will be removed in future releases.
- Optimizations for processors with Intel SSE4.1 support and Intel AVX support are deprecated and will be removed in the future releases.
Breaking Changes
- Updated minimal supported Arm Compute Library version to v53.1.0 (was v52.7.0).
Thanks to our Contributors
This release contains contributions from the project core team as well as Alexandre de Limas Santana @alexandrelimassantana, Andrei Hutu @Anndrey24, Anna Sztukowska @asztukow, @bhanuprasad14, Fadi Arafeh @fadara01, George Nash @georgen117, Georgii Zagoruiko @AstonMartin-one-77, Henry Gardiner @henry-gar, Kamil Wieloch @kwieloch-intel, Keanu Czirjak @keanucz, Michał Patronik @mikita12, Qize Li @Ga1axy0, Rohan @Rohanjames1997, @velonica0, and Xiuchuan Zhai @azhai219.
v3.12.3
This is a patch release containing the following changes to v3.12.2:
- Fixed potential memory corruption in Graph API for logical tensors with number of dimensions exceeding 12 (1470adb)
- Enabled matmul and inner product primitives compatibility with SYCL Graph native recording mode on Intel GPUs (817bf1f, a16c22f, ad55577)
- Fixed performance regression in SDPA subgraph with head size 64 on Intel GPUs based on Xe2 architecture (486c7f7, ee77a16)
v3.13-rc
Performance Optimizations
Intel 64/AMD64 Processors
- Improved performance on future Intel Core Ultra processors with Intel AVX10.2 instruction set support (codename Nova Lake).
- Improved performance of matmul on processors with Intel AMX instruction set support.
- Improved performance of
bf16/f16/f32matmul with unit M/N or K dimentions (GEMV-like) on processors with Intel AVX2 instruction set support. - Improved performance of
u8/s8matmul withu4/s4weights and grouped scales. - Improved peformance of
bf16/f16matmul withf8weights. - Improved performance of
f8quantized Scaled Dot Product Attention (SDPA) subgraph with Graph API.
Intel Graphics
- Improved performance for future integrated GPUs based on Xe3p-LPG architecture (codename Nova Lake P).
- Improved
u8/s8convolution performance on Intel Arc B-series graphics. - Improved
f16andu8/s8matmul performance withu8/s8andu4/s4weights in non-transposed layout.
AArch64 Processors
- Improved
u8/s8matmuls withu8/s8/f16/s32outputs - Improved
u8/s8convolutions withbf16/f32outputs - Improved
u8/s8lnorm performance - Improved performance of convolution training on platforms with 128-bit SVE
- Improved performance of pooling on platforms with 128-bit SVE
- Improved performance of
bf16inner-product - Improved multi-threaded bnorm performance
- Improved binary operator, and post-op performance
- Improved performance of
gelu_erfactivations
RISC-V ProcessorsExpand commentComment on line R28Resolved
- Improved
f32convolution, matmul, inner product, binary, eltwise, pooling, batch normalization, and group normalization primitive performance on processors withVextension support. - Improved
f16matmul, binary, eltwise, pooling, softmax, and layer normalization primitive performance on processors withZvfhextension support. - Improved
bf16matmul primitive performance on processors withZvfbfwmaextension support.
Functionality
Functional API
- [experimental] Introduced support for eltwise and binary post-ops in matmul with grouped memory. Optimized implementation is available on Intel GPUs.
- [experimental] Extended grouped matmul with NVFP4 quantization scheme, including support for
f4_e2m1tensors withf8_e4m3grouped scales and per-group binary post-op to implement globalfp32scale. This is an experimental feature that requires opt-in withONEDNN_EXPERIMENTAL_GROUPED_MEMORY=ONbuild option.
Graph API
- Introduced support for device-side seed, offset, and probability arguments for
Dropoutoperation.
Usability
Common
- Introduced user-managed scratchpad support in Graph API.
Intel Graphics
- Refactored verbose profiling on Intel GPUs to avoid spurious synchronizations with SYCL or OpenCL runtimes. The new implementation reports device time instead of host time and is compatible with SYCL Graph record/replay mode.
- Reduced memory consumption of Gated MLP subgraph with Graph API.
- Enabled interoperability with SYCL Graph native recording mode for Intel GPUs.
- Introduced
ONEDNN_ZE_INCLUDE_DIRandONEDNN_OCL_INCLUDE_DIRbuild knobs to use Level Zero or OpenCL headers from a user-defined location instead of the vendored headers. - [experimental] Introduced support for persistent cache with Level Zero runtime on GPU. Level Zero support is experimental.
AArch64 Processors
- Fixed a correctness issue with leaky ReLU with alpha > 1
- Fixed an issue where convolutions could be accumulated in a lower precision than intended
- Reduced baseline stack-space usage across all operators
Validation
- Extended SYCL Graph validation mode in benchdnn with support for native recording mode. This mode is enabled using
--execution-mode=native_graphknob. - Enabled SYCL recording mode validation
--execution-mode=graphfor benchdnn--graphdriver. - Introduced benchdnn knob
--mode=Sto improve performance validation speed in simulation or emulation environments. - Improved GPU performance reporting for
--mode=Fby stabilizing measurement methodology and reducing inaccuracies caused by cache effects and run-to-run variability.
Deprecated Functionality
- The BLAS-like API, including
dnnl::sgemm,dnnl::gemm_u8s8s32, anddnnl::gemm_s8s8s32, is deprecated
and will be removed in future releases. If you are using this API, consider switching to the matmul primitive. f4_e3m0data type is deprecated and will be removed in future releases.- Optimizations for Intel Iris Xe MAX Graphics and Intel Graphics included with 11th-14th Generation Intel Core Processors are deprecated and will be removed in future releases.
Breaking Changes
- The minimum version of Arm® Compute Library is now v53.1.0
Thanks to our Contributors
This release contains contributions from the project core team as well as Alexandre de Limas Santana @alexandrelimassantana, Andrei Hutu @Anndrey24, Anna Sztukowska @asztukow, @bhanuprasad14, Fadi Arafeh @fadara01, George Nash @georgen117, Georgii Zagoruiko @AstonMartin-one-77, Henry Gardiner @henry-gar, Kamil Wieloch @kwieloch-intel, Keanu Czirjak @keanucz, Michał Patronik @mikita12, Qize Li @Ga1axy0, Rohan @Rohanjames1997, velonica0 @velonica0 and Xiuchuan Zhai @azhai219.