Current states by using latest dawn and latest tint on DX12:
As we can see, the second one with post-multiplication used 3 multiply and 3 add instructions where the first one used 4 x dp4.
Proposed optimization for tint:
- Using "mad" instruction (multiply and add) of Shader Module 5 assembly to optimize the post-multiplication from second one, the best we can get is:
mul r0.xyzw, v0.yyyy, CB0[0][1].xyzw
mad r0.xyzw, v0.xxxx, CB0[0][0].xyzw, r0.xyzw
mad r0.xyzw, v0.zzzz, CB0[0][2].xyzw, r0.xyzw
add r0.xyzw, r0.xyzw, CB0[0][3].xyzw
Instead of using 6 instructions, now we are using 4.
Proposed wgsl feature:
- Adding a "variable-type modifier" to wgsl which is similar to HLSL. e.g.
struct GlobalParams
{
row_major viewMatrix : mat4x4<f32>;
column_major projMatrix : mat4x4<f32>;
};
- So now multiplication for matrix becomes:
- Pre-multiplications:
-
// Pre- multiplication with raw-major matrix
var transformedPos:vec4<f32> = vec4<f32>(input.position, 1.0) * uParams.viewMatrix;
// DXBC:
// 1 mul + 2 mad + 1 add, as described above.
-
// Pre- multiplication with column-major matrix
var transformedPos:vec4<f32> = vec4<f32>(input.position, 1.0) * uParams.projMatrix ;
// DXBC:
// 4 x dp4, as described above.
- Post-multiplications:
-
// Post- multiplication with raw-major matrix
var transformedPos:vec4<f32> = uParams.viewMatrix * vec4<f32>(input.position, 1.0);
// DXBC:
// 4 x dp4, as described above.
-
// Post- multiplication with column-major matrix
var transformedPos:vec4<f32> = uParams.projMatrix * vec4<f32>(input.position, 1.0) ;
// DXBC:
// 1 mul + 2 mad + 1 add, as described above.
- Benifiets:
- It is not changing the constant buffer memory layout order, but the generated shader assembly instructions. It only gives compiler a hint to generate desired instructions according to the variable-type modifier.
- Gives a more explicit way for the user to control their matrix binding strategy without depending on changing the constant buffer layout on CPU or sometimes need a transpose for different major order matrix on CPU. e.g. a row major based matrix lib on CPU side can now use it in shader with pre-multiplication and without transpose before uploaded to GPU buffer.
// viewMatrix as decalred as row_major
// Now passing raw-major matrix from cpu without transposing would work the same as post-multiplication(matrix * vector )
var transformedPos:vec4<f32> = vec4<f32>(input.position, 1.0) * uParams.viewMatrix;
- Let user to determine where should be used with either "4xdp4" or "1 mul + 2 mad + 1 add".
- Refactor current documentations of matrix multiplication, as it always suggesting to use post multiplication, which is not always true, as when running on DX12, current implementation has a bit performance overhead as described above.
Current states by using latest dawn and latest tint on DX12:
Pre-multiplication:
And after tint compiling it to DXBC (using PIX for DX12 as profiler for dawn):

Post-multiplication:
and DXBC from tint:

As we can see, the second one with post-multiplication used 3 multiply and 3 add instructions where the first one used 4 x dp4.
Proposed optimization for tint:
Proposed wgsl feature: