Currently in the spec, workgroup_size is optional. It defaults to (1,1,1).
Using (1,1,1) is usually a terrible choice, for performance.
It will always work but almost always the worst choice (except for the ones where you'd exceed a register budget).
I understand this has already caused some pain in bringing up some applications. cc: @kainino0x
So the proposal is: A compute shader must have a workgroup_size attribute.
(Don't default the size in the x dimension. Still allow defaulting to 1 for the y and z dimensions.)
A rejected alternative is to have the implementation pick a workgroup size. The thinking is that drivers probably don't have a great idea of what will work best. Instead, pull the application into the decision loop.
Currently in the spec, workgroup_size is optional. It defaults to (1,1,1).
Using (1,1,1) is usually a terrible choice, for performance.
It will always work but almost always the worst choice (except for the ones where you'd exceed a register budget).
I understand this has already caused some pain in bringing up some applications. cc: @kainino0x
So the proposal is: A compute shader must have a workgroup_size attribute.
(Don't default the size in the x dimension. Still allow defaulting to 1 for the y and z dimensions.)
A rejected alternative is to have the implementation pick a workgroup size. The thinking is that drivers probably don't have a great idea of what will work best. Instead, pull the application into the decision loop.