Native APIs provide different constraints and features when it comes to resource copies and clears, where resources can be buffers or images. In this issue, we'll try to find a common ground (a least common denominator API) that is usable and efficient on all backends.
In Metal, all of the copy/clear operations are done via the MTLBlitCommandEncoder.
In Vulkan, these are transfer operations, supported on any queue type. They require TRANSFER_SRC flag on the source and TRANSFER_DST flag on the destination.
Operation table
Buffer Updates
In D3D12, the only way to update a buffer with new data coming from CPU is to use a staging buffer (that is mapped, filled, then copied to the destination).
In Metal, similar effect can be achieved by creating a buffer with makeBuffer that re-uses the existing storage.
In Vulkan, the implementation may have a fast-path for small buffer updates by in-lining the data right into the command buffer space. The implementation can fall back to a staging-like scheme for larger updates.
Image Blitting
Image blits are different from image copies for allowing format conversion and arbitrary scaling with filtering. A typical use case for blitting is mipmap generation. It is not clear to me why/how Vulkan provides this on a transfer-only queue, but other APIs are far more (and reasonably) limited with regards to where and how they can blit surfaces.
Alignment rules
Vulkan
VkPhysicalDeviceLimits has optimal alignments for buffer data when transferring to/from image:
optimalBufferCopyOffsetAlignment is the optimal buffer offset alignment in bytes for vkCmdCopyBufferToImage and vkCmdCopyImageToBuffer
optimalBufferCopyRowPitchAlignment is the optimal buffer row pitch alignment in bytes for vkCmdCopyBufferToImage and vkCmdCopyImageToBuffer
These are not enforced by the validation layers but are recommended for optimal performance.
D3D12
MSDN section lists the following restrictions:
- linear subresource copying must be aligned to
D3D12_TEXTURE_DATA_PLACEMENT_ALIGNMENT (512) bytes
- row pitch aligned to
D3D12_TEXTURE_DATA_PITCH_ALIGNMENT (256) bytes
Proposed API
Clears
D3D12 model appears to be the least common denominator. If we have the concept of views, we can have API calls to clear them. In Vulkan, these calls would trivially translate into direct clears. In Metal, we'd need to run a compute shader to clear the resources. Supporting multiple cear rectangles seems to complicate this scheme quite a bit, so I suggest only doing the full-slice clears.
Updates
Given the limited support of resource updates, I suggest not providing this API at all in favor of requiring the user to use staging resources manually.
Copies
All 3 APIs appear to provide the copy capability between buffers and textures. The difference is mostly about the alignment requirements. I suggest having device flags to the minimum offset/pitch required:
- D3D12: equal to D3D12 constants
- Vulkan: equal to optimal alignment features
- Metal: some reasonable default selected by Apple
Blits
D3D12 doesn't support any sort of blitting, I'm inclined to propose no workarounds here. Users doing simple render passes for blitting textures shouldn't be slower than emulating this in the API, anyway.
Afterword
This analysis may be incomplete, corrections are welcome to go directly as the issue edits.
Native APIs provide different constraints and features when it comes to resource copies and clears, where resources can be buffers or images. In this issue, we'll try to find a common ground (a least common denominator API) that is usable and efficient on all backends.
In Metal, all of the copy/clear operations are done via the
MTLBlitCommandEncoder.In Vulkan, these are transfer operations, supported on any queue type. They require
TRANSFER_SRCflag on the source andTRANSFER_DSTflag on the destination.Operation table
Buffer Updates
In D3D12, the only way to update a buffer with new data coming from CPU is to use a staging buffer (that is mapped, filled, then copied to the destination).
In Metal, similar effect can be achieved by creating a buffer with makeBuffer that re-uses the existing storage.
In Vulkan, the implementation may have a fast-path for small buffer updates by in-lining the data right into the command buffer space. The implementation can fall back to a staging-like scheme for larger updates.
Image Blitting
Image blits are different from image copies for allowing format conversion and arbitrary scaling with filtering. A typical use case for blitting is mipmap generation. It is not clear to me why/how Vulkan provides this on a transfer-only queue, but other APIs are far more (and reasonably) limited with regards to where and how they can blit surfaces.
Alignment rules
Vulkan
VkPhysicalDeviceLimits has optimal alignments for buffer data when transferring to/from image:
optimalBufferCopyOffsetAlignmentis the optimal buffer offset alignment in bytes forvkCmdCopyBufferToImageandvkCmdCopyImageToBufferoptimalBufferCopyRowPitchAlignmentis the optimal buffer row pitch alignment in bytes forvkCmdCopyBufferToImageandvkCmdCopyImageToBufferThese are not enforced by the validation layers but are recommended for optimal performance.
D3D12
MSDN section lists the following restrictions:
D3D12_TEXTURE_DATA_PLACEMENT_ALIGNMENT(512) bytesD3D12_TEXTURE_DATA_PITCH_ALIGNMENT(256) bytesProposed API
Clears
D3D12 model appears to be the least common denominator. If we have the concept of views, we can have API calls to clear them. In Vulkan, these calls would trivially translate into direct clears. In Metal, we'd need to run a compute shader to clear the resources. Supporting multiple cear rectangles seems to complicate this scheme quite a bit, so I suggest only doing the full-slice clears.
Updates
Given the limited support of resource updates, I suggest not providing this API at all in favor of requiring the user to use staging resources manually.
Copies
All 3 APIs appear to provide the copy capability between buffers and textures. The difference is mostly about the alignment requirements. I suggest having device flags to the minimum offset/pitch required:
Blits
D3D12 doesn't support any sort of blitting, I'm inclined to propose no workarounds here. Users doing simple render passes for blitting textures shouldn't be slower than emulating this in the API, anyway.
Afterword
This analysis may be incomplete, corrections are welcome to go directly as the issue edits.