-
Notifications
You must be signed in to change notification settings - Fork 386
Multiple Queues skeleton proposal. #95
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,83 @@ | ||
| # Multiple Queues | ||
|
|
||
| Multiple queues is important for: | ||
|
|
||
| - Explicit submission onto less capable queues, such as pure copy queues. | ||
| - Explicit distribution of submissions on machines with multiple queues of the same family. | ||
|
|
||
| ## QueueFamily | ||
|
|
||
| There are multiple families of queues. | ||
| Availible families of queues are surfaced via `sequence<QueueFamily> Adapter.queueFamilies`. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. type: Availible
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. typo: type 🤣 |
||
| Support for `graphics`, `compute`, and `copy` operations are surfaced by the respective attributes on the QueueFamily. | ||
| Queue families with `graphics` or `compute` will always have `copy`. | ||
| There may be queue families in the future like D3D12's video decode queues, which do not support copy operations. | ||
| If there are any families that support graphics, there is at least one family that supports both compute and graphics. | ||
| The most capable family is always first in the `Adapter.queueFamilies` list. | ||
| (It will either be a graphics+compute, or at least compute) | ||
| Exposed adapters must have at least a compute family. | ||
| Users can also create `QueueFamily`s. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. What does this mean? There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Users shouldn't create QueueFamilies, they are an intrinsic part of the device and an only be queried (Vulkan). There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. In Vulkan there will always be one queue family that can do everything basic like graphics, transfer and compute.
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
@jdashg keeps saying that this is only true if there is a queue family that supports graphics at all. |
||
|
|
||
| ## Queue Creation | ||
|
|
||
| Queues are created at Device creation time, and are exposed via `sequence<Queue> Device.queues`. | ||
| Each element in `sequence<QueueFamily> WebGPUDeviceDescriptor.queueRequests` results in a corresponding element in `Device.queues`. | ||
| If an user-provided QueueFamily does not match a family in `Adapter.queueFamilies`, a more capable family is used. | ||
| If no available family can satisfy an user-provided family, device creation fails. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Will it be an error if the developer passes an empty sequence for |
||
|
|
||
| ## Synchronization | ||
|
|
||
| Queues can have fences inserted into them, and any queue can wait on that fence to be complete. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Some links: Metal
D3D12
VulkanFencesEventsSemaphoresThere was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Any queue cannot wait on a fence signalled on another queue... that will crash badly in Vulkan. For cross-queue synch you need a semaphore. |
||
|
|
||
| Synchronizing access to resources across queues should be done by telling a queue to wait until a fence is complete, followed by submitting command buffers that depend on that fence. | ||
| Resource data may have multiple concurrent readers, or exclusively one writer, at a time. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. how does this work w.r.t different types of read access? Say, one queue uses resource A as a vertex buffer, another queue uses it as a uniform buffer. Both only read, but neither knows about what exact state this resource is expected to be in. There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Big issue here, fences and events in vulkan cannot be signalled in one queue and waited upon in another. You need a VkSemaphore for that. |
||
| Resources may have multiple concurrent readers and writers, so long as all subranges satisfy many-read/single-write exclusion. [1] | ||
| If different commands would violate this exclusion, the implmentation injects synchronization. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. implmentation There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Again, make the distinction between:
|
||
| If a command is submitted that will never be able to synchronize for exclusion without subsequent user commands, that Submit is refused. | ||
|
|
||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I think I know what you're trying to get at here. But, in case I am misunderstanding, please provide an example that would get rejected by the API. |
||
| [1]: For MVP, we require reads/write exclusion at whole-resource granularity, instead of allowing subranges. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Or we can just say for MVP that we only expose a single queue? |
||
|
|
||
| ### How Is Implicit Synchronization Done | ||
|
|
||
| Implicit Synchronization is clearly possible via the degenerate approach of strict serialization. | ||
| Viability depends on how difficult it is to inject synchronization with minimal overhead, and in particular minimal excess serialization. | ||
|
|
||
| Each CommandBuffer knows all reads and writes to its resources, and satisfies reads/write exclusion internally. | ||
| (out of scope of multi-queue discussion) | ||
| Further, baked CommandBuffers know their required starting memory barrier requirements, as well as their end state. | ||
|
|
||
| Each Resource effectively has: | ||
|
|
||
| ~~~ | ||
| struct LastAccess { | ||
| CommandBuffer cb; | ||
| AccessBits access; | ||
| }; | ||
| list<LastAccess> last_accesses; | ||
| ~~~ | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. If this information is kept inside of each resource, how are we going to handle race conditions where one resource is used by multiple queues from multiple web workers? Should we make resources become read-only once they're transferred between workers? Even if we do, @kvark 's point about there being different types of reading (shader resource vs UAV, etc) still holds. D3D's D3D12_RESOURCE_FLAG_ALLOW_SIMULTANEOUS_ACCESS is relevant here. |
||
|
|
||
| Each CommandBuffer effectively has: | ||
|
|
||
| ~~~ | ||
| struct ResourceAccess { | ||
| Resource res; | ||
| AccessBits src_access; | ||
| AccessBits dst_access; | ||
| }; | ||
| list<ResourceAccess> res_accesses; | ||
| ~~~ | ||
|
|
||
| For a `submit(sequence<WebGPUCommandBuffer> buffers)`, implementations inject any required synchronization. | ||
| On submit, each CommandBuffer in turn traverses its resources and identifies any outstanding synchronization requirements. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 👍 There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. wait.. this means that you will require and expect submitted command buffers to execute in-order of submission!?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. What this is saying is that we (WebGPU implementation) will insert command buffers doing the memory barriers / resource transitions between some of those in a sequence. That doesn't serialize the execution (by GPU) of the command buffers more than the user would do in Vulkan. E.g. if one command buffer is writing into a UAV and another one is using the same resource as a vertex buffer, WebGPU implementation will insert the transition. The user would do the same in Vulkan/D3D12. There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Ok, so you will go on a quest to find dependency chains between the uses of resources? I hate to tell you that, but it is exactly what OpenGL has to do... and that makes creating a bugless and performant OpenGL implementation (driver) very very hard.
I think a dependency graph of the resources used in the command buffers you will be processing for submission will be an enormous DAG and quite a challenge to analyze. The Vulkan approach of having the user explicitly provide the transitions ahead of time and having them baked into the command buffer, while not trying to insert them at runtime is most probably the biggest reason (jointly with offline shader compilation) why Vulkan applications are seeing the often quoted 60% overall-CPU-utilization reductions on mobile devices even when multithreading and worker-thread command buffer generation is accounted for.
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Yes, that's pretty much where the group is heading at the moment.
I'm not as pessimistic. There aren't that many resources that change their usage through the frame: say, a hundred render targets plus a bunch of buffers we write as UAV. This isn't an enormous DAG to analyze.
The good thing for WebGPU (contrary to OpenGL) is that we have command buffers. So the barriers inside command buffers will also need to be computed only once (either at recording, or the first submission, depending on implementation). It's just the barriers between command buffers that we'll need to compute and insert on every submission at runtime. There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The whole point of the "render pass" system added in the latest D3D12 is to help track and validate these transitions, Metal already does this tracking implicitly, and it's now a common tactic for engines to track them it themselves as well ( see https://www.gdcvault.com/play/1024656/Advanced-Graphics-Tech-Moving-to and http://32ipi028l5q82yhj72224m8j.wpengine.netdna-cdn.com/wp-content/uploads/2017/03/GDC2017-D3D12-And-Vulkan-Done-Right.pdf ). There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. well yes, so since engines do it themselves.... maybe you should let them and make your lives easier?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. There is great deal of historical context to this discussion. Basically, since the browser implementations have to track the lifetime and usage of resources for validation anyway, it's not a big step to derive the barriers from this info. There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. lifetime resource tracking is easiest, instead of deleting (where the usual JS GC would do) put it on a list with an associated API fence inserted after last API use (or instead of where you'd want the delete)... I do that already. barriers are easy (in OpenGL you basically slam a glMemoryBarrier before the first read from a modified resource) begin/end dependencies are a bit harder, and I would really like to see an implementation that is able to insert them and validate them before command submit to virtual-queue (webGPU queue, not VK queue) time so we can benefit from pre-validated and pre-compiled command buffers like we go in D3D12 and VK. |
||
| (i.e. a CommandBuffer with a Texture read knows that some other previously-submitted CommandBuffer last wrote to that Texture, and establishes a dependency on this write) | ||
| Upon submission, each CommandBuffers tags its resources with the relevant synchronization info for later CommandBuffers to check against. | ||
|
|
||
| In Vulkan, synchronization injection takes the form of synthesizing VkSubmitInfos, synchronizing via VkSemaphores and VkPipelineStageFlags, as well as submitting synthesized CommandBuffers containing memory barriers and queue family transfers. | ||
| (Queue family transfers may require submitting synthesized CommandBuffers to other queues, as well) | ||
|
|
||
| Implementation may warn users about synchronization overhead. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Implementations |
||
|
|
||
| ## Further work | ||
|
|
||
| Overhead could be reduced by allowing users to provide access/memory barrier hints and queue transfer hints. | ||
| Providing bad hints may cause worse performance, which implementations should warn users about. | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -580,8 +580,11 @@ interface WebGPUFence { | |
|
|
||
| // Queue | ||
| interface WebGPUQueue { | ||
| readonly attribute WebGPUQueueFamily family; | ||
|
|
||
| void submit(sequence<WebGPUCommandBuffer> buffers); | ||
| WebGPUFence insertFence(); | ||
| void wait(WebGPUFence); | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Queues can wait but they can't signal?
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. I think this commit is missing a rebase. |
||
| }; | ||
|
|
||
| // SwapChain / RenderingContext | ||
|
|
@@ -615,6 +618,7 @@ interface WebGPUDevice { | |
| readonly attribute WebGPUExtensions extensions; | ||
| readonly attribute WebGPULimits limits; | ||
| readonly attribute WebGPUAdapter adapter; | ||
| readonly attribute sequence<WebGPUQueue> queues; | ||
|
|
||
| WebGPUBuffer createBuffer(WebGPUBufferDescriptor descriptor); | ||
| WebGPUTexture createTexture(WebGPUTextureDescriptor descriptor); | ||
|
|
@@ -635,23 +639,28 @@ interface WebGPUDevice { | |
| WebGPUCommandBuffer createCommandBuffer(WebGPUCommandBufferDescriptor descriptor); | ||
| WebGPUFence createFence(WebGPUFenceDescriptor descriptor); | ||
|
|
||
| WebGPUQueue getQueue(); | ||
|
|
||
| attribute WebGPULogCallback onLog; | ||
| WebGPUObjectStatusQuery getObjectStatus(StatusableObject statusableObject); | ||
| }; | ||
|
|
||
| interface WebGPUQueueFamily { | ||
| readonly attribute boolean graphics; | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 8 different possibilities? The DX12 style is easier to work with. |
||
| readonly attribute boolean compute; | ||
| readonly attribute boolean copy; | ||
| }; | ||
|
|
||
| dictionary WebGPUDeviceDescriptor { | ||
| WebGPUExtensions extensions; | ||
| //WebGPULimits limits; Don't expose higher limits for now. | ||
|
|
||
| // TODO are other things configurable like queues? | ||
| sequence<WebGLQueueFamily> queueRequests; | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. nit: WebGPUQueueFamily |
||
| }; | ||
|
|
||
| interface WebGPUAdapter { | ||
| readonly attribute DOMString name; | ||
| readonly attribute WebGPUExtensions extensions; | ||
| //readonly attribute WebGPULimits limits; Don't expose higher limits for now. | ||
| readonly attribute sequence<WebGLQueueFamily> queueFamilies; | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. ditto WebGPUQueueFamily |
||
|
|
||
| WebGPUDevice createDevice(WebGPUDeviceDescriptor descriptor); | ||
| }; | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
What is the benefit of exposing families? Why not make each queue independent, having its own capabilities?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Because multiple queues from the same family do not need to transfer ownership between each other
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
They also allow for less idling when complex execution dependencies are present
https://mynameismjp.wordpress.com/2018/06/17/breaking-down-barriers-part-3-multiple-command-processors/
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I don't think we've discussed the technique for ownership transfer. Depending on how that works, the implementation could be the one determining which barriers are issued, so it would know internally if the two queues happen to be in the same family and omit barriers accordingly.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Yes, but this transfer will be costly so the user has to have exposed queue family information, so that they can avoid using resources across queues from different families.