Sitelet https://github.com/gpuweb/gpuweb/pull/95/files
Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 83 additions & 0 deletions design/MultipleQueues.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# Multiple Queues

Multiple queues is important for:

- Explicit submission onto less capable queues, such as pure copy queues.
- Explicit distribution of submissions on machines with multiple queues of the same family.

## QueueFamily

There are multiple families of queues.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the benefit of exposing families? Why not make each queue independent, having its own capabilities?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Because multiple queues from the same family do not need to transfer ownership between each other

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They also allow for less idling when complex execution dependencies are present
https://mynameismjp.wordpress.com/2018/06/17/breaking-down-barriers-part-3-multiple-command-processors/

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we've discussed the technique for ownership transfer. Depending on how that works, the implementation could be the one determining which barriers are issued, so it would know internally if the two queues happen to be in the same family and omit barriers accordingly.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, but this transfer will be costly so the user has to have exposed queue family information, so that they can avoid using resources across queues from different families.

Availible families of queues are surfaced via `sequence<QueueFamily> Adapter.queueFamilies`.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

type: Availible

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

typo: type 🤣

Support for `graphics`, `compute`, and `copy` operations are surfaced by the respective attributes on the QueueFamily.
Queue families with `graphics` or `compute` will always have `copy`.
There may be queue families in the future like D3D12's video decode queues, which do not support copy operations.
If there are any families that support graphics, there is at least one family that supports both compute and graphics.
The most capable family is always first in the `Adapter.queueFamilies` list.
(It will either be a graphics+compute, or at least compute)
Exposed adapters must have at least a compute family.
Users can also create `QueueFamily`s.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this mean?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Users shouldn't create QueueFamilies, they are an intrinsic part of the device and an only be queried (Vulkan).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In Vulkan there will always be one queue family that can do everything basic like graphics, transfer and compute.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In Vulkan there will always be one queue family that can do everything basic like graphics, transfer and compute.

@jdashg keeps saying that this is only true if there is a queue family that supports graphics at all.


## Queue Creation

Queues are created at Device creation time, and are exposed via `sequence<Queue> Device.queues`.
Each element in `sequence<QueueFamily> WebGPUDeviceDescriptor.queueRequests` results in a corresponding element in `Device.queues`.
If an user-provided QueueFamily does not match a family in `Adapter.queueFamilies`, a more capable family is used.
If no available family can satisfy an user-provided family, device creation fails.

@RafaelCintron RafaelCintron Oct 29, 2018 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will it be an error if the developer passes an empty sequence for queueRequest?


## Synchronization

Queues can have fences inserted into them, and any queue can wait on that fence to be complete.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any queue cannot wait on a fence signalled on another queue... that will crash badly in Vulkan.

For cross-queue synch you need a semaphore.


Synchronizing access to resources across queues should be done by telling a queue to wait until a fence is complete, followed by submitting command buffers that depend on that fence.
Resource data may have multiple concurrent readers, or exclusively one writer, at a time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how does this work w.r.t different types of read access? Say, one queue uses resource A as a vertex buffer, another queue uses it as a uniform buffer. Both only read, but neither knows about what exact state this resource is expected to be in.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Big issue here, fences and events in vulkan cannot be signalled in one queue and waited upon in another.

You need a VkSemaphore for that.

Resources may have multiple concurrent readers and writers, so long as all subranges satisfy many-read/single-write exclusion. [1]
If different commands would violate this exclusion, the implmentation injects synchronization.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

implmentation

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again, make the distinction between:

  • Pipeline Barrier
  • Fence
  • Event
  • Semaphore

If a command is submitted that will never be able to synchronize for exclusion without subsequent user commands, that Submit is refused.

@RafaelCintron RafaelCintron Oct 29, 2018 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I know what you're trying to get at here. But, in case I am misunderstanding, please provide an example that would get rejected by the API.

[1]: For MVP, we require reads/write exclusion at whole-resource granularity, instead of allowing subranges.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Or we can just say for MVP that we only expose a single queue?


### How Is Implicit Synchronization Done

Implicit Synchronization is clearly possible via the degenerate approach of strict serialization.
Viability depends on how difficult it is to inject synchronization with minimal overhead, and in particular minimal excess serialization.

Each CommandBuffer knows all reads and writes to its resources, and satisfies reads/write exclusion internally.
(out of scope of multi-queue discussion)
Further, baked CommandBuffers know their required starting memory barrier requirements, as well as their end state.

Each Resource effectively has:

~~~
struct LastAccess {
CommandBuffer cb;
AccessBits access;
};
list<LastAccess> last_accesses;
~~~

@RafaelCintron RafaelCintron Oct 29, 2018 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this information is kept inside of each resource, how are we going to handle race conditions where one resource is used by multiple queues from multiple web workers? Should we make resources become read-only once they're transferred between workers? Even if we do, @kvark 's point about there being different types of reading (shader resource vs UAV, etc) still holds. D3D's D3D12_RESOURCE_FLAG_ALLOW_SIMULTANEOUS_ACCESS is relevant here.


Each CommandBuffer effectively has:

~~~
struct ResourceAccess {
Resource res;
AccessBits src_access;
AccessBits dst_access;
};
list<ResourceAccess> res_accesses;
~~~

For a `submit(sequence<WebGPUCommandBuffer> buffers)`, implementations inject any required synchronization.
On submit, each CommandBuffer in turn traverses its resources and identifies any outstanding synchronization requirements.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wait.. this means that you will require and expect submitted command buffers to execute in-order of submission!?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What this is saying is that we (WebGPU implementation) will insert command buffers doing the memory barriers / resource transitions between some of those in a sequence. That doesn't serialize the execution (by GPU) of the command buffers more than the user would do in Vulkan.

E.g. if one command buffer is writing into a UAV and another one is using the same resource as a vertex buffer, WebGPU implementation will insert the transition. The user would do the same in Vulkan/D3D12.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, so you will go on a quest to find dependency chains between the uses of resources?

I hate to tell you that, but it is exactly what OpenGL has to do... and that makes creating a bugless and performant OpenGL implementation (driver) very very hard.
You'd have to do amazing work on par with the engineers at Nvidia and AMD to both:

  1. insert the transitions correctly to not get undefined behaviour, artefacts, other problems
  2. not insert too many to get subpar performance and serrialization
  3. do it fast

I think a dependency graph of the resources used in the command buffers you will be processing for submission will be an enormous DAG and quite a challenge to analyze.

The Vulkan approach of having the user explicitly provide the transitions ahead of time and having them baked into the command buffer, while not trying to insert them at runtime is most probably the biggest reason (jointly with offline shader compilation) why Vulkan applications are seeing the often quoted 60% overall-CPU-utilization reductions on mobile devices even when multithreading and worker-thread command buffer generation is accounted for.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, that's pretty much where the group is heading at the moment.

I think a dependency graph of the resources used in the command buffers you will be processing for submission will be an enormous DAG and quite a challenge to analyze.

I'm not as pessimistic. There aren't that many resources that change their usage through the frame: say, a hundred render targets plus a bunch of buffers we write as UAV. This isn't an enormous DAG to analyze.

The Vulkan approach of having the user explicitly provide the transitions ahead of time and having them baked into the command buffer, while not trying to insert them at runtime

The good thing for WebGPU (contrary to OpenGL) is that we have command buffers. So the barriers inside command buffers will also need to be computed only once (either at recording, or the first submission, depending on implementation). It's just the barriers between command buffers that we'll need to compute and insert on every submission at runtime.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The whole point of the "render pass" system added in the latest D3D12 is to help track and validate these transitions, Metal already does this tracking implicitly, and it's now a common tactic for engines to track them it themselves as well ( see https://www.gdcvault.com/play/1024656/Advanced-Graphics-Tech-Moving-to and http://32ipi028l5q82yhj72224m8j.wpengine.netdna-cdn.com/wp-content/uploads/2017/03/GDC2017-D3D12-And-Vulkan-Done-Right.pdf ).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

well yes, so since engines do it themselves.... maybe you should let them and make your lives easier?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is great deal of historical context to this discussion. Basically, since the browser implementations have to track the lifetime and usage of resources for validation anyway, it's not a big step to derive the barriers from this info.

@devshgraphicsprogramming devshgraphicsprogramming Nov 16, 2018 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lifetime resource tracking is easiest, instead of deleting (where the usual JS GC would do) put it on a list with an associated API fence inserted after last API use (or instead of where you'd want the delete)... I do that already.

barriers are easy (in OpenGL you basically slam a glMemoryBarrier before the first read from a modified resource)

begin/end dependencies are a bit harder, and I would really like to see an implementation that is able to insert them and validate them before command submit to virtual-queue (webGPU queue, not VK queue) time so we can benefit from pre-validated and pre-compiled command buffers like we go in D3D12 and VK.

(i.e. a CommandBuffer with a Texture read knows that some other previously-submitted CommandBuffer last wrote to that Texture, and establishes a dependency on this write)
Upon submission, each CommandBuffers tags its resources with the relevant synchronization info for later CommandBuffers to check against.

In Vulkan, synchronization injection takes the form of synthesizing VkSubmitInfos, synchronizing via VkSemaphores and VkPipelineStageFlags, as well as submitting synthesized CommandBuffers containing memory barriers and queue family transfers.
(Queue family transfers may require submitting synthesized CommandBuffers to other queues, as well)

Implementation may warn users about synchronization overhead.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Implementations


## Further work

Overhead could be reduced by allowing users to provide access/memory barrier hints and queue transfer hints.
Providing bad hints may cause worse performance, which implementations should warn users about.
15 changes: 12 additions & 3 deletions design/sketch.webidl
Original file line number Diff line number Diff line change
Expand Up @@ -580,8 +580,11 @@ interface WebGPUFence {

// Queue
interface WebGPUQueue {
readonly attribute WebGPUQueueFamily family;

void submit(sequence<WebGPUCommandBuffer> buffers);
WebGPUFence insertFence();
void wait(WebGPUFence);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Queues can wait but they can't signal?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this commit is missing a rebase.

};

// SwapChain / RenderingContext
Expand Down Expand Up @@ -615,6 +618,7 @@ interface WebGPUDevice {
readonly attribute WebGPUExtensions extensions;
readonly attribute WebGPULimits limits;
readonly attribute WebGPUAdapter adapter;
readonly attribute sequence<WebGPUQueue> queues;

WebGPUBuffer createBuffer(WebGPUBufferDescriptor descriptor);
WebGPUTexture createTexture(WebGPUTextureDescriptor descriptor);
Expand All @@ -635,23 +639,28 @@ interface WebGPUDevice {
WebGPUCommandBuffer createCommandBuffer(WebGPUCommandBufferDescriptor descriptor);
WebGPUFence createFence(WebGPUFenceDescriptor descriptor);

WebGPUQueue getQueue();

attribute WebGPULogCallback onLog;
WebGPUObjectStatusQuery getObjectStatus(StatusableObject statusableObject);
};

interface WebGPUQueueFamily {
readonly attribute boolean graphics;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

8 different possibilities? The DX12 style is easier to work with.

readonly attribute boolean compute;
readonly attribute boolean copy;
};

dictionary WebGPUDeviceDescriptor {
WebGPUExtensions extensions;
//WebGPULimits limits; Don't expose higher limits for now.

// TODO are other things configurable like queues?
sequence<WebGLQueueFamily> queueRequests;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: WebGPUQueueFamily

};

interface WebGPUAdapter {
readonly attribute DOMString name;
readonly attribute WebGPUExtensions extensions;
//readonly attribute WebGPULimits limits; Don't expose higher limits for now.
readonly attribute sequence<WebGLQueueFamily> queueFamilies;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ditto WebGPUQueueFamily


WebGPUDevice createDevice(WebGPUDeviceDescriptor descriptor);
};
Expand Down