Repository navigation
Conversation
Wrote a short tutorial on non-determinism in TensorFlow due to GPU reductions.
|
Can one of the admins verify this patch? |
|
This looks fine to me, but I will pass to @zheng-xq in case there is a different best practice to ensure determinism when it is required. |
|
Erich is working on a design to enable better deterministic support. Adding him to make sure the language is consistent with the direction we are taking. |
|
@ekelsen any updates? |
|
Reductions will soon be deterministic on the GPU without performance loss. For some types (float16, complex64, complex128, bool) the performance will increase by 30-1000x. |
|
This is an awesome doc and thank you for taking the time to write it. The non-deterministic nature of the GPU reductions has been confusing people for a long time. However, I'm not sure that it makes sense to pull this in for the brief period of time before the new reductions go in. |
|
Sounds awesome, deterministic reductions with no performance loss. If determinism is guaranteed from then on, then perhaps it doesn't make sense to pull this in for the brief period. It might make sense to still pull this in and add more explanation on how your side has made it deterministic with no performance loss, for those who intend to understand it. Feel free to close this PR if it's the best option at this point. |
|
@drpngx cheers! Shall look for the next PR to do :) |
|
@ekelsen Hi, do you maybe have an update for when the deterministic functionality will be available? Is it something implemented in TF itself or a feature of CUDNN? Thank you |
|
Reductions are deterministic now. l2loss is deterministic. softmax will be shortly. cross entropy sometime soon. |
|
The backward pass of the convolutions is not deterministic, but this basically NVIDIA's fault. There is possibly to make it deterministic, but it is 3x slower. |
|
For up-to-date status of work on making TensorFlow operate deterministically on GPUs, please see the following repo: https://github.com/NVIDIA/tensorflow-determinism. |
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476 PiperOrigin-RevId: 617425257
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476 PiperOrigin-RevId: 617425257
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476 PiperOrigin-RevId: 617425257
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476 PiperOrigin-RevId: 617496055
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476 PiperOrigin-RevId: 617496055
Imported from GitHub PR openxla/xla#10636 Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId(). The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function. The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths. openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs. Copybara import of the project: -- ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>: Add annotation of stream id for copy-start and its use of copy-done instruction -- 64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>: Enable the annotator for copy-start/copy-done in gpu compiler -- 7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>: Add the annotator pass after HLO rematerialization pass -- 2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>: Add the dependency in BUILD -- db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>: Use a function to annotate copy-start and add description -- b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>: remove the bool var copy_start from the StreamAttributeAnnotator class -- 11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>: Fixes according to the code review Merging this change closes #10636 PiperOrigin-RevId: 617550093
Fixes #2732. Wrote a short tutorial on non-determinism in TensorFlow due to GPU reductions.