Sitelet https://github.com/tensorflow/tensorflow/pull/10636
Skip to content

Non-determinism Docs (#2732) - #10636

Closed
jkschin wants to merge 1 commit into
tensorflow:masterfrom
jkschin:non_determinism
Closed

jkschin wants to merge 1 commit into
tensorflow:masterfrom
jkschin:non_determinism

Conversation

@jkschin

@jkschin jkschin commented Jun 11, 2017 •

Copy link
Copy Markdown

Fixes #2732. Wrote a short tutorial on non-determinism in TensorFlow due to GPU reductions.

Wrote a short tutorial on non-determinism in TensorFlow due to GPU reductions.
@tensorflow-jenkins

Copy link
Copy Markdown
Collaborator

Can one of the admins verify this patch?

@martinwicke martinwicke added the awaiting review Pull request awaiting review label Jun 13, 2017
@michaelisard

Copy link
Copy Markdown

This looks fine to me, but I will pass to @zheng-xq in case there is a different best practice to ensure determinism when it is required.

@zheng-xq zheng-xq assigned ekelsen and unassigned zheng-xq Jun 15, 2017
@zheng-xq

Copy link
Copy Markdown
Contributor

Erich is working on a design to enable better deterministic support. Adding him to make sure the language is consistent with the direction we are taking.

@jkschin

jkschin commented Jun 21, 2017

Copy link
Copy Markdown
Author

@ekelsen any updates?

@ekelsen

ekelsen commented Jun 21, 2017

Copy link
Copy Markdown
Contributor

Reductions will soon be deterministic on the GPU without performance loss. For some types (float16, complex64, complex128, bool) the performance will increase by 30-1000x.

@ekelsen

ekelsen commented Jun 21, 2017

Copy link
Copy Markdown
Contributor

This is an awesome doc and thank you for taking the time to write it. The non-deterministic nature of the GPU reductions has been confusing people for a long time.

However, I'm not sure that it makes sense to pull this in for the brief period of time before the new reductions go in.

@jkschin

jkschin commented Jun 22, 2017

Copy link
Copy Markdown
Author

Sounds awesome, deterministic reductions with no performance loss. If determinism is guaranteed from then on, then perhaps it doesn't make sense to pull this in for the brief period.

It might make sense to still pull this in and add more explanation on how your side has made it deterministic with no performance loss, for those who intend to understand it.

Feel free to close this PR if it's the best option at this point.

@drpngx

drpngx commented Jun 26, 2017

Copy link
Copy Markdown
Contributor

@ekelsen I am closing this PR since it's going to be obsolete pretty soon. Feel free to re-open if you think it's worth having in the meantime.
@tfboyd just FYI
@jkschin nice doc! I wish we had earlier :-)

@drpngx drpngx closed this Jun 26, 2017
@jkschin

jkschin commented Jun 27, 2017

Copy link
Copy Markdown
Author

@drpngx cheers! Shall look for the next PR to do :)

@alexklibisz

Copy link
Copy Markdown

@ekelsen Hi, do you maybe have an update for when the deterministic functionality will be available? Is it something implemented in TF itself or a feature of CUDNN? Thank you

@ekelsen

ekelsen commented Sep 13, 2017

Copy link
Copy Markdown
Contributor

Reductions are deterministic now. l2loss is deterministic. softmax will be shortly. cross entropy sometime soon.

@ekelsen

ekelsen commented Sep 13, 2017

Copy link
Copy Markdown
Contributor

The backward pass of the convolutions is not deterministic, but this basically NVIDIA's fault. There is possibly to make it deterministic, but it is 3x slower.

@duncanriach

Copy link
Copy Markdown
Contributor

For up-to-date status of work on making TensorFlow operate deterministically on GPUs, please see the following repo: https://github.com/NVIDIA/tensorflow-determinism.

copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476
PiperOrigin-RevId: 617425257
copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476
PiperOrigin-RevId: 617425257
copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476
PiperOrigin-RevId: 617425257
copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476
PiperOrigin-RevId: 617496055
copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

FUTURE_COPYBARA_INTEGRATE_REVIEW=openxla/xla#10636 from zhenying-liu:offloading/annotator 11148fe63588f58c85ab2530bda64807bf400476
PiperOrigin-RevId: 617496055
copybara-service Bot pushed a commit that referenced this pull request Mar 20, 2024
Imported from GitHub PR openxla/xla#10636

Add stream id in the backend_config of copy-start instruction. The stream id is obtained from hlo_query::NextChannelId().
The corresponding copy-done instruction which is the use of copy-start instruction will be traversed and added the stream id in the backend_config too. This part is automatically done by the subsequent AnnotateStreamAttributesForUsers() existing in the function.
The bool data member copy_start_done_ is used to differentiate copy-start/copy-done from other collective instructions and go through two different paths.
openxla/xla#10450 is split and the current PR is the first 1 out of 3 PRs.
Copybara import of the project:

--
ff99c161a634b868d4265204d10d0b80adf2e772 by Jane Liu <janeliu@nvidia.com>:

Add annotation of stream id for copy-start and its use of copy-done instruction

--
64c746a5de6aa4c9370034ed712192dfd78e3a6e by Jane Liu <janeliu@nvidia.com>:

Enable the annotator for copy-start/copy-done in gpu compiler

--
7972b987e45cd165834483c4350e953094b5dbe8 by Jane Liu <janeliu@nvidia.com>:

Add the annotator pass after HLO rematerialization pass

--
2a9e317aac85e4de3f0ff5f1558885408ab4562d by Jane Liu <janeliu@nvidia.com>:

Add the dependency in BUILD

--
db5ed79f0ec90a1032252beff5108e4f01b61c4a by Jane Liu <janeliu@nvidia.com>:

Use a function to annotate copy-start and add description

--
b4b3932bfedfabc7f3e22a8766ecbc1fa1188402 by Jane Liu <janeliu@nvidia.com>:

remove the bool var copy_start from the StreamAttributeAnnotator class

--
11148fe63588f58c85ab2530bda64807bf400476 by Jane Liu <janeliu@nvidia.com>:

Fixes according to the code review

Merging this change closes #10636

PiperOrigin-RevId: 617550093
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review Pull request awaiting review cla: yes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mention that GPU reductions are nondeterministic in docs

10 participants