Adding usage tip for ignore cpu affinity #1948
Merged
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Details
Do not mention proprietary info or link to internal work items in this PR.
Work item: "Internal", or link to GitHub issue (if applicable).
https://ontrack-internal.amd.com/browse/LWPCLPAT-559
What were the changes?
One sentence describing the work done.
Added recommendation for setting NCCL_IGNORE_CPU_AFFINITY for multi-node in usage tips
Why were the changes made?
Explain the motivation behind the work. Provide any publicly-available historical context.
It has been observed that NCCL_IGNORE_CPU_AFFINITY=1 improves performance over baseline
OpenMPI-5 on multi-node scales
How was the outcome achieved?
Technical details behind the work. Explain any publicly-available hardware peculiarities.
Comparison of performance between OpenMPI-4 and OpenMPI-5 with NCCL_IGNORE_CPU_AFFINITY on / off
Additional Documentation:
What else should the reviewer know?
Only documentation was added with no change to the default setting or no code modifications
Approval Checklist
Do not approve until these items are satisfied.