Tensorflow 2.21 and related changes - #10789
Conversation
|
A new Pull Request was created by @smuzaffar for branch IB/CMSSW_20_1_X/master. @akritkbehera, @cmsbuild, @iarspider, @raoatifshad, @smuzaffar can you please review it and eventually sign? Thanks. |
|
cms-bot internal usage |
|
please build |
|
-1 Failed Tests: Build The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic: You can see more details here: Failed BuildI found compilation error when building: /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc14/external/gcc/14.3.1-724da22786638848892aa9ded8fcd995/bin/../lib/gcc/x86_64-redhat-linux-gnu/14.3.1/../../../../x86_64-redhat-linux-gnu/bin/ld.bfd: :(.text+0x742): undefined reference to `__tfcompile_tfaot_model_test_simple_bs1_Constant_2_constant_buffer_contents' /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc14/external/gcc/14.3.1-724da22786638848892aa9ded8fcd995/bin/../lib/gcc/x86_64-redhat-linux-gnu/14.3.1/../../../../x86_64-redhat-linux-gnu/bin/ld.bfd: :(.text+0x75b): undefined reference to `__tfcompile_tfaot_model_test_simple_bs1_Constant_3_constant_buffer_contents' /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc14/external/gcc/14.3.1-724da22786638848892aa9ded8fcd995/bin/../lib/gcc/x86_64-redhat-linux-gnu/14.3.1/../../../../x86_64-redhat-linux-gnu/bin/ld.bfd: :(.text+0x774): undefined reference to `__tfcompile_tfaot_model_test_simple_bs1_Constant_4_constant_buffer_contents' /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc14/external/gcc/14.3.1-724da22786638848892aa9ded8fcd995/bin/../lib/gcc/x86_64-redhat-linux-gnu/14.3.1/../../../../x86_64-redhat-linux-gnu/bin/ld.bfd: :(.text+0x790): undefined reference to `__tfcompile_tfaot_model_test_simple_bs1_Constant_5_constant_buffer_contents' /data/cmsbld/jenkins/workspace/ib-run-pr-tests/testBuildDir/el9_amd64_gcc14/external/gcc/14.3.1-724da22786638848892aa9ded8fcd995/bin/../lib/gcc/x86_64-redhat-linux-gnu/14.3.1/../../../../x86_64-redhat-linux-gnu/bin/ld.bfd: :(.text+0x7af): undefined reference to `__tfcompile_tfaot_model_test_simple_bs1_Constant_6_constant_buffer_contents' collect2: error: ld returned 1 exit status >> Deleted: tmp/el9_amd64_gcc14/src/PhysicsTools/TensorFlowAOT/test/testTFAOTInterface/testTFAOTInterface gmake: *** [tmp/el9_amd64_gcc14/src/PhysicsTools/TensorFlowAOT/test/testTFAOTInterface/testTFAOTInterface] Error 1 >> Leaving Package PhysicsTools/TensorFlowAOT >> Package PhysicsTools/TensorFlowAOT built >> Entering Package RecoTracker/MkFitCMS |
|
please test with cms-sw/cmssw#51713 |
|
enable gpu |
|
-1 Failed Tests: RelVals-AMD_MI300X DAS Queries: The DAS query tests failed, see the summary page for details. Failed RelVals-AMD_MI300X
Comparison SummarySummary:
AMD_W7900 Comparison SummaryThere are some workflows for which there are errors in the baseline: Summary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_H100@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_L40S@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
please test with cms-sw/cmssw#51713 |
|
-1 Failed Tests: RelVals-AMD_W7900 The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic:
You can see more details here: DAS Queries: The DAS query tests failed, see the summary page for details. Failed RelVals-AMD_W7900
Comparison SummarySummary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_H100@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_L40S@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
please test with cms-sw/cmssw#51713 |
|
-1 Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-69da77/55447/summary.html DAS Queries: The DAS query tests failed, see the summary page for details. Comparison SummarySummary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_H100@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_L40S@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
please test with cms-sw/cmssw#51713 |
|
-1 Failed Tests: UnitTests RelVals-AMD_MI300X The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic:
You can see more details here: Failed Unit TestsI found 14 errors in the following unit tests: ---> test testAlignmentStats had ERRORS ---> test DiMuonVertex had ERRORS ---> test DiElectronVertex had ERRORS and more ... Failed RelVals-AMD_MI300X
Comparison SummarySummary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_H100@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_L40S@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
please test with cms-sw/cmssw#51713 |
|
-1 Failed Tests: UnitTests The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic: You can see more details here: Failed Unit TestsI found 15 errors in the following unit tests: ---> test testAlignmentStats had ERRORS ---> test SagittaBiasNtuplizer had ERRORS ---> test DiMuonVertex had ERRORS and more ... Comparison SummarySummary:
AMD_MI300X Comparison SummaryThere are some workflows for which there are errors in the baseline: Summary:
AMD_W7900 Comparison SummaryThere are some workflows for which there are errors in the baseline: Summary:
NVIDIA_H100 Comparison SummarySummary:
NVIDIA_L40S Comparison SummarySummary:
NVIDIA_T4 Comparison SummarySummary:
Max Memory Comparisons exceeding threshold@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_H100@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_L40S@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
Max Memory Comparisons exceeding threshold NVIDIA_T4@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold: Expand to see workflows ...
|
|
+externals good to go in 20.1.X IBs along with cms-sw/cmssw#51713 |
|
This pull request is fully signed and it will be integrated in one of the next IB/CMSSW_20_1_X/master IBs (but tests are reportedly failing). This pull request will now be reviewed by the release team before it's merged. @ftenchini, @mandrenguyen, @sextonkennedy (and backports should be raised in the release meeting by the corresponding L2) |
|
NOTE that this PR and cms-sw/cmssw#51713 should go togather in same IB |
This backports Tensorflow 2.21.0 from CMSSW_20_1_TF_X IBs. This includes following changes
This also requires CMSSW update for Protobuf