New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
test_svd_dense_opencl fails on Ubuntu 20.04 LTS using NVIDIA OpenCL on AWS g3s.xlarge instance #3147
Comments
|
@christopher-gee Can you please run the test using gdb and share the stack trace. Error from issue description indicates that is not accuracy error but routine failed to launch totally. Also, kindly run the program with the following environment variables set to see more detailed info from the library. export AF_TRACE=all
export AF_PRINT_ERRORS=1Update: I just noticed your comment, It is odd that no additional output was produced with these set. |
|
While running tests for another PR, I encountered exactly the same failure
as you described @christopher-gee.
@9prady9
The error persisted independent of the used library: oneMKL vs fftw3 &
openblas. (windows10 & MSVC)
You can find the log file with (AF_TRACE=all and AF_PRINT_ERRORS=1)
attached.
I also noticed some strange behavior for the first time in the compiler log.
I would expect the CL_TARGET_OPENCL_VERSION to be 120, although it is
somewhere not defined resulting in version 220. (also included)
I think this is collateral damage of the recent changes done on some
"CMakeLists" files
For completeness I also include clinfo:
- AMD Driver supports OCL 2.1 / AMD Device support OCL 2.0
- CUDA 11.2 support OCL 3.0 / GTX 750 Ti support OCL 1.2
-
[log openblas.txt](https://github.com/arrayfire/arrayfire/files/6689481/log.openblas.txt)
[begin of compiler log.txt](https://github.com/arrayfire/arrayfire/files/6689484/begin.of.compiler.log.txt)
[clinfo.txt](https://github.com/arrayfire/arrayfire/files/6689486/clinfo.txt)
|
|
@willyborn I will look into it this week, thanks for the logs. Update: Note that, CPU backend compute upstreams won't have any effect on OpenCL backend. OpenCL backend uses combination of custom kernels (magma interface etc) + CLBLast I think for LAPACK routines. So, change in oneMKL/fftw3/openblas isn't expected to change this, otherwise that would be a major bug :) |
|
I was unable to reproduce the svd_dense failure on windows so far both in debug and relwithdebinfo configurations using MSVC toolchain. Nevertheless, will try to reproduce on linux if not then will have to setup ubuntu 20.04 based docker image. @willyborn We do set I think they are related to CLBlast project and we don't have control over it. Nevertheless, our installers use OpenCL 1.2 only so they won't cause issues on any user systems. If you are building on your target system though, that is completely different scenario. In that case, the user who is building has to take care of such differences. ArrayFire codebase doesn't use any OpenCL 2.* features. Therefore, for ArrayFire, it doesn't matter which version of OpenCL as long as it is greater than 1.2. @willyborn How recent is your oneMKL installation on windows ? I am curious because I have the one that seems to be current and I cannot reproduce failure of svd_dense with that compute upstream. Also please note that, on Windows, with openblas you need to provide lapacke lib also for LA routines and often this is not realiably available on windows and ArrayFire builds without LA support without complaining during cmake stage (needs to be fixed soon, on our to do list). I wonder if your failure while using openblas is due to missing LA routine support. |
|
The oneMKL dates from 17/june. From the package, only the MKL module is installed.
The tests are run in RelWithDebInfo (which is the std config)
In my cmake configuration I am not doing something special for LA (I even
do not know what it should be).
I include the windows bat file using the cmake:
`
SET COMP_C_RELEASE_FLAGS="/MD /O2 /DNDEBUG /MP /GS- /GL /arch:AVX"
SET COMP_CXX_RELEASE_FLAGS="/MD /O2 /DNDEBUG /MP /GS- /GL /GT /arch:AVX"
SET LINK_EXE_RELEASE_FLAGS="/INCREMENTAL:NO /LTCG:incremental"
SET LINK_MODULE_RELEASE_FLAGS="/LTCG:incremental"
SET LINK_STATIC_RELEASE_FLAGS="/LTCG:incremental"
REM REFUSE TO DELETE, when ONLY compile options are changed. Do REBUILD
afterwards however.
RD /S build
MKDIR build
CD build
REM oneAPI/mkl based --> vcpkg.json dependencies "intel-mkl" package
CALL "%ONEAPI_ROOT%\setvars.bat"
cmake -DAF_BUILD_CPU=ON -DAF_BUILD_CUDA=ON -DAF_BUILD_OPENCL=ON
-DAF_BUILD_UNIFIED=OFF -Wno-dev -G "Visual Studio 15 2017 Win64" -Thost=x64 -DUSE_CPU_MKL=1 -DUSE_OPENCL_MKL=1
-DCMAKE_C_FLAGS_RELEASE=%COMP_C_RELEASE_FLAGS%
-DCMAKE_CXX_FLAGS_RELEASE=%COMP_CXX_RELEASE_FLAGS%
-DCMAKE_EXE_LINKER_FLAGS_RELEASE=%LINK_EXE_RELEASE_FLAGS%
-DCMAKE_MODULE_LINKER_FLAGS_RELEASE=%LINK_MODULE_RELEASE_FLAGS%
-DCMAKE_SHARED_LINKER_FLAGS_RELEASE=%LINK_MODULE_RELEASE_FLAGS%
-DCMAKE_STATIC_LINKER_FLAGS_RELEASE=%LINK_STATIC_RELEASE_FLAGS%
-DCMAKE_TOOLCHAIN_FILE="%VCPKG_DIR%/scripts/buildsystems/vcpkg.cmake" ..
@Rem end oneAPI/mkl
@Rem fftw3, openblas based --> vcpkg.json dependencies "fftw3", "openblas"
@Rem cmake -DAF_BUILD_CPU=ON -DAF_BUILD_CUDA=ON -DAF_BUILD_OPENCL=ON
-DAF_BUILD_UNIFIED=OFF -Wno-dev -G "Visual Studio 15 2017 Win64" -Thost=x64
-DCMAKE_C_FLAGS_RELEASE=%COMP_C_RELEASE_FLAGS%
-DCMAKE_CXX_FLAGS_RELEASE=%COMP_CXX_RELEASE_FLAGS%
-DCMAKE_EXE_LINKER_FLAGS_RELEASE=%LINK_EXE_RELEASE_FLAGS%
-DCMAKE_MODULE_LINKER_FLAGS_RELEASE=%LINK_MODULE_RELEASE_FLAGS%
-DCMAKE_SHARED_LINKER_FLAGS_RELEASE=%LINK_MODULE_RELEASE_FLAGS%
-DCMAKE_STATIC_LINKER_FLAGS_RELEASE=%LINK_STATIC_RELEASE_FLAGS%
-DCMAKE_TOOLCHAIN_FILE="%VCPKG_DIR%/scripts/buildsystems/vcpkg.cmake" ..
@Rem end fftw3, openblas
CD ..
`
PS: I 'M STILL SEARCHING for a problem in CUDA, where reduce_first_kernel is wrongly called when using the LTCG (LTO on unix). The usage of "scalar<To>(nanval)" performs memory overwrites, resulting in wrong function called and exceptions later.
OPENCL, compiles fine and gives an remarkable extra performance boost.
As soon as I find the cause, I will create a new PR.
|
|
I have debugged this issue on several occasions without any resolutions. The problem is that one of the lapack functions are failing for no determinable reason. Additionally if you change the order in which the tests run (for example run doubles before floats in the svd test) the tests pass. I suspect is an internal state issue but I cannot isolate the issue. I have even tried to run this code with address sanitizers and other debugging utilities. Any help regarding this test is appreciated. |
Test 340
test_svd_dense_openclfails with ArrayFire compiled on Ubuntu 20.04 LTS using NVIDIA OpenCL on AWS g3s.xlarge instance.Description
[ArrayFire was built from source code on AWS EC2 g3s instance (NVIDIA Tesla M60 GPU)]
[OpenCL]
[No]
[Yes]
[Expect that Test 340 will pass successfully]
variables set.
[Variables were set, but no additional output was produced]
Reproducible Code and/or Steps
System Information
[Master Branch, commit 9738a31]
[NVIDIA Tesla M60 GPU]
Linux:
Checklist
The text was updated successfully, but these errors were encountered: