Repository navigation
[BUG] Getting STATUS_STACK_BUFFER_OVERRUN on any operations #285
Description
Activity
* (actually checking this now; regardless, I believe the fact that there is no human-readable error in any mode and instead it crashes internally, is a bug)UPD: updating driver helped, so at least that's good.
I've noticed that docs don't seem to match the reality. The docs on Backend suggest that default backend / first choice would be OpenCL, falling back to others.
@RReverser Can you please point me to this location in docs.
UPD: about driver.
I am confused, what do you mean ? Does updating driver resolved the problem ? or It helped with partially but not completely - if so, can you please share how far it did help the program to progress ?
Are you a Rust only developer ? If not, can you please try C++ helloworld example from ArrayFire and let me know how that runs on your system.
@RReverser Can you please point me to this location in docs.
It's here: https://docs.rs/arrayfire/3.8.0/arrayfire/enum.Backend.html.
Default backend order: OpenCL -> CUDA -> CPU
I am confused, what do you mean ? Does updating driver resolved the problem ?
Yes, it resolved the problem and ArrayFire works now (so I can't repro with C++ either), but as I said above:
regardless, I believe the fact that there is no human-readable error in any mode and instead it crashes internally, is a bug
Surely, there has to be a way to print a human-readable error in case of incompatible driver version or something like that instead of crashing.
It's here: https://docs.rs/arrayfire/3.8.0/arrayfire/enum.Backend.html.
Thanks, I will correct that.
Yes, it resolved the problem and ArrayFire works now (so I can't repro with C++ either), but as I said above:
The problem is resolved with driver update, cool.
Surely, there has to be a way to print a human-readable error in case of incompatible driver version or something like that instead of crashing.
We do some checks for driver and cuda runtime compatibility and log them too. I wonder if they are captured by cargo output 🤔
https://github.com/arrayfire/arrayfire/blob/master/src/backend/cuda/device_manager.cpp#L459We do some checks for driver and cuda runtime compatibility and log them too. I wonder if they are captured by cargo output 🤔
https://github.com/arrayfire/arrayfire/blob/master/src/backend/cuda/device_manager.cpp#L459I don't think it was, and either way it's probably worth encoding them as a runtime panic. Anything's better than a stack buffer overrun indicating a memory corruption in apparently safe Rust code.
We do some checks for driver and cuda runtime compatibility and log them too. I wonder if they are captured by cargo output thinking
https://github.com/arrayfire/arrayfire/blob/master/src/backend/cuda/device_manager.cpp#L459I don't think it was, and either way it's probably worth encoding them as a runtime panic. Anything's better than a stack buffer overrun indicating a memory corruption in apparently safe Rust code.
I am not really sure if the cause of that is arrayfire code base. Can you tell me your driver version when it caused issue.
Can you tell me your driver version when it caused issue.
It's in the report above: 452.66
Can you tell me your driver version when it caused issue.
It's in the report above: 452.66
Okay, thank you. I will check with that driver version and see if I can reproduce the issue. If I can, I shall move this issue upstream to address it correctly.
@RReverser I just realized that you are trying to use CUDA 11.2 based ArrayFire with driver 450 series. CUDA 11.2 requires 460.82 minimum. It is not a bug rather, wrong driver version was being used with ArrayFire that was built with CUDA 11.2
Okay. I still think that it should provide better error message on version mismatch, but as it's not affecting me personally, I don't have strong opinion on this.
That is true, it should give an error message rather than a silent seg fault. I was only letting you know that incorrect driver was the reason which I didn't realize until today when I was going through the conversation again.
Reacted by Ingvar StepanyanI was only letting you know that incorrect driver was the reason
Oh yeah, I assumed it was the case as soon as updating the driver fixed the problem :)
@RReverser I have reviewed the issue and code once again today. We do throw an error here https://github.com/arrayfire/arrayfire/blob/master/src/backend/cuda/device_manager.cpp#L459 under this call. Oddly, somewhere in that call or in CUDA API calls, the hang is taking place. Can you please try running the code (with the old driver) through Visual Studio Debugger and share the call stack with me. I am interested in the line that is causing the hang.
@RReverser I have reviewed the issue and code once again today. We do throw an error here https://github.com/arrayfire/arrayfire/blob/master/src/backend/cuda/device_manager.cpp#L459 under this call. Oddly, somewhere in that call or in CUDA API calls, the hang is taking place. Can you please try running the code (with the old driver) through Visual Studio Debugger and share the call stack with me. I am interested in the line that is causing the hang.
@RReverser Did you get a chance to look into this ?
Can you please try running the code (with the old driver)
Yeah, sorry, as I said above I already upgraded the driver which resolved the blocker for me and wouldn't want to look for ways to downgrade it back or return to the issue at this point.
That is alright. I have checked our code base and the necessary log statements for the given scenario are present. I suspect it is either the driver that is causing the hang in which case we cannot gracefully catch the issue to give an appropriate error message.
Until we can figure out some evidence that our log statements aren't working due to logical error in our code base, I don't think we can investigate this further, so closing the issue for now. For any other future/current who may encounter similar issue, if you think you have some additional info to share, feel free to reopen the issue.
Thank you.
Fair enough, thanks.
running into the same issue on windows 10 right now. i tried updating my drivers and cuda and i tried both the 10.x and 11.x arrayfire binaries to no avail. i also tried manually setting the back-end like OP but it didn't seem to fix anything; OpenCL, CUDA, and CPU all throw
STATUS_STACK_BUFFER_OVERRUN. in fact it appears simply runningset_backendand nothing else causes an overrun.the download link on github for the cinfo binary seems to be down right now but heres the output from nvidia-smi
Fri Feb 4 02:19:12 2022 +-----------------------------------------------------------------------------+ | NVIDIA-SMI 511.65 Driver Version: 511.65 CUDA Version: 11.6 | |-------------------------------+----------------------+----------------------+ | GPU Name TCC/WDDM | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 NVIDIA GeForce ... WDDM | 00000000:01:00.0 On | N/A | | 0% 49C P0 126W / 350W | 1056MiB / 24576MiB | 2% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=============================================================================| | 0 N/A N/A 1580 C+G N/A | | 0 N/A N/A 2872 C+G ...icrosoft VS Code\Code.exe N/A | | 0 N/A N/A 5536 C+G N/A | | 0 N/A N/A 6932 C+G ...kyb3d8bbwe\Calculator.exe N/A | | 0 N/A N/A 7884 C+G C:\Windows\explorer.exe N/A | | 0 N/A N/A 12184 C+G ...artMenuExperienceHost.exe N/A | | 0 N/A N/A 12364 C+G ...ekyb3d8bbwe\YourPhone.exe N/A | | 0 N/A N/A 12584 C+G ...5n1h2txyewy\SearchApp.exe N/A | | 0 N/A N/A 13440 C+G ...4__htrsf667h5kn2\AWCC.exe N/A | | 0 N/A N/A 14248 C+G ...ekyb3d8bbwe\HxOutlook.exe N/A | | 0 N/A N/A 15588 C+G ...108.43\msedgewebview2.exe N/A | | 0 N/A N/A 16044 C+G ...perience\NVIDIA Share.exe N/A | | 0 N/A N/A 16160 C+G ...perience\NVIDIA Share.exe N/A | | 0 N/A N/A 16676 C+G N/A | | 0 N/A N/A 16712 C+G ...zilla Firefox\firefox.exe N/A | | 0 N/A N/A 16968 C+G ...nputApp\TextInputHost.exe N/A | | 0 N/A N/A 21680 C+G ...lack\app-4.23.0\slack.exe N/A | | 0 N/A N/A 22944 C+G ...zilla Firefox\firefox.exe N/A | | 0 N/A N/A 23416 C+G ...y\ShellExperienceHost.exe N/A | | 0 N/A N/A 23752 C+G ...lPanel\SystemSettings.exe N/A | +-----------------------------------------------------------------------------+and heres
AF_TRACE=all[unified][1643970448][5736] [ ..\src\api\unified\symbol_manager.cpp(141) ] Attempting: Default System Paths [unified][1643970448][5736] [ ..\src\api\unified\symbol_manager.cpp(144) ] Found: afcpu.dll [unified][1643970448][5736] [ ..\src\api\unified\symbol_manager.cpp(151) ] Device Count: 1. [unified][1643970448][5736] [ ..\src\api\unified\symbol_manager.cpp(141) ] Attempting: Default System Paths [unified][1643970448][5736] [ ..\src\api\unified\symbol_manager.cpp(144) ] Found: afopencl.dll [platform][1643970448][5736] [ ..\src\backend\common\DependencyModule.cpp(99) ] Attempting to load: forge.dll [platform][1643970448][5736] [ ..\src\backend\common\DependencyModule.cpp(102) ] Found: forge.dll [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(218) ] Found 3 OpenCL platforms [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(230) ] Found 1 devices on platform NVIDIA CUDA [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(235) ] Found device NVIDIA GeForce RTX 3090 on platform NVIDIA CUDA [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(230) ] Found 1 devices on platform Intel(R) OpenCL [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(235) ] Found device 11th Gen Intel(R) Core(TM) i9-11900K @ 3.50GHz on platform Intel(R) OpenCL [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(230) ] Found 1 devices on platform Experimental OpenCL 2.1 CPU Only Platform [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(235) ] Found device 11th Gen Intel(R) Core(TM) i9-11900K @ 3.50GHz on platform Experimental OpenCL 2.1 CPU Only Platform [platform][1643970448][5736] [ ..\src\backend\opencl\device_manager.cpp(240) ] Found 3 OpenCL devices error: process didn't exit successfully: `target\debug\rust_playground.exe` (exit code: 0xc0000409, STATUS_STACK_BUFFER_OVERRUN)i elected to file a issue since it seems like the root cause is different from OP. feel free to close the issue if its a duplicate
Description
Any operation, including simple
arrayfire::info()orArray::new(...)seems to be taking a very long time, and eventually fails with:I've used official Windows installer for ArrayFire 3.8 with CUDA 11.2 from here: https://arrayfire.s3.amazonaws.com/3.8.0/ArrayFire-v3.8.0-CUDA-11.2.exe.
It's happening on CUDA backend.
Actually, while trying to see which backend is experiencing this issue, I've noticed that docs don't seem to match the reality. The docs on Backend suggest that default backend / first choice would be OpenCL, falling back to others. However, if I set it explicitly via
arrayfire::set_backend(Backend::OPENCL), then everything works, so I guess the choice is done differently?Yes, I can set backend to OPENCL or CPU and then everything works.
Yes, every time with every API I tried.
All those calls to succeed I guess.
AF_PRINT_ERRORS=1doesn't help / doesn't print anything.AF_TRACE=allproduces following output:Reproducible Code and/or Steps
System Information
Windows:
Download clinfo from https://github.com/Oblomov/clinfo
If you have NVIDIA GPUs. Run nvidia-smi usually located in
C:\Program Files\NVIDIA Corporation\NVSMI
Checklist