我在任何地方都找不到答案,我可能忽略了这一点,但似乎不能使用
__constant__
内存(以及
cudaMemcpyToSymbol
)和UVA的点对点访问。
我已经尝试过simpleP2P nvidia示例代码,它在我的4 NV100上运行良好,但只要我在内核中声明因子2为:
__constant__ float M_; // in global space
float M = 2.0;
cudaMemcpyToSymbol(M_, &M, sizeof(float), 0, cudaMemcpyDefault);
结果基本上是零。如果我使用C预处理器(例如
#define M_ 2.0
)来定义它,它会工作得很好。
所以我想知道,这是真的还是我做错了什么?还有其他类型的内存也不能以这种方式访问吗(例如纹理内存)?
发布于 2020-03-10 23:33:10
你的问题为什么“结果基本上是零”和UVA的P2P访问之间的关系对我来说不是很清楚。
是真的吗?还是我做错了什么?
很难说,因为你的问题有点含糊,也没有完整的例子。
__constant__ float M_
在
所有
CUDA可见设备的
常量内存
上分配变量
M_
。为了在多个设备上设置值,您应该执行如下操作:
__constant__ float M_; // <= This declares M_ on the constant memory of all CUDA visible devices
__global__ void showMKernel() {
printf("****** M_ = %f\n", M_);
int main()
float M = 2.0;
// Make sure that the return values are properly checked for cudaSuccess ...
int deviceCount = -1;
cudaGetDeviceCount(&deviceCount);
// Set M_ on the constant memory of each device:
for (int i = 0; i < deviceCount; i++) {
cudaSetDevice(i);
cudaMemcpyToSymbol(M_, &M, sizeof(float), 0, cudaMemcpyDefault);
// Now, run a kernel to show M_:
for (int i = 0; i < deviceCount; i++)
cudaSetDevice(i);
printf("Device %g :\n", i);
showMKernel<<<1,1>>>();
cudaDeviceSynchronize();
}
它返回:
Device 0 :
****** M = 2.000000
Device 1 :
****** M = 2.000000
// so on for other devices
现在,如果我把
// Set M_ on the constant memory of each device:
for (int i = 0; i < deviceCount; i++) {
cudaSetDevice(i);