to_device ( seq , stream = stream ) for seq in seqs ) d_T = cuda . To make multiple shared memory arrays, the full dynamic shared memory can be sliced using the array[start:end] operator. + if l < A : + hashes [ l * t + k ] = global_hashes [ l ][ k ] + signs [ l * t + k ] = global_signs [ l ][ k ] + for ll in range ( l , D , D // L ): + Tin [ 0 * plane + 0 * D + ll ] = 0 + Tin [ 0 * plane + ( 0 + 1 ) * D + ll ] = 0 + Tin [ 1 * plane + 0 * D + ll ] = 0 + Tin [ 1 * plane + ( 0 + 1 ) * D + ll ] = 0 cuda .