Determines whether to use the composite or specialized CPU kernel. When the total size of the tensor is larger than the cache size and the batch size is large compared to the smallest matrix dimension, then the composite implementation is inefficient since it has to read the entire
(fast, tensor_shape)
source not stored for this graph (policy: none)
no test coverage detected