REINFORCE loss. Not used right now. Args: cond: dict with key state/rgb; more recent obs at the end state: (B, To, Do) rgb: (B, To, C, H, W) chains: (B, K+1, Ta, Da) reward (to go): (b,)
(self, cond, chains, reward)
source not stored for this graph (policy: none)
nothing calls this directly
no test coverage detected