This simulator provides basic cues and rewards according to the network's choice, as described in the Backpropamine paper: https://openreview.net/pdf?id=r1lrAiA5Ym :param epdur: int: duration (timesteps) of an episode; default = 200 :param cuebits: int: max number of bits t
| 9 | |
| 10 | |
| 11 | class CueRewardSimulator: |
| 12 | """ |
| 13 | This simulator provides basic cues and rewards according to the |
| 14 | network's choice, as described in the Backpropamine paper: |
| 15 | https://openreview.net/pdf?id=r1lrAiA5Ym |
| 16 | |
| 17 | :param epdur: int: duration (timesteps) of an episode; default = 200 |
| 18 | :param cuebits: int: max number of bits to hold a cue (max value = 2^n for n bits) |
| 19 | :param seed: real: random seed |
| 20 | :param zprob: real: probability of zero vectors in each trial. |
| 21 | """ |
| 22 | |
| 23 | def __init__(self, **kwargs) -> None: |
| 24 | self.ep_duration = kwargs.get("epdur", 200) # episode duration in timesteps |
| 25 | self.cuebits = kwargs.get("cuebits", 20) |
| 26 | self.seed = int(kwargs.get("seed", time())) |
| 27 | self.zeroprob = kwargs.get("zprob", 0.6) |
| 28 | |
| 29 | # zero array consists of the binary cue vector + four other fields. |
| 30 | self.zeroArray = np.zeros(self.cuebits + 4, dtype="int32") |
| 31 | |
| 32 | assert ( |
| 33 | 0.0 <= self.zeroprob and self.zeroprob <= 1.0 |
| 34 | ), "zprob must be valid probability" |
| 35 | |
| 36 | def make(self, name): |
| 37 | self.reset() # Simply reset according to grid definition. |
| 38 | |
| 39 | def step(self, action): |
| 40 | """ |
| 41 | Every trial, we randomly select two of the four cues to provide to the |
| 42 | network. Every timestep within that trial we either randomly display |
| 43 | only zeros, or we alternate between the two cues in the pair. |
| 44 | |
| 45 | At the end of a trial, we provide the response cue, for which the network |
| 46 | must respond 1 if the target was one of the provided cues or 0 if it was |
| 47 | not. The next timestep, we evaluate the response, giving a reward of 1 |
| 48 | for correct and -1 for incorrect. |
| 49 | |
| 50 | :param action: network's decision if the target cue was displayed. |
| 51 | :return obs: observation of vector with binary cue and the following fields: |
| 52 | - time since start of episode |
| 53 | - one-hot-encoded value for a response of 0 in previous timestep |
| 54 | - one-hot-encoded value for a response of 1 in previous timestep |
| 55 | - reward of previous timestep |
| 56 | :return reward: 1 for correct response; -1 for incorrect response. |
| 57 | :return done: indicates termination of simulation |
| 58 | :return info: dictionary including values for debugging purposes. |
| 59 | """ |
| 60 | self.tstep += 1 # increment episode timestep |
| 61 | self.trialTime -= 1 # decrement current trial timestep |
| 62 | |
| 63 | # Populate base fields of observation. |
| 64 | self.obs = self.zeroArray # default to empty array |
| 65 | self.obs[-4] = self.tstep # time since start of episode. |
| 66 | self.obs[-3] = int(self.response == 0) # response = 0 for previous timestep |
| 67 | self.obs[-2] = int(self.response == 1) # response = 1 for previous timestep |
| 68 | self.obs[-1] = self.reward[0] # reward of previous timestep |