Update the `Q` function using the TD(0) Q-learning update: Q[s, a] <- Q[s, a] + lr * ( r + temporal_discount * max_a { Q[s', a] } - Q[s, a] ) Parameters ---------- s : int as returned by `self._obs2num` The id for
(self, s, a, r, s_)
| 1196 | Q[(s, a)] = qsa + self.lr * (r + self.temporal_discount * E_Q - qsa) |
| 1197 | |
| 1198 | def _off_policy_update(self, s, a, r, s_): |
| 1199 | """ |
| 1200 | Update the `Q` function using the TD(0) Q-learning update: |
| 1201 | |
| 1202 | Q[s, a] <- Q[s, a] + lr * ( |
| 1203 | r + temporal_discount * max_a { Q[s', a] } - Q[s, a] |
| 1204 | ) |
| 1205 | |
| 1206 | Parameters |
| 1207 | ---------- |
| 1208 | s : int as returned by `self._obs2num` |
| 1209 | The id for the state/observation at timestep `t-1` |
| 1210 | a : int as returned by `self._action2num` |
| 1211 | The id for the action taken at timestep `t-1` |
| 1212 | r : float |
| 1213 | The reward after taking action `a` in state `s` at timestep `t-1` |
| 1214 | s_ : int as returned by `self._obs2num` |
| 1215 | The id for the state/observation at timestep `t` |
| 1216 | """ |
| 1217 | Q, E = self.parameters["Q"], self.env_info |
| 1218 | n_actions = np.prod(E["n_actions_per_dim"]) |
| 1219 | |
| 1220 | qsa = Q[(s, a)] |
| 1221 | Qs_ = [Q[(s_, aa)] for aa in range(n_actions)] if s_ else [0] |
| 1222 | Q[(s, a)] = qsa + self.lr * (r + self.temporal_discount * np.max(Qs_) - qsa) |
| 1223 | |
| 1224 | def update(self): |
| 1225 | """Update the parameters of the model online after each new state-action.""" |