MCPcopy Create free account
hub / github.com/ddbourgin/numpy-ml / _off_policy_update

Method _off_policy_update

numpy_ml/rl_models/agents.py:1198–1222  ·  view source on GitHub ↗

Update the `Q` function using the TD(0) Q-learning update: Q[s, a] <- Q[s, a] + lr * ( r + temporal_discount * max_a { Q[s', a] } - Q[s, a] ) Parameters ---------- s : int as returned by `self._obs2num` The id for

(self, s, a, r, s_)

Source from the content-addressed store, hash-verified

1196 Q[(s, a)] = qsa + self.lr * (r + self.temporal_discount * E_Q - qsa)
1197
1198 def _off_policy_update(self, s, a, r, s_):
1199 """
1200 Update the `Q` function using the TD(0) Q-learning update:
1201
1202 Q[s, a] <- Q[s, a] + lr * (
1203 r + temporal_discount * max_a { Q[s&#x27;, a] } - Q[s, a]
1204 )
1205
1206 Parameters
1207 ----------
1208 s : int as returned by `self._obs2num`
1209 The id for the state/observation at timestep `t-1`
1210 a : int as returned by `self._action2num`
1211 The id for the action taken at timestep `t-1`
1212 r : float
1213 The reward after taking action `a` in state `s` at timestep `t-1`
1214 s_ : int as returned by `self._obs2num`
1215 The id for the state/observation at timestep `t`
1216 """
1217 Q, E = self.parameters["Q"], self.env_info
1218 n_actions = np.prod(E["n_actions_per_dim"])
1219
1220 qsa = Q[(s, a)]
1221 Qs_ = [Q[(s_, aa)] for aa in range(n_actions)] if s_ else [0]
1222 Q[(s, a)] = qsa + self.lr * (r + self.temporal_discount * np.max(Qs_) - qsa)
1223
1224 def update(self):
1225 """Update the parameters of the model online after each new state-action."""

Callers 1

updateMethod · 0.95

Calls

no outgoing calls

Tested by

no test coverage detected