"Pull" (i.e., sample from) a given arm's payoff distribution. Parameters ---------- arm_id : int The integer ID of the arm to sample from context : :py:class:`ndarray ` of shape `(D,)` or None The context vector for the
(self, arm_id, context=None)
| 45 | pass |
| 46 | |
| 47 | def pull(self, arm_id, context=None): |
| 48 | """ |
| 49 | "Pull" (i.e., sample from) a given arm's payoff distribution. |
| 50 | |
| 51 | Parameters |
| 52 | ---------- |
| 53 | arm_id : int |
| 54 | The integer ID of the arm to sample from |
| 55 | context : :py:class:`ndarray <numpy.ndarray>` of shape `(D,)` or None |
| 56 | The context vector for the current timestep if this is a contextual |
| 57 | bandit. Otherwise, this argument is unused and defaults to None. |
| 58 | |
| 59 | Returns |
| 60 | ------- |
| 61 | reward : float |
| 62 | The reward sampled from the given arm's payoff distribution |
| 63 | """ |
| 64 | assert arm_id < self.n_arms |
| 65 | |
| 66 | self.step += 1 |
| 67 | return self._pull(arm_id, context) |
| 68 | |
| 69 | def reset(self): |
| 70 | """Reset the bandit step and action counters to zero.""" |