Instantiate and evaluate a model on a list of tasks. :param model: Union[str, LM] Name of model, transformers.PreTrainedModel object, or LM object, see lm_eval.models.get_model :param model_args: Optional[str] String arguments for each model class, see LM.create_from_arg_str
(
model,
model_args=None,
tasks=[],
num_fewshot=0,
batch_size=None,
max_batch_size=None,
device=None,
no_cache=False,
limit=None,
bootstrap_iters=100000,
description_dict=None,
check_integrity=False,
decontamination_ngrams_path=None,
write_out=False,
output_base_path=None,
test_set=True
)
| 15 | |
| 16 | @positional_deprecated |
| 17 | def simple_evaluate( |
| 18 | model, |
| 19 | model_args=None, |
| 20 | tasks=[], |
| 21 | num_fewshot=0, |
| 22 | batch_size=None, |
| 23 | max_batch_size=None, |
| 24 | device=None, |
| 25 | no_cache=False, |
| 26 | limit=None, |
| 27 | bootstrap_iters=100000, |
| 28 | description_dict=None, |
| 29 | check_integrity=False, |
| 30 | decontamination_ngrams_path=None, |
| 31 | write_out=False, |
| 32 | output_base_path=None, |
| 33 | test_set=True |
| 34 | ): |
| 35 | """Instantiate and evaluate a model on a list of tasks. |
| 36 | |
| 37 | :param model: Union[str, LM] |
| 38 | Name of model, transformers.PreTrainedModel object, or LM object, see lm_eval.models.get_model |
| 39 | :param model_args: Optional[str] |
| 40 | String arguments for each model class, see LM.create_from_arg_string. |
| 41 | Ignored if `model` argument is a LM object. |
| 42 | :param tasks: list[Union[str, Task]] |
| 43 | List of task names or Task objects. Task objects will be taken to have name task.EVAL_HARNESS_NAME if defined and type(task).__name__ otherwise. |
| 44 | :param num_fewshot: int |
| 45 | Number of examples in few-shot context |
| 46 | :param batch_size: int or str, optional |
| 47 | Batch size for model |
| 48 | :param max_batch_size: int, optional |
| 49 | Maximal batch size to try with automatic batch size detection |
| 50 | :param device: str, optional |
| 51 | PyTorch device (e.g. "cpu" or "cuda:0") for running models |
| 52 | :param no_cache: bool |
| 53 | Whether or not to cache |
| 54 | :param limit: int or float, optional |
| 55 | Limit the number of examples per task (only use this for testing), If <1, limit is a percentage of the total number of examples. |
| 56 | :param bootstrap_iters: |
| 57 | Number of iterations for bootstrap statistics |
| 58 | :param description_dict: dict[str, str] |
| 59 | Dictionary of custom task descriptions of the form: `task_name: description` |
| 60 | :param check_integrity: bool |
| 61 | Whether to run the relevant part of the test suite for the tasks |
| 62 | :param write_out: bool |
| 63 | If True, write details about prompts and logits to json for all tasks |
| 64 | :param output_base_path: str, optional |
| 65 | Directory to which detailed eval info will be written. Defaults to present working dir. |
| 66 | :return |
| 67 | Dictionary of results |
| 68 | """ |
| 69 | print("test_set", test_set) |
| 70 | random.seed(1234) |
| 71 | np.random.seed(1234) |
| 72 | |
| 73 | assert tasks != [], "No tasks specified" |
| 74 |
nothing calls this directly
no test coverage detected