A context manager to partition the model parameters during the model construction with MiCS partition strategy. Model states are partitioned to the number of devices specified via ``mics_shard_size`` field in the deepspeed config json file. The context manager also introduces
(self,
module=None,
data_parallel_group=None,
mem_efficient_linear=True,
remote_device=None,
pin_memory=False,
config_dict_or_path=None,
config=None,
enabled=True,
dtype=None,
mpu=None)
| 54 | class MiCS_Init(Init): |
| 55 | |
| 56 | def __init__(self, |
| 57 | module=None, |
| 58 | data_parallel_group=None, |
| 59 | mem_efficient_linear=True, |
| 60 | remote_device=None, |
| 61 | pin_memory=False, |
| 62 | config_dict_or_path=None, |
| 63 | config=None, |
| 64 | enabled=True, |
| 65 | dtype=None, |
| 66 | mpu=None): |
| 67 | """A context manager to partition the model parameters during the model |
| 68 | construction with MiCS partition strategy. Model states are partitioned |
| 69 | to the number of devices specified via ``mics_shard_size`` field in the |
| 70 | deepspeed config json file. The context manager also introduces |
| 71 | hierarchical communication method to reduce the cost of inter-node |
| 72 | communications, which can be enabled with |
| 73 | ``mics_hierarchical_params_gather`` field in deepspeed config. |
| 74 | |
| 75 | Args: |
| 76 | module (``torch.nn.Module``, optional): If provided, partition the model as |
| 77 | if it was constructed in the context. |
| 78 | data_parallel_group (``deepspeed.comm`` process group, optional): |
| 79 | The group of processes to partition among. Defaults to all processes. |
| 80 | mem_efficient_linear (bool, optional): Replace |
| 81 | torch.nn.functional.linear with an implementation that allows |
| 82 | DeepSpeed to partition parameters. Defaults to ``True``. |
| 83 | remote_device (string, optional): The initial device to store model |
| 84 | weights e.g., ``cpu``, ``nvme``. Passing ``"cpu"`` will create the model in CPU |
| 85 | memory. The model may still be moved to GPU based on the |
| 86 | offload settings for training. Defaults to param offload device if a config is |
| 87 | defined, otherwise GPU. |
| 88 | pin_memory (bool, optional): Potentially increase performance by |
| 89 | using pinned memory for model weights. ``remote_device`` must be |
| 90 | ``"cpu"``. Defaults to pin_memory value in config, otherwise ``False``. |
| 91 | config_dict_or_path (dict or ``json file``, optional): If provided, provides configuration |
| 92 | for swapping fp16 params to NVMe. |
| 93 | config (dict or ``json file``, optional): Deprecated, use config_dict_or_path instead. |
| 94 | enabled (bool, optional): If ``False``, this context has no |
| 95 | effect. Defaults to ``True``. |
| 96 | dtype (``dtype``, optional): Can be used to change the data type of the parameters. |
| 97 | Supported options are ``torch.half`` and ``torch.float``. Defaults to ``None`` |
| 98 | mpu (``object``, optional): A model parallelism unit object that implements get_{model,data}_parallel_{rank,group,world_size}. |
| 99 | |
| 100 | This context follows the same logic as ``deepspeed.zero.Init()``, but |
| 101 | with the modification for partition size of each parameter. |
| 102 | |
| 103 | Examples |
| 104 | -------- |
| 105 | |
| 106 | #. Allocate a model and partition it among all processes: |
| 107 | |
| 108 | .. code-block:: python |
| 109 | # the config_dict_or_path is required to let the context manager know |
| 110 | # how partition the parameters. |
| 111 | # The configuration has to include the field ``mics_shard_size`` |
| 112 | with deepspeed.zero.MiCS_Init(config_dict_or_path=ds_config): |
| 113 | model = MyLargeModel() |
no test coverage detected