Log selected cluster and runtime environment variables. The output includes two blocks: - Cluster environment variables (scheduler/runtime IDs). - Runtime environment variables (CUDA/NCCL/OMP/PATH, etc.). Environment variables collected by default: Cluster prefixes: |
(
logger: logging.Logger,
cluster_prefixes: Iterable[str] | None = None,
runtime_prefixes: Iterable[str] | None = None,
runtime_keys: Iterable[str] | None = None,
)
| 546 | |
| 547 | |
| 548 | def log_env_metadata( |
| 549 | logger: logging.Logger, |
| 550 | cluster_prefixes: Iterable[str] | None = None, |
| 551 | runtime_prefixes: Iterable[str] | None = None, |
| 552 | runtime_keys: Iterable[str] | None = None, |
| 553 | ) -> None: |
| 554 | """Log selected cluster and runtime environment variables. |
| 555 | |
| 556 | The output includes two blocks: |
| 557 | - Cluster environment variables (scheduler/runtime IDs). |
| 558 | - Runtime environment variables (CUDA/NCCL/OMP/PATH, etc.). |
| 559 | |
| 560 | Environment variables collected by default: |
| 561 | |
| 562 | Cluster prefixes: |
| 563 | | Prefix | Purpose | |
| 564 | |-----------|------------------------------------------------------| |
| 565 | | `SLURM_` | Slurm job/step metadata (job id, task id, node info) | |
| 566 | | `PBS_` | PBS/Torque job metadata | |
| 567 | | `LSF_` | LSF job metadata | |
| 568 | | `SGE_` | SGE job metadata | |
| 569 | | `COBALT_` | Cobalt job metadata | |
| 570 | | `OMPI_` | Open MPI runtime metadata | |
| 571 | | `PMI_` | PMI (Process Management Interface) metadata | |
| 572 | | `MPI_` | MPI runtime metadata (generic prefix) | |
| 573 | |
| 574 | Runtime prefixes: |
| 575 | | Prefix | Purpose | |
| 576 | |-------------|-------------------------------------| |
| 577 | | `NCCL_` | NCCL configuration (multi-GPU comms)| |
| 578 | | `CUDA_` | CUDA runtime configuration | |
| 579 | | `ROCM_` | ROCm runtime configuration | |
| 580 | | `OMP_` | OpenMP threading configuration | |
| 581 | | `MKL_` | Intel MKL configuration | |
| 582 | | `OPENBLAS_` | OpenBLAS configuration | |
| 583 | | `UCX_` | UCX communication configuration | |
| 584 | | `NVIDIA_` | NVIDIA runtime configuration | |
| 585 | |
| 586 | Explicit runtime keys: |
| 587 | | Key | Purpose | |
| 588 | |-----------------------|----------------------------------| |
| 589 | | `PATH` | Executable search path | |
| 590 | | `PYTHONPATH` | Python module search path | |
| 591 | | `LD_LIBRARY_PATH` | Shared library search path | |
| 592 | | `CUDA_VISIBLE_DEVICES`| GPU visibility mask | |
| 593 | | `RANK` | Global rank (distributed) | |
| 594 | | `LOCAL_RANK` | Local rank on node | |
| 595 | | `NODE_RANK` | Node rank in job | |
| 596 | | `WORLD_SIZE` | Total process count | |
| 597 | | `MASTER_ADDR` | Distributed master address | |
| 598 | | `MASTER_PORT` | Distributed master port | |
| 599 | |
| 600 | Example usage: |
| 601 | ```python |
| 602 | from pathlib import Path |
| 603 | from espnet3.utils.logging_utils import configure_logging, log_env_metadata |
| 604 | |
| 605 | logger = configure_logging(log_dir=Path("exp/run1")) |
no test coverage detected
searching dependent graphs…