MCPcopy Create free account
hub / github.com/allenai/OLMoASR / long_form_eval

Function long_form_eval

scripts/eval/eval.py:1906–2139  ·  view source on GitHub ↗

Evaluate ASR model performance on long-form transcription tasks. This evaluation function handles variable-length audio segments (typically longer than 30 seconds) that require different processing than short-form evaluation. It uses the model's built-in transcribe() method with beam se

(
    batch_size: int,
    num_workers: int,
    ckpt: str,
    eval_set: Literal[
        "tedlium",
        "meanwhile",
        "rev16",
        "earnings21",
        "earnings22",
        "coraal",
        "kincaid46",
    ],
    log_dir: str,
    n_mels: int = DEFAULT_N_MELS,
    bootstrap: bool = False,
    wandb_log: bool = False,
    wandb_log_dir: str = "wandb",
    eval_dir: str = "data/eval",
    hf_token: Optional[str] = None,
)

Source from the content-addressed store, hash-verified

1904
1905
1906def long_form_eval(
1907 batch_size: int,
1908 num_workers: int,
1909 ckpt: str,
1910 eval_set: Literal[
1911 "tedlium",
1912 "meanwhile",
1913 "rev16",
1914 "earnings21",
1915 "earnings22",
1916 "coraal",
1917 "kincaid46",
1918 ],
1919 log_dir: str,
1920 n_mels: int = DEFAULT_N_MELS,
1921 bootstrap: bool = False,
1922 wandb_log: bool = False,
1923 wandb_log_dir: str = "wandb",
1924 eval_dir: str = "data/eval",
1925 hf_token: Optional[str] = None,
1926) -> None:
1927 """Evaluate ASR model performance on long-form transcription tasks.
1928
1929 This evaluation function handles variable-length audio segments (typically longer
1930 than 30 seconds) that require different processing than short-form evaluation.
1931 It uses the model's built-in transcribe() method with beam search and timestamp
1932 generation for optimal long-form performance.
1933
1934 Key differences from short-form evaluation:
1935 - No mel-spectrogram preprocessing (handled internally by transcribe())
1936 - Uses beam search decoding with best-of selection
1937 - Processes full audio length without padding/trimming
1938 - Enables timestamp generation for alignment verification
1939 - Single-sample batch processing for stability
1940
1941 Args:
1942 batch_size (int): Number of audio samples per batch (typically 1 for long-form)
1943 num_workers (int): Number of worker processes for data loading
1944 ckpt (str): Path to model checkpoint file
1945 eval_set (Literal): Name of evaluation dataset (see supported datasets below)
1946 log_dir (str): Directory for saving evaluation results and logs
1947 n_mels (int, optional): Number of mel-spectrogram bins (unused in long-form).
1948 Defaults to DEFAULT_N_MELS.
1949 bootstrap (bool, optional): Enable bootstrap sampling for confidence intervals.
1950 Defaults to False.
1951 wandb_log (bool, optional): Enable Weights & Biases logging. Defaults to False.
1952 wandb_log_dir (str, optional): Directory for WandB logs. Defaults to "wandb".
1953 eval_dir (str, optional): Root directory for evaluation datasets.
1954 Defaults to "data/eval".
1955 hf_token (Optional[str], optional): HuggingFace token for private datasets.
1956 Defaults to None.
1957
1958 Supported Datasets:
1959 - tedlium: TED talks with longer audio segments
1960 - meanwhile: Meanwhile podcast corpus for long-form evaluation
1961 - rev16: Rev.com transcription corpus subset
1962 - earnings21: Corporate earnings call transcripts (2021)
1963 - earnings22: Corporate earnings call transcripts (2022)

Callers

nothing calls this directly

Calls 7

gen_inf_ckptFunction · 0.90
load_modelFunction · 0.90
EvalDatasetClass · 0.85
_add_to_wandb_tableFunction · 0.85
_log_resultsFunction · 0.85
init_wandbMethod · 0.80
deviceMethod · 0.45

Tested by

no test coverage detected