Evaluate ASR model performance on long-form transcription tasks. This evaluation function handles variable-length audio segments (typically longer than 30 seconds) that require different processing than short-form evaluation. It uses the model's built-in transcribe() method with beam se
(
batch_size: int,
num_workers: int,
ckpt: str,
eval_set: Literal[
"tedlium",
"meanwhile",
"rev16",
"earnings21",
"earnings22",
"coraal",
"kincaid46",
],
log_dir: str,
n_mels: int = DEFAULT_N_MELS,
bootstrap: bool = False,
wandb_log: bool = False,
wandb_log_dir: str = "wandb",
eval_dir: str = "data/eval",
hf_token: Optional[str] = None,
)
| 1904 | |
| 1905 | |
| 1906 | def long_form_eval( |
| 1907 | batch_size: int, |
| 1908 | num_workers: int, |
| 1909 | ckpt: str, |
| 1910 | eval_set: Literal[ |
| 1911 | "tedlium", |
| 1912 | "meanwhile", |
| 1913 | "rev16", |
| 1914 | "earnings21", |
| 1915 | "earnings22", |
| 1916 | "coraal", |
| 1917 | "kincaid46", |
| 1918 | ], |
| 1919 | log_dir: str, |
| 1920 | n_mels: int = DEFAULT_N_MELS, |
| 1921 | bootstrap: bool = False, |
| 1922 | wandb_log: bool = False, |
| 1923 | wandb_log_dir: str = "wandb", |
| 1924 | eval_dir: str = "data/eval", |
| 1925 | hf_token: Optional[str] = None, |
| 1926 | ) -> None: |
| 1927 | """Evaluate ASR model performance on long-form transcription tasks. |
| 1928 | |
| 1929 | This evaluation function handles variable-length audio segments (typically longer |
| 1930 | than 30 seconds) that require different processing than short-form evaluation. |
| 1931 | It uses the model's built-in transcribe() method with beam search and timestamp |
| 1932 | generation for optimal long-form performance. |
| 1933 | |
| 1934 | Key differences from short-form evaluation: |
| 1935 | - No mel-spectrogram preprocessing (handled internally by transcribe()) |
| 1936 | - Uses beam search decoding with best-of selection |
| 1937 | - Processes full audio length without padding/trimming |
| 1938 | - Enables timestamp generation for alignment verification |
| 1939 | - Single-sample batch processing for stability |
| 1940 | |
| 1941 | Args: |
| 1942 | batch_size (int): Number of audio samples per batch (typically 1 for long-form) |
| 1943 | num_workers (int): Number of worker processes for data loading |
| 1944 | ckpt (str): Path to model checkpoint file |
| 1945 | eval_set (Literal): Name of evaluation dataset (see supported datasets below) |
| 1946 | log_dir (str): Directory for saving evaluation results and logs |
| 1947 | n_mels (int, optional): Number of mel-spectrogram bins (unused in long-form). |
| 1948 | Defaults to DEFAULT_N_MELS. |
| 1949 | bootstrap (bool, optional): Enable bootstrap sampling for confidence intervals. |
| 1950 | Defaults to False. |
| 1951 | wandb_log (bool, optional): Enable Weights & Biases logging. Defaults to False. |
| 1952 | wandb_log_dir (str, optional): Directory for WandB logs. Defaults to "wandb". |
| 1953 | eval_dir (str, optional): Root directory for evaluation datasets. |
| 1954 | Defaults to "data/eval". |
| 1955 | hf_token (Optional[str], optional): HuggingFace token for private datasets. |
| 1956 | Defaults to None. |
| 1957 | |
| 1958 | Supported Datasets: |
| 1959 | - tedlium: TED talks with longer audio segments |
| 1960 | - meanwhile: Meanwhile podcast corpus for long-form evaluation |
| 1961 | - rev16: Rev.com transcription corpus subset |
| 1962 | - earnings21: Corporate earnings call transcripts (2021) |
| 1963 | - earnings22: Corporate earnings call transcripts (2022) |
nothing calls this directly
no test coverage detected