Browse by type
This Python-based tool is designed for transcribing YouTube videos and playlists into text. It integrates various technologies like faster-whisper for transcription, SpaCy for natural language processing, and CUDA for GPU acceleration, aimed at processing video content efficiently. The script is capable of handling both individual videos and entire playlists, outputting accurate transcripts along with metadata.
![]() |
|---|
| Bulk Transcripts Have Never Been This Easy! |
pytube to download the audio from YouTube videos or playlists.faster_whisper.WhisperModel for converting audio to text. This model is a variant of OpenAI's Whisper designed for speed and accuracy.tqdm for displaying progress bars during transcription.convert_single_video flag.It sets up necessary directories for storing downloaded audio, transcripts, and metadata.
Environment Configuration:
Configures the number of workers for transcription based on the CPU core count.
Video Processing:
It ensures unique naming for each audio file to avoid overwrites.
Transcription:
Transcription results are split into sentences, either using SpaCy or a custom regex-based splitter.
Metadata Generation:
Along with the transcript, the script generates metadata including timestamps and confidence scores for each segment.
Output:
The transcripts are saved in plain text, CSV, and JSON formats, providing both the raw transcript and structured metadata.
Display/Read:
transcript_reader.html, which does further clean up and offers a "Reader Mode" where you can choose the font, text size, text width, and toggle dark mode. Simply open this html file in your browser and paste in the transcript text from one of the generated files in the generated_transcript_combined_texts folder.![]() |
|---|
| Screenshot of it in Action |
![]() |
![]() |
|---|---|
| Paste Transcript Text into the Transcript Reader HTML File | Reader using Dark Mode and Cambria Font |
git clone https://github.com/Dicklesworthstone/bulk_transcribe_youtube_videos_from_playlist.git
cd bulk_transcribe_youtube_videos_from_playlist
If you prefer to use pyenv for managing Python versions, follow these steps:
# Install pyenv if not already installed
if ! command -v pyenv &> /dev/null; then
git clone https://github.com/pyenv/pyenv.git ~/.pyenv
echo 'export PYENV_ROOT="$HOME/.pyenv"' >> ~/.bashrc
echo 'export PATH="$PYENV_ROOT/bin:$PATH"' >> ~/.bashrc
echo 'eval "$(pyenv init --path)"' >> ~/.bashrc
source ~/.bashrc
fi
# Update pyenv and install Python 3.12
cd ~/.pyenv && git pull && cd -
pyenv install 3.12
# Set local Python version for the project
pyenv local 3.12
Choose one of the following methods:
# Create a virtual environment
python3 -m venv venv
# Activate the virtual environment
source venv/bin/activate
# Upgrade pip and install wheel
python3 -m pip install --upgrade pip
python3 -m pip install wheel
# Install dependencies
pip install -r requirements.txt
# Create a virtual environment
python -m venv venv
# Activate the virtual environment
source venv/bin/activate
# Upgrade pip and install wheel
python -m pip install --upgrade pip
python -m pip install wheel
# Install dependencies
pip install -r requirements.txt
Open the bulk_transcribe_youtube_videos_from_playlist.py file and set the following variables according to your preferences:
convert_single_video: Set to 1 for a single video, 0 for a playlistuse_spacy_for_sentence_splitting: Set to 1 to use SpaCy, 0 for regex-based splittinguse_openai_api_for_transcription: Set to 1 to use OpenAI API, 0 for local Whisper modelopenai_api_key: If using OpenAI API, replace with your API keysingle_video_url: URL of the single video (if convert_single_video is 1)playlist_url: URL of the playlist (if convert_single_video is 0)Execute the script with Python:
python bulk_transcribe_youtube_videos_from_playlist.py
If you encounter any issues during setup or execution, please check the project's issue tracker on GitHub or create a new issue with details about the problem.
convert_single_video flag. This choice dictates which URL (either single_video_url or playlist_url) will be used for downloading content.add_to_system_path function adds new paths to the system environment, ensuring that dependencies like CUDA Toolkit are accessible. For Windows systems, it also handles the case where the new path contains spaces, enclosing it in quotes.get_cuda_toolkit_path locates the CUDA Toolkit directory, crucial for GPU acceleration. It checks the Anaconda packages directory for the toolkit's installation path.download_audio asynchronously downloads audio from YouTube videos. It ensures unique naming for each audio file by appending a counter if a file with the same name already exists. This function returns the path to the downloaded audio file and the filename.compute_transcript_with_whisper_from_audio_func configures either the local WhisperModel or the OpenAI API for transcription. For local inference, it checks CUDA availability and sets the device and compute type accordingly.use_spacy_for_sentence_splitting flag, the script either uses SpaCy or a custom regex-based method for sentence splitting. This is important for structuring the transcript into readable sentences.clean_filename sanitizes video titles for use as filenames, removing special characters and replacing spaces with underscores.remove_pagination_breaks cleans up the transcript text by removing hyphens at line breaks and correcting line break issues, improving readability.normalize_logprobs normalizes the log probabilities of transcription segments, useful for assessing the model's confidence in its transcription.__main__ block, where it selects the URL to process (single video or playlist) and initiates the process_video_or_playlist coroutine.process_video_or_playlist handles the asynchronous downloading and transcription of videos. It creates a semaphore to limit the number of simultaneous downloads based on max_simultaneous_youtube_downloads. Important Note: The OpenAI API version of Whisper uses an older and less accurate version of the Whisper model. Users will achieve significantly higher transcription accuracy using local inference with the latest Whisper model.
model.transcribe on a separate thread using asyncio.to_thread to maintain the asynchronous nature of the script. For the OpenAI API, sends a request to the transcription endpoint.beam_size of 10 and activates the vad_filter. The beam_size parameter affects the trade-off between accuracy and speed during transcription - a higher value can lead to more accurate results but requires more computational resources. The vad_filter (Voice Activity Detection filter) helps in ignoring non-speech segments in the audio, focusing the transcription process on relevant audio parts.sophisticated_sentence_splitter.$ claude mcp add bulk_transcribe_youtube_videos_from_playlist \
-- python -m otcore.mcp_server <graph>