Encode audios into latents. Audios should already be preprocesed by preprocess_audio_for_encoder. If chunked is True, split the audio into chunks of a given maximum size chunk_size, with given overlap. Overlap and chunk_size params are both measured in number of latents (not
(self, audio, chunked=False, overlap=32, chunk_size=128, **kwargs)