Stream Merge. It merges the frame-level processed audio chunks in the streaming *simulation*. It is noted that, in real applications, the processed audio should be sent to the output channel frame by frame. You may refer to this function to manage your streaming outp
(self, chunks: torch.Tensor, ilens: torch.tensor = None)
| 41 | )[0] |
| 42 | |
| 43 | def streaming_merge(self, chunks: torch.Tensor, ilens: torch.tensor = None): |
| 44 | """Stream Merge. |
| 45 | |
| 46 | It merges the frame-level processed audio chunks |
| 47 | in the streaming *simulation*. It is noted that, in real applications, |
| 48 | the processed audio should be sent to the output channel frame by frame. |
| 49 | You may refer to this function to manage your streaming output buffer. |
| 50 | |
| 51 | Args: |
| 52 | chunks: List [(B, frame_size),] |
| 53 | ilens: [B] |
| 54 | Returns: |
| 55 | merge_audio: [B, T] |
| 56 | """ |
| 57 | hop_size = self.stride |
| 58 | frame_size = self.kernel_size |
| 59 | |
| 60 | num_chunks = len(chunks) |
| 61 | batch_size = chunks[0].shape[0] |
| 62 | audio_len = ( |
| 63 | int(hop_size * num_chunks + frame_size - hop_size) |
| 64 | if not ilens |
| 65 | else ilens.max() |
| 66 | ) |
| 67 | |
| 68 | output = torch.zeros((batch_size, audio_len), dtype=chunks[0].dtype).to( |
| 69 | chunks[0].device |
| 70 | ) |
| 71 | |
| 72 | for i, chunk in enumerate(chunks): |
| 73 | output[:, i * hop_size : i * hop_size + frame_size] += chunk |
| 74 | |
| 75 | return output |
| 76 | |
| 77 | |
| 78 | if __name__ == "__main__": |