Methodforward vis: b, 512, h, w txt: b, L, 512 pad_mask: b, L
modeling/vl_model/layers.py:160
Methodforward vis: 26*26, b, 512 txt: L, b, 512 vis_pos: 26*26, 1, 512 txt_pos: L, 1, 512 pad_mask: b,
modeling/vl_model/layers.py:230
Methodforward img: b, 3, h, w word: b, words word_mask: b, words mask: b, 1, h, w
modeling/vl_model/models.py:81
Methodforward(self, lang, frames, actions, lens_lang, lens_frames, pos=None)
modeling/ns_model/nn/encodings.py:27
Functionprocess(data_is_valid, object_id_to_class, class_to_area_thresholds, classes_list, split, met_path)
modeling/vision_model/prepare_clean_data.py:19