During each decoder layer, each 3D poses of each queries will be projected into each camera view to aggregate features and update coarse projected 2D poses. Then, a triangulation will be performed to get updated 3D poses for each queries. Also, the feature of each query will
(self, tgt, query_pos, reference_points, src_views,
src_spatial_shapes,
level_start_index, meta, src_padding_mask=None, rgb_views = None,
output_dir='./', frame_id = None, indices=None, threshold=0.5, indices_all=None)
source not stored for this graph (policy: none)
nothing calls this directly
no test coverage detected