Elite360M: Efficient 360 Multi-task Learning via Bi-projection Fusion and Cross-task Collaboration
Hao Ai, Lin Wang

TL;DR
Elite360M introduces a multi-task learning framework for 360 images that effectively combines geometry and semantics by addressing spherical distortion and enhancing global perception through innovative projections and cross-task collaboration.
Contribution
The paper presents a novel end-to-end multi-task learning framework that fuses geometry and semantics in 360 images using a bi-projection fusion and cross-task collaboration, improving global perception and task integration.
Findings
Effective multi-task learning of 3D structures and semantics from 360 images.
Enhanced global perception through ICOSAP and ERP fusion.
Superior performance demonstrated on extensive experiments.
Abstract
360 cameras capture the entire surrounding environment with a large FoV, exhibiting comprehensive visual information to directly infer the 3D structures, e.g., depth and surface normal, and semantic information simultaneously. Existing works predominantly specialize in a single task, leaving multi-task learning of 3D geometry and semantics largely unexplored. Achieving such an objective is, however, challenging due to: 1) inherent spherical distortion of planar equirectangular projection (ERP) and insufficient global perception induced by 360 image's ultra-wide FoV; 2) non-trivial progress in effectively merging geometry and semantics among different tasks to achieve mutual benefits. In this paper, we propose a novel end-to-end multi-task learning framework, named Elite360M, capable of inferring 3D structures via depth and surface normal estimation, and semantics via semantic…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsBrain Tumor Detection and Classification · Online Learning and Analytics · Face and Expression Recognition
MethodsBilinear Attention
