Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection (2407.15354v1)

Published 22 Jul 2024 in cs.CV and cs.RO

Abstract: The Bird's-Eye-View (BEV) representation is a critical factor that directly impacts the 3D object detection performance, but the traditional BEV grid representation induces quadratic computational cost as the spatial resolution grows. To address this limitation, we present a new camera-based 3D object detector with high-resolution vector representation: VectorFormer. The presented high-resolution vector representation is combined with the lower-resolution BEV representation to efficiently exploit 3D geometry from multi-camera images at a high resolution through our two novel modules: vector scattering and gathering. To this end, the learned vector representation with richer scene contexts can serve as the decoding query for final predictions. We conduct extensive experiments on the nuScenes dataset and demonstrate state-of-the-art performance in NDS and inference time. Furthermore, we investigate query-BEV-based methods incorporated with our proposed vector representation and observe a consistent performance improvement.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Zhili Chen (51 papers)
  2. Shuangjie Xu (16 papers)
  3. Maosheng Ye (12 papers)
  4. Zian Qian (5 papers)
  5. Xiaoyi Zou (5 papers)
  6. Dit-Yan Yeung (78 papers)
  7. Qifeng Chen (188 papers)

Summary

We haven't generated a summary for this paper yet.