The field of autonomous driving increasingly demands high-quality annotated training data. In this paper we propose Panacea an innovative approach to generate panoramic and controllable videos in driving scenarios capable of yielding an unlimited numbers of diverse annotated samples pivotal for autonomous driving advancements. Panacea addresses two critical challenges:‘Consistency’and’Controllability.‘Consistency ensures temporal and cross-view coherence while Controllability ensures the alignment of generated content with corresponding annotations. Our approach integrates a novel 4D attention and a two-stage generation pipeline to maintain coherence supplemented by the ControlNet framework for meticulous control by the Bird’s-Eye-View (BEV) layouts. Extensive qualitative and quantitative evaluations of Panacea on the nuScenes dataset prove its effectiveness in generating high-quality multi-view driving-scene videos. This work notably propels the field of autonomous driving by effectively augmenting the training dataset used for advanced BEV perception techniques.
本文提出Panacea,一种面向自动驾驶场景的全景可控视频生成方法。该方法通过创新的4D注意力机制和两阶段生成流程,有效解决了生成视频中的时间与跨视角一致性问题;同时引入ControlNet框架,利用鸟瞰图(BEV)布局对生成内容进行精细控制。在nuScenes数据集上的实验表明,Panacea能够生成高质量的多视角驾驶视频,为BEV感知任务提供丰富的训练数据增强,推动自动驾驶感知技术的发展。