ACM MM 2026

LATTE: Language-Driven Construction of 3D Driving Scenes for Closed-Loop AD Simulation

Constructing explicit 3D driving scenarios from open-world language requests, with asset grounding, layout optimization, and closed-loop autonomous-driving evaluation.

Zaiyi Hu1, Haifeng Wu1, Yiyi Liao2, Lixin Duan1, Wen Li1
1University of Electronic Science and Technology of China 2Zhejiang University

Motivation

Comparison of video generative models, agentic 3D editing, and LATTE.
LATTE links language understanding, open-world asset acquisition, structurally coherent layout optimization, and closed-loop simulation.

Closed-loop simulation is central to autonomous driving safety evaluation, yet turning natural-language requests into valid 3D test scenarios remains difficult. LATTE decomposes this problem into two coupled stages: Image-Grounded Asset Construction (IGAC), which acquires insertion-ready 3D assets for missing objects from visual evidence, and Structurally Coherent Layout Optimization (SCLO), which turns language-guided placements into simulation-suitable layouts under physical and semantic constraints. The resulting edits are written back to a real-scene-reconstructed simulator and evaluated through closed-loop AD rollouts. Experiments show that LATTE supports open-world asset insertion, produces more coherent scene layouts, and creates behaviorally effective test cases that measurably alter downstream AD performance.

Method

Results

Limitations & Future Directions

Citation

@inproceedings{hu2026latte,
  title     = {LATTE: Language-Driven Construction of 3D Driving Scenes for Closed-Loop AD Simulation},
  author    = {Hu, Zaiyi and Wu, Haifeng and Liao, Yiyi and Duan, Lixin and Li, Wen},
  booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
  year      = {2026}
}