Agentic scene construction framework.
LATTE uses a Manager Agent, Scene Editor Agent, and Layout Agent to translate language requests into executable 3D scene edit programs, then commits the edited scene to a closed-loop simulation environment.
ACM MM 2026
Constructing explicit 3D driving scenarios from open-world language requests, with asset grounding, layout optimization, and closed-loop autonomous-driving evaluation.
Closed-loop simulation is central to autonomous driving safety evaluation, yet turning natural-language requests into valid 3D test scenarios remains difficult. LATTE decomposes this problem into two coupled stages: Image-Grounded Asset Construction (IGAC), which acquires insertion-ready 3D assets for missing objects from visual evidence, and Structurally Coherent Layout Optimization (SCLO), which turns language-guided placements into simulation-suitable layouts under physical and semantic constraints. The resulting edits are written back to a real-scene-reconstructed simulator and evaluated through closed-loop AD rollouts. Experiments show that LATTE supports open-world asset insertion, produces more coherent scene layouts, and creates behaviorally effective test cases that measurably alter downstream AD performance.
A compact view of the language-agentic construction pipeline.
Qualitative, asset, and closed-loop evidence in one compact panel.
Paired render and BEV simulation videos for each language-driven edit.
LATTE turns language requests into executable 3D driving-scene edits, while still inheriting open challenges from agent reasoning, open-world asset construction, and scene rendering.
@inproceedings{hu2026latte,
title = {LATTE: Language-Driven Construction of 3D Driving Scenes for Closed-Loop AD Simulation},
author = {Hu, Zaiyi and Wu, Haifeng and Liao, Yiyi and Duan, Lixin and Li, Wen},
booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
year = {2026}
}