端到端自动驾驶规划方法研究综述:场景表示、规划生成与评估验证
A Review of End-to-End Autonomous Driving Planning Methods: Scene Representation, Planning Generation, and Evaluation and Validation
摘要: 端到端自动驾驶规划通过减少人工规则和模块接口限制,使模型能够围绕最终驾驶任务进行联合优化。然而,场景表示、辅助任务和模型结构日益复杂,并不意味着其实际驾驶能力得到充分验证,场景信息能否有效支撑规划、不同技术机制如何生成驾驶行为,以及现有评估能否可靠反映模型能力,仍需系统分析。为此,本文围绕“场景表示–规划生成–评估验证”这一主线,对端到端自动驾驶规划研究进行综述。首先,界定端到端规划的任务边界与分类依据,归纳鸟瞰视角表示、占据表示、向量化场景表示以及语义增强与潜在世界状态表示,分析其从空间关系、三维占用、道路拓扑、高层语义和未来演化等方面对规划的支撑作用。其次,依据规划结果生成与改进的主导机制,将现有方法划分为模仿学习、多任务联合学习、生成模型和视觉–语言–动作模型四类,比较其规划依据、生成机制、输出形式及能力边界。随后,比较开环评估、闭环仿真评估和语言–动作一致性评估的评价对象、主要指标与适用场景,分析其在衡量离线规划质量、连续驾驶能力和语义–行为一致性方面的适用性与局限。最后,归纳现有研究的共性挑战,并展望面向闭环可信与可验证安全的发展方向。
Abstract: End-to-end autonomous driving planning reduces reliance on manually designed rules and module interfaces, enabling models to be jointly optimized toward the final driving task. However, increasingly complex scene representations, auxiliary tasks, and model architectures do not necessarily imply that actual driving capabilities have been adequately validated. Whether scene information can effectively support planning, how different technical mechanisms generate driving behaviors, and whether existing evaluation methods can reliably reflect model capability still require systematic investigation. To address these issues, this paper reviews end-to-end autonomous driving planning along the main line of “scene representation - planning generation - evaluation and validation”. First, the task boundaries and classification criteria of end-to-end planning are defined. Major scene representations, including bird’s-eye-view representation, occupancy representation, vectorized scene representation, semantic-enhanced representation, and latent world-state representation, are summarized, with emphasis on how they support planning through spatial relationships, three-dimensional occupancy, road topology, high-level semantics, and future evolution. Second, according to the dominant mechanisms used to generate and refine planning results, existing methods are categorized into four groups: imitation learning, multi-task joint learning, generative models, and vision-language-action models. Their planning basis, generation mechanisms, output forms, and capability boundaries are then compared. Subsequently, open-loop evaluation, closed-loop simulation evaluation, and language-action consistency evaluation are compared in terms of evaluation objects, major metrics, and applicable scenarios, and their applicability and limitations in measuring offline planning quality, continuous driving capability, and semantic-behavior consistency are analyzed. Finally, the common challenges faced by existing studies are summarized, and future research directions toward closed-loop trustworthiness and verifiable safety are discussed.
文章引用:禹鑫. 端到端自动驾驶规划方法研究综述:场景表示、规划生成与评估验证[J]. 人工智能与机器人研究, 2026, 15(5): 1188-1202. https://doi.org/10.12677/airr.2026.155108

参考文献

[1] Chen, L., Wu, P., Chitta, K., Jaeger, B., Geiger, A. and Li, H. (2024) End-to-End Autonomous Driving: Challenges and Frontiers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46, 10164-10183.
https://doi.org/10.1109/tpami.2024.3435937
[2] Tampuu, A., Matiisen, T., Semikin, M., Fishman, D. and Muhammad, N. (2022) A Survey of End-to-End Driving: Architectures and Training Methods. IEEE Transactions on Neural Networks and Learning Systems, 33, 1364-1384.
https://doi.org/10.1109/tnnls.2020.3043505
[3] Codevilla, F., Muller, M., López, A., Koltun, V. and Dosovitskiy, A. (2018) End-to-End Driving via Conditional Imitation Learning. 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, 21-25 May 2018, 4693-4700.
https://doi.org/10.1109/icra.2018.8460487
[4] Codevilla, F., Santana, E., Lopez, A. and Gaidon, A. (2019) Exploring the Limitations of Behavior Cloning for Autonomous Driving. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, 27 October-2 November 2019, 9329-9338.
https://doi.org/10.1109/iccv.2019.00942
[5] Hu, S., Chen, L., Wu, P., Li, H., Yan, J. and Tao, D. (2022) ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning. In: Computer Vision-ECCV 2022, Springer Nature Switzerland, 533-549.
https://doi.org/10.1007/978-3-031-19839-7_31
[6] Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., et al. (2023) Planning-Oriented Autonomous Driving. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 18-22 June 2023, 17853-17862.
https://doi.org/10.1109/cvpr52729.2023.01712
[7] Zhai, J.T., Feng, Z., Du, J., et al. (2023) Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes. arXiv: 2305.10430.
https://arxiv.org/abs/2305.10430
[8] Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., et al. (2022) BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers. In: Computer Vision-ECCV 2022, Springer Nature Switzerland, 1-18.
https://doi.org/10.1007/978-3-031-20077-9_1
[9] Li, Y., Ge, Z., Yu, G., Yang, J., Wang, Z., Shi, Y., et al. (2023) BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object Detection. Proceedings of the AAAI Conference on Artificial Intelligence, 37, 1477-1485.
https://doi.org/10.1609/aaai.v37i2.25233
[10] Chang, M., Zhang, X., Zhang, R., Zhao, Z., He, G. and Liu, S. (2024) RecurrentBEV: A Long-Term Temporal Fusion Framework for Multi-View 3D Detection. In: Computer Vision-ECCV 2024, Springer Nature Switzerland, 131-147.
https://doi.org/10.1007/978-3-031-73220-1_8
[11] Liu, Z., Tang, H., Amini, A., Yang, X., Mao, H., Rus, D.L., et al. (2023) BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation. 2023 IEEE International Conference on Robotics and Automation (ICRA), London, 29 May-2 June 2023, 2774-2781.
https://doi.org/10.1109/icra48891.2023.10160968
[12] Jia, X., Wu, P., Chen, L., Xie, J., He, C., Yan, J., et al. (2023) Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 18-22 June 2023, 21983-21994.
https://doi.org/10.1109/cvpr52729.2023.02105
[13] Huang, Y., Zheng, W., Zhang, Y., Zhou, J. and Lu, J. (2023) Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 18-22 June 2023, 9223-9232.
https://doi.org/10.1109/cvpr52729.2023.00890
[14] Zhang, Y., Zhu, Z. and Du, D. (2023) Occformer: Dual-Path Transformer for Vision-Based 3D Semantic Occupancy Prediction. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, 2-3 October 2023, 9433-9443.
https://doi.org/10.1109/iccv51070.2023.00865
[15] Wei, Y., Zhao, L., Zheng, W., Zhu, Z., Zhou, J. and Lu, J. (2023) SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Driving. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, 2-3 October 2023, 21729-21740.
https://doi.org/10.1109/iccv51070.2023.01986
[16] Li, Y., Yu, Z., Choy, C., Xiao, C., Alvarez, J.M., Fidler, S., et al. (2023) VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene Completion. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 18-22 June 2023, 9087-9098.
https://doi.org/10.1109/cvpr52729.2023.00877
[17] Wang, Y., Chen, Y., Liao, X., Fan, L. and Zhang, Z. (2024) PanoOcc: Unified Occupancy Representation for Camera-Based 3D Panoptic Segmentation. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 16-22 June 2024, 17158-17168.
https://doi.org/10.1109/cvpr52733.2024.01624
[18] Liu, Y., Yuan, T., Wang, Y., et al. (2023) VectorMapNet: End-to-End Vectorized HD Map Learning. Proceedings of the 40th International Conference on Machine Learning. PMLR, 202, 22352-22369.
[19] Liao, B., Chen, S., Wang, X., et al. (2023) MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction. International Conference on Learning Representations, Kigali, 1-5 May 2023, 710-727.
[20] Yuan, T., Liu, Y., Wang, Y., Wang, Y. and Zhao, H. (2024) StreamMapNet: Streaming Mapping Network for Vectorized Online HD Map Construction. 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, 3-8 January 2024, 7356-7365.
https://doi.org/10.1109/wacv57701.2024.00719
[21] Wu, D., Chang, J., Jia, F., et al. (2024) TopoMLP: A Simple Yet Strong Pipeline for Driving Topology Reasoning. International Conference on Learning Representations, Vienna, 7-11 May 2024, 13809-13820.
[22] Choudhary, T., Dewangan, V., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A.K., et al. (2024) Talk2BEV: Language-Enhanced Bird’s-Eye View Maps for Autonomous Driving. 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, 13-17 May 2024, 16345-16352.
https://doi.org/10.1109/icra57147.2024.10611485
[23] Pan, C., Yaman, B., Nesti, T., Mallik, A., Allievi, A.G., Velipasalar, S., et al. (2024) VLP: Vision Language Planning for Autonomous Driving. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 16-22 June 2024, 14760-14769.
https://doi.org/10.1109/cvpr52733.2024.01398
[24] Schmidt, F., Enzweiler, M. and Valada, A. (2025) GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving. arXiv: 2511.11266.
https://arxiv.org/abs/2511.11266
[25] Zheng, W., Chen, W., Huang, Y., Zhang, B., Duan, Y. and Lu, J. (2024) OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving. In: Computer Vision-ECCV 2024, Springer Nature Switzerland, 55-72.
https://doi.org/10.1007/978-3-031-72624-8_4
[26] Min, C., Zhao, D., Xiao, L., Zhao, J., Xu, X., Zhu, Z., et al. (2024) DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vienna, 7-11 May 2024, 15522-15533.
https://doi.org/10.1109/cvpr52733.2024.01470
[27] Gao, S., Yang, J., Chen, L., Chitta, K., Qiu, Y., Geiger, A., et al. (2024) Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability. Advances in Neural Information Processing Systems 37, Vancouver, 10-15 December 2024, 91560-91596.
https://doi.org/10.52202/079017-2906
[28] Chitta, K., Prakash, A., Jaeger, B., Yu, Z., Renz, K. and Geiger, A. (2023) Transfuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45, 12878-12895.
https://doi.org/10.1109/tpami.2022.3200245
[29] Wu, P., Jia, X., Chen, L., Yan, J., Li, H. and Qiao, Y. (2022) Trajectory-Guided Control Prediction for End-to-End Autonomous Driving: A Simple Yet Strong Baseline. Advances in Neural Information Processing Systems 35, New Orleans, 28 November-9 December 2022, 6119-6132.
https://doi.org/10.52202/068431-0443
[30] Casas, S., Sadat, A. and Urtasun, R. (2021) MP3: A Unified Model to Map, Perceive, Predict and Plan. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19-25 June 2021, 14403-14412.
https://doi.org/10.1109/cvpr46437.2021.01417
[31] Shao, H., Wang, L., Chen, R., et al. (2023) Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer. Proceedings of the 6th Conference on Robot Learning. PMLR, 205, 726-737.
[32] Jiang, B., Chen, S., Xu, Q., Liao, B., Chen, J., Zhou, H., et al. (2023) VAD: Vectorized Scene Representation for Efficient Autonomous Driving. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, 1-6 October 2023, 8340-8350.
https://doi.org/10.1109/iccv51070.2023.00766
[33] Sun, W., Lin, X., Shi, Y., Zhang, C., Wu, H. and Zheng, S. (2025) SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation. 2025 IEEE International Conference on Robotics and Automation (ICRA), Atlanta, 19-23 May 2025, 8795-8801.
https://doi.org/10.1109/icra55743.2025.11128800
[34] Weng, X., Ivanovic, B., Wang, Y., Wang, Y. and Pavone, M. (2024) PARA-Drive: Parallelized Architecture for Real-Time Autonomous Driving. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 16-22 June 2024, 15449-15458.
https://doi.org/10.1109/cvpr52733.2024.01463
[35] Jia, X., You, J., Zhang, Z., et al. (2025) DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving. International Conference on Learning Representations, Singapore, 24-28 April 2025, 57196-57212.
[36] Zheng, W., Song, R., Guo, X., Zhang, C. and Chen, L. (2024) GenAD: Generative End-to-End Autonomous Driving. In: Computer Vision-ECCV 2024, Springer Nature Switzerland, 87-104.
https://doi.org/10.1007/978-3-031-73650-6_6
[37] Liao, B., Chen, S., Yin, H., Jiang, B., Wang, C., Yan, S., et al. (2025) DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 11-15 June 2025, 12037-12047.
https://doi.org/10.1109/cvpr52734.2025.01124
[38] Li, Y., Fan, L., He, J., et al. (2025) Enhancing End-to-End Autonomous Driving with Latent World Model. International Conference on Learning Representations, Singapore, 24-28 April 2025, 34849-34866.
[39] Zheng, Y., Yang, P., Xing, Z., Zhang, Q., Zheng, Y., Gao, Y., et al. (2025) World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-23 October 2025, 28632-28642.
https://doi.org/10.1109/iccv51701.2025.02659
[40] Zhang, B., Song, N., Li, J., Zhu, X., Deng, J. and Zhang, L. (2025) Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution. Advances in Neural Information Processing Systems 38, San Diego, 2-7 December 2025, 11358-11383.
https://doi.org/10.52202/085713-0345
[41] Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.K., et al. (2024) DriveGPT4: Interpretable End-to-End Autonomous Driving via Large Language Model. IEEE Robotics and Automation Letters, 9, 8186-8193.
https://doi.org/10.1109/lra.2024.3440097
[42] Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., et al. (2024) DriveLM: Driving with Graph Visual Question Answering. In: Computer Vision-ECCV 2024, Springer Nature Switzerland, 256-274.
https://doi.org/10.1007/978-3-031-72943-0_15
[43] Tian, X., Gu, J., Li, B., et al. (2025) DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models. Proceedings of the 8th Conference on Robot Learning. PMLR, 270, 4698-4726.
[44] Hwang, J.J., Xu, R., Lin, H., et al. (2025) EMMA: End-to-End Multimodal Model for Autonomous Driving. Transactions on Machine Learning Research, 1-31.
[45] Zhang, Q., Zhu, M. and Yang, H.F. (2024) Think-Driver: From Driving-Scene Understanding to Decision-Making with Vision Language Models. European Conference on Computer Vision Workshop on Autonomous Vehicles Meet Multimodal Foundation Models, Milan, 29 September 2024, 1-14.
[46] Zhou, X., Han, X., Yang, F., Ma, Y., Tresp, V. and Knoll, A. (2026) OpenDriveVLA: Towards End-to-End Autonomous Driving with Large Vision Language Action Model. Proceedings of the AAAI Conference on Artificial Intelligence, 40, 13782-13790.
https://doi.org/10.1609/aaai.v40i16.38386
[47] Zhou, Z., Cai, T., Zhao, S., Zhang, Y., Huang, Z., Zhou, B., et al. (2025) AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning. Advances in Neural Information Processing Systems 38, San Diego, 2-7 December 2025, 31725-31761.
https://doi.org/10.52202/085713-0942
[48] Renz, K., Chen, L., Arani, E. and Sinavski, O. (2025) SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 11-15 June 2025, 11993-12003.
https://doi.org/10.1109/cvpr52734.2025.01120
[49] Fang, S., Cui, Y., Liang, H., et al. (2025) CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine.
https://arxiv.org/abs/2509.15968
[50] Xie, S., Kong, L., Dong, Y., Sima, C., Zhang, W., Chen, Q.A., et al. (2025) Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-20 October 2025, 6585-6597.
https://doi.org/10.1109/iccv51701.2025.00621
[51] Fu, H., Zhang, D., Zhao, Z., Cui, J., Liang, D., Zhang, C., et al. (2025) ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-20 October 2025, 24823-24834.
https://doi.org/10.1109/iccv51701.2025.02302
[52] Liu, C., Zhu, M., Zhang, Z., Song, L., Zhao, X., Luo, Q., et al. (2025) TAD-E2E: A Large-Scale End-to-End Autonomous Driving Dataset. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-20 October 2025, 26600-26609.
https://doi.org/10.1109/iccv51701.2025.02469
[53] Dosovitskiy, A., Ros, G., Codevilla, F., et al. (2017) CARLA: An Open Urban Driving Simulator. Proceedings of the 1st Annual Conference on Robot Learning. PMLR, 78, 1-16.
[54] Caesar, H., Kabzan, J., Tan, K.S., et al. (2021) nuPlan: A Closed-Loop ML-Based Planning Benchmark for Autonomous Vehicles. CVPR Workshop on Autonomous Driving: Perception, Prediction and Planning, 19-25 June 2021, 1-5.
[55] Jia, X., Yang, Z., Li, Q., Zhang, Z. and Yan, J. (2024) Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-to-End Autonomous Driving. Advances in Neural Information Processing Systems 37, Vancouver, 10-15 December 2024, 819-844.
https://doi.org/10.52202/079017-0025
[56] Daza, I.G., Izquierdo, R., Martínez, L.M., Benderius, O. and Llorca, D.F. (2023) Sim-to-Real Transfer and Reality Gap Modeling in Model Predictive Control for Autonomous Driving. Applied Intelligence, 53, 12719-12735.
https://doi.org/10.1007/s10489-022-04148-1
[57] Qian, T., Chen, J., Zhuo, L., Jiao, Y. and Jiang, Y. (2024) NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario. Proceedings of the AAAI Conference on Artificial Intelligence, 38, 4542-4550.
https://doi.org/10.1609/aaai.v38i5.28253
[58] Wang, J., Pun, A., Tu, J., Manivasagam, S., Sadat, A., Casas, S., et al. (2021) AdvSim: Generating Safety-Critical Scenarios for Self-Driving Vehicles. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19-25 June 2021, 9909-9918.
https://doi.org/10.1109/cvpr46437.2021.00978
[59] Gil, R.C., Ornia, D.J., Mustafa, K.A. and Alonso Mora, J. (2025) Predictability Awareness for Efficient and Robust Multi-Agent Coordination. International Joint Conference on Autonomous Agents and Multiagent Systems, Detroit, 19-23 May 2025, 886-894.
https://doi.org/10.65109/hhpp1298