Aeronáutica y espacio

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
- Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is...
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
- Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning...
Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving
- Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for...
DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving
- Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across perception, language, and planning. However,...
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
- Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical...
Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving
- Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world models reduce the reliance on costly environment...
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
- Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language...
4D-WAM: 4D Consistent World Modeling for Autonomous Driving
- Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs...
The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software
- For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can provide strong statistical support for...
Unsupervised Point Cloud Registration with Self-Distillation
- Rigid point cloud registration is a fundamental problem and highly relevant in robotics and autonomous driving. Nowadays deep learning methods can be trained to match a pair of point clouds, given...
FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
- Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical...
Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives
- Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP)...
GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
- Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry...
Beyond the Plane: Coupling Planar Vehicle Dynamics with Three-Dimensional Road Geometry
- Simulation is crucial for developing and testing autonomous driving systems. In particular, the development of localization and control algorithms relies on an accurate vehicle dynamics simulation....
DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving
- Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions...
CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
- Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a...
How Roadside Units Enhance Intersection Safety? Cooperative Autonomous Driving System Design and A Proof of Concept
- Intersections remain one of the most hazardous locations in urban road networks, where heterogeneous traffic participants and limited visibility frequently lead to severe traffic conflicts. In this...
UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes
- End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize poorly to unstructured mining environments. In...
Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis
- Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine...
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving
- Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed...
Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles
- Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow...