``
Home › Our People › Doctoral Candidates › DC12 – Muhammad Shahid
DC12

Multi-Modal Fusion for Road Infrastructure Analysis: An AI-Driven Framework for Heterogeneous Data

Muhammad Shahid
Safer mobility
for a brighter
tomorrow
Faculty of Transport and Traffic Science (FPZ) logo
iRAP logo
♧   Work Package 6
AI for Proactive Infrastructure Safety Management

Research at a glance

My research focuses on automated road safety assessment and infrastructure analysis utilizing heterogeneous road datasets, situated at the intersection of deep multimodal learning, video understanding, and trustworthy artificial intelligence. Recently, my work has advanced from standard visual classifiers toward multimodal Vision-Language Models (VLMs), teaching these systems to visually interpret road environments and generate structured safety records for a more flexible and holistic understanding of complex scenes. A core priority of my methodology is ensuring these models are dependable for real-world deployment; to achieve this, I design solutions that prevent invalid predictions, provide reliable confidence scores, and enable seamless human-in-the-loop workflows where ambiguous scenes are automatically flagged for expert review. Furthermore, my research tackles the inherent noise of real-world infrastructure data by adaptively prioritizing rare, high-risk safety hazards while intelligently filtering out human labelling inconsistencies without discarding valuable visual evidence.

Research objectives

  • Advance Automated Road Safety Auditing through Vision-Language Models: Transition road infrastructure assessment from rigid, closed-set classifiers to multimodal Vision-Language Models (VLMs), enabling the automated generation of comprehensive and structured records directly from road survey imagery.
  • Develop Robust Spatio-Temporal Video Modeling Frameworks: Design deep learning architectures that effectively capture both persistent infrastructure characteristics (e.g., lane configurations and medians) and transient, localized features (e.g., roadside hazards).
  • Tackle Extreme Data Imbalance and Imperfect Real-World Annotations: Formulate adaptive optimization strategies and field-level noise-filtering techniques that prioritize rare, safety-critical hazards over common classes while systematically mitigating human labeling inconsistencies without sacrificing usable visual data.
  • Integrate Heterogeneous Geospatial Data for Cross-Regional Generalizability: Combine complementary data sources, such as street-level RGB video, 3D LiDAR point clouds, and GIS mapping data, to resolve visual and depth ambiguities, ensuring the developed models generalize reliably across diverse, international road networks.
  • Establish Trustworthy AI Systems: Incorporate reliable confidence estimation into multimodal models to enable practical “human-in-the-loop” deployments, allowing the AI to autonomously process high-confidence road segments while systematically escalating ambiguous, high-risk cases for expert human review.

Publications

TitleAuthorsVenueYearLink
Multi-scale spatio-temporal feature aggregation for road safety attributes classification from videosShahid, M., Ševrović, M., Hassani, A., Olyslagers, M.IEEE Access2026
A Spatio-Temporal Multi-Task Framework for iRAP Road Attribute Classification from Street-Level VideosShahid, M., Ševrović, M., Olyslagers, M., Hassani, A.ACAI 2025 (IEEE)
pp. 1–6
2025
Optimizing Car Collision Detection Using Large Dashcam-Based Datasets: A Comparative Study of Pre-Trained Models and Hyperparameter ConfigurationsShahid, M., Gregurić, M., Hassani, A., Ševrović, M.Applied Sciences
15(13), 7001
2025
Auditing iRAP’s ViDA Risk Engine: A Two-Stage Surrogate Learning and Orthogonalized Heterogeneity Framework for Modelled Road SafetyShahid, M.
Co-author, with Hassani, A., Abramović, B., Ševrović, M.
Infrastructures
Vol. 11, No. 129
2026
EfficientNet-Swin Transformer for Automated iRAP Road Safety Attribute Extraction in Low- and Middle-Income CountriesShahid, M.
Co-author, with Hassani, A., Rozi, O., Ševrović, M.
ICTR 20252025

Conference contributions

TitleConferenceDateLink
Ensemble Learning for Multi-Task Road Safety Attributes ClassificationhEART Conference (14th Symposium of the European Association for Research in Transportation)
Paris, France
Sep–Oct 2026
Impact of Camera–LiDAR Data Fusion on Road Attribute Identification Using Deep Learning: An Empirical Investigation8th IRTAD International Conference
Athens, Greece
Apr 2026
A Spatio-Temporal Multi-Task Framework for iRAP Road Attribute Classification from Street-Level Videos8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI)
Nanjing, China
Dec 2025

Secondments & collaborations

Host organisationPeriodPurposeStatus
Verne (P3M)