DANCE: A Lightweight Dynamic Adapter Network for Real-Time Traffic Accident Anticipation
Ieee AccessPeer ReviewedHao-Yu Yen +22026Magazines
Traffic accident anticipation is critical for proactive safety systems, yet remains challenging due to rapidly evolving scenes, transient accident precursors, and strict low-latency inference requirements. In this paper, we propose Dynamic Adapters for Neural CLIP Enhancement (DANCE), a lightweight end-to-end framework for real-time traffic accident anticipation. DANCE adopts a dual-branch architecture based on Contrastive Language–Image Pre-training (CLIP) and integrates a set of dynamic adapters to enable efficient task adaptation. To accommodate diverse driving scenarios, we introduce two input-aware mechanisms within the adapters: Dynamic Gating (DyG) and Dynamic Scaling (DyS). DyG acts as a content-conditioned soft routing mechanism that selectively emphasizes salient tokens while suppressing irrelevant noise, whereas DyS adaptively modulates the strength of residual contributions to balance pre-trained knowledge preservation and task-specific learning. Operating synergistically, these mechanisms enhance temporal sensitivity and semantic stability without compromising computational efficiency. In addition, lightweight adapters in the text encoder improve the alignment between risk-related textual queries and visual tokens, while a temporal refinement head stabilizes predictions over time. Two auxiliary objectives—a vision–text Contrastive Loss (CLoss) and a Gate Sparsity Loss (GSLoss)—are introduced to strengthen cross-modal consistency and improve feature selectivity.We also provide a discussion on framework properties, including the computational complexity, convergence, and stability of the proposed framework. Extensive experiments on standard benchmark datasets demonstrate that DANCE achieves competitive accuracy with improved early anticipation performance using only RGB inputs. Importantly, the proposed method achieves 41.19 FPS on the NVIDIA Jetson Orin Nano, demonstrating a favorable trade-off between accuracy and efficiency among CLIP-based RGB–text approaches for real-time edge deployment.
The content you want is available to Zendy users.
Already have an account? Sign inHaving issues? Contact support