Intangible cultural heritage is highly process-oriented, embodied, and context-dependent. Conventional text, video, and static 3D displays have difficulty presenting the continuous relationships among craft actions, object states, and cultural meanings. To address dynamic representation in immersive exhibitions, publicly available images, videos, and craft descriptions of Yangliuqing Woodblock New Year Pictures from 2018 to 2025 were used to construct a multimodal dataset. A lightweight digital twin state model was developed through 3D data processing, MediaPipe-based pose extraction, Dynamic Time Warping, and finite state modelling. A controlled VR experiment then compared a digital twin immersive condition with a conventional 3D exhibition condition. Simulated validation results showed that action recognition achieved a Macro F1 of 0.908 ± 0.026, while state transition accuracy reached 95.6%. Within-stage DTW distances were substantially lower than cross-stage distances. Participants in the digital twin condition also achieved greater knowledge gains, shorter task completion times, and fewer operational errors without a significant increase in cognitive load. These results indicate that organising geometric objects, action sequences, and semantic information into a continuous state chain can improve the computational representation of intangible cultural heritage processes. The approach provides a feasible methodological basis for immersive digital exhibitions that balance technical implementation, learning efficiency, and cultural knowledge communication.
Research Article
Open Access