Emotion recognition is key tech for human-computer interaction and mental health check, it got big value in affective computing and smart healthcare. Old single-modal way only use one data source, so noise mess it up easy and it hard to catch full emotional meaning. But multimodal method mix visual, voice and text data together, use how they help each other to make result much better. Still, it have some trouble: feature talk bad between modes, meaning not match well, and too much noise stay around. The LRD method split weights and fused modal features side by side, which cut down model params a lot. Tests on CMU-MOSI and CMU-MOSEI datasets show the new model boost accuracy in multimodal emotion recognition and also keep model simpler.
Show more