Multimodal Group Emotion Recognition In-the-wild Using Privacy-Compliant Features

Multimodal Intelligence FGCMA AVEL CV ViT FREA
ICMI '23: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION, Oct 2023, Paris, France. pp.750-754
2023年12月06日
本文探讨了在EmotiW Challenge 2023中进行符合隐私的群体情绪识别,该识别是在自然环境中进行的。群体情绪识别在许多领域中都非常有用,包括社交机器人、对话代理、电子辅导和学习分析。本研究仅使用全局特征来实现,避免使用个体特征,即所有可用于在视频中识别或跟踪人员的特征(面部标记、身体姿势、音频分离等)。所提出的多模型模型由视频和音频分支组成,模态之间具有交叉注意力。视频分支基于微调的ViT架构。音频分支提取Mel频谱图,并通过CNN块馈送到变压器编码器中。我们的训练范式包括生成的合成数据集,以数据驱动的方式提高我们的模型对图像中面部表情的敏感度。广泛的实验显示了我们方法的重要性。我们的符合隐私的提议在EmotiW挑战赛中表现良好,最佳模型在验证集和测试集上的准确率分别为79.24%和75.13%。值得注意的是,我们的研究结果表明,仅使用视频中均匀分布的5帧,并使用符合隐私的特征,就可以达到这种准确性水平。
This paper explores privacy-compliant group-level emotion recognition ''in-the-wild'' within the EmotiW Challenge 2023. Group-level emotion recognition can be useful in many fields including social robotics, conversational agents, e-coaching and learning analytics. This research imposes itself using only global features avoiding individual ones, i.e. all features that can be used to identify or track people in videos (facial landmarks, body poses, audio diarization, etc.). The proposed multimodal model is composed of a video and an audio branches with a cross-attention between modalities. The video branch is based on a fine-tuned ViT architecture. The audio branch extracts Mel-spectrograms and feed them through CNN blocks into a transformer encoder. Our training paradigm includes a generated synthetic dataset to increase the sensitivity of our model on facial expression within the image in a data-driven way. The extensive experiments show the significance of our methodology. Our privacy-compliant proposal performs fairly on the EmotiW challenge, with 79.24% and 75.13% of accuracy respectively on validation and test set for the best models. Noticeably, our findings highlight that it is possible to reach this accuracy level with privacy-compliant features using only 5 frames uniformly distributed on the video.
许愿