Offre de thèse
Modélisation et génération sémantiquement informées des gestes co-verbaux à l'aide de modèles de langage multimodaux
Date limite de candidature
15-09-2026
Date de début de contrat
01-10-2026
Directeur de thèse
OUNI Slim
Encadrement
Suivi régulier, réunions hebdomadaires. Proposition de participation à des écoles thématiques. Participation à des conférences du domaine.
Type de contrat
école doctorale
équipe
MULTISPEECHcontexte
To interact naturally with Humans, Embodied Conversational Agents (ECAs) must master multimodal communication where speech is intrinsically linked to gestures, facial expressions, and posture. While recent generative models, particularly those based on diffusion, have enabled spectacular advances in terms of motor fluidity and physical realism (Alexanderson et al., 2020; Deichler et al., 2023), they still struggle to capture the deep semantics of gestures. A major bottleneck persists: semantic appropriateness. Most current systems effectively generate rhythmic gestures but often fail to produce iconic or metaphoric gestures relevant to the discourse content, as highlighted by the results of the international GENEA challenge (Kucherenko et al., 2025). This limitation stems primarily from the scarcity of semantically annotated data on a large scale. The emergence of Multimodal Large Language Models (MLLM), capable of jointly processing video, audio, and textual streams, now offers a breakthrough opportunity to automate the understanding of human gesture and drive generation (Liu et al., 2024).spécialité
Informatiquelaboratoire
LORIA - Laboratoire Lorrain de Recherche en Informatique et ses Applications
Mots clés
communication multimodale, gestes co-verbaux
Détail de l'offre
Cette thèse vise à développer de nouvelles méthodes pour la génération de gestes co-verbaux sémantiquement pertinents pour des agents conversationnels incarnés (Embodied Conversational Agents). La recherche s'appuiera sur un corpus multimodal unique de 20 heures d'interactions humaines spontanées, acquis dans le cadre du projet ANR SYNCOGEST, combinant parole, vidéo et capture de mouvement haute précision.
La thèse explorera l'utilisation des grands modèles de langage multimodaux (MLLM) comme couche intermédiaire de représentation sémantique entre la communication humaine et la génération de gestes. Elle s'articulera autour de trois axes complémentaires : (1) l'utilisation des MLLM pour annoter automatiquement les gestes selon des taxonomies linguistiques et pragmatiques établies ; (2) l'apprentissage de l'inférence d'interprétations sémantiques de niveau expert à partir de descriptions perceptives des gestes, obtenues à faible coût ; et (3) l'intégration de ces représentations sémantiques dans des modèles génératifs afin de produire des gestes davantage en adéquation avec le sens du discours.
L'objectif global est de dépasser les approches actuelles, qui génèrent principalement des gestes rythmiques ou prosodiques, pour parvenir à une génération de gestes contrôlée sémantiquement, capable de produire, en contexte, des gestes iconiques, métaphoriques et déictiques appropriés au contenu du discours.
Keywords
multimodal communication., co-verbal gesture
Subject details
This PhD aims to develop new methods for generating semantically appropriate co-speech gestures for Embodied Conversational Agents. The research will leverage a unique 20-hour multimodal corpus of spontaneous human interactions, acquired within the ANR SYNCOGEST project, combining speech, video, and high-precision motion capture. The thesis will explore Multimodal Large Language Models (MLLMs) as an intermediate semantic layer between human communication and gesture generation. It will investigate three complementary directions: (1) using MLLMs to automatically annotate gestures according to established linguistic and pragmatic taxonomies; (2) learning to infer expert-level semantic interpretations from low-cost, perceptual gesture descriptions; and (3) incorporating these semantic representations into generative models to produce gestures that are more closely aligned with discourse meaning. The overall objective is to move beyond current approaches that primarily generate prosodic or rhythmic gestures, toward semantically controlled gesture generation, capable of producing appropriate iconic, metaphoric, and deictic gestures in context.
Profil du candidat
Experience with multimodal learning, vision-LLMs or LLMs; Knowledge of
human motion modeling, gesture analysis, or embodied conversational agents ; Familiarity with generative models (e.g., diffusion, sequence-to-sequence) ; Experience with data annotation, evaluation, or human-in-the-loop methods
Candidate profile
Experience with multimodal learning, vision-LLMs or LLMs; Knowledge of
human motion modeling, gesture analysis, or embodied conversational agents ; Familiarity with generative models (e.g., diffusion, sequence-to-sequence) ; Experience with data annotation, evaluation, or human-in-the-loop methods
Référence biblio
1. Alexanderson, S., Henter, G. E., Kemantas, T., & Beskow, J. (2020). Style-controllable speech-driven gesture synthesis using normalising flows. Computer Graphics Forum, 39(2), 487–496.
2. Abel, L (2025). Co-speech gesture synthesis: Towards a controllable and interpretable model using a graph deterministic approach. Thesis.
3. Ao, T., Zhang, Z., & Liu, L. (2023).GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents. ACM Transactions on Graphics (TOG), 42(4), Article 135.
4. Deichler, A., Alexanderson, S., Mehta, S., & Beskow, J. (2023). Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation. Proceedings of the 25th International Conference on Multimodal Interaction (ICMI '23), 32–40.
5. Ghaleb, E., Burenko, I., Rasenberg, M., Pouw, W., ... & Fernández, R. (2024). Co-Speech Gesture Detection through Multi-phase Sequence Labeling. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 4007-4015.
6. Kucherenko, T., Delbosc, A., Nagy, R., Hensel, L. B., Yoon, Y., Celiktutan, O., & Henter, G. E. (2025, October).GENEA Workshop 2025: The 6th Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. In Proceedings of the 33rd ACM International Conference on Multimedia (pp. 14305-14307).
7. Laurençon, H., Tronchon, L., Cord, M., & Sanh, V. (2024).What matters when building vision-language models?. Advances in Neural Information Processing Systems, 37, 87874-87907.
8. Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, Q., & Yuan, L. (2024). Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 23681-23691.
9. Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2024). Visual Instruction Tuning. Advances in Neural Information Processing Systems (NeurIPS 2023), 36.
10. Rohrer, P. L., Tütüncübasi, U., Florit-Pons, J., Vilà-Giménez, I., Esteve-Gibert, N., Ren-Mitchell, A.,... & Prieto, P. (2025).Multidimensional Labeling of Gesture in Communication: the M3D Proposal: PL Rohrer et al. Corpus Pragmatics, 1-23.
11. Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., ... & Lin, J. (2024). Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191.
12. Zhang, J., Zhang, Y., Cui, X., Li, S., & Wang, L. (2024). MotionGPT: Human Motion Generation with Large Language Models. arXiv preprint arXiv:2401.03456.
13. Zhang, H., Li, X., & Bing, L. (2023).Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858. Zhu, W., Cao, J., Xie, J., Yang, S., & Pang, Y.
(2024). CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation. IEEE Transactions on Circuits and Systems for Video Technology.
14. Zhu, W., Cao, J., Xie, J., Yang, S., & Pang, Y. (2024). CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation. IEEE Transactions on Circuits and Systems for Video Technology.

