Modélisation et génération sémantiquement informées des gestes co-verbaux à l'aide de modèles de langage multimodaux

Offre de thèse

Modélisation et génération sémantiquement informées des gestes co-verbaux à l'aide de modèles de langage multimodaux

Date limite de candidature

15-09-2026

Date de début de contrat

01-10-2026

Directeur de thèse

OUNI Slim

Encadrement

Suivi régulier, réunions hebdomadaires. Proposition de participation à des écoles thématiques. Participation à des conférences du domaine.

Type de contrat

ANR Financement d'Agences de financement de la recherche

école doctorale

IAEM - INFORMATIQUE - AUTOMATIQUE - ELECTRONIQUE - ELECTROTECHNIQUE - MATHEMATIQUES

équipe

MULTISPEECH

contexte

To interact naturally with Humans, Embodied Conversational Agents (ECAs) must master multimodal communication where speech is intrinsically linked to gestures, facial expressions, and posture. While recent generative models, particularly those based on diffusion, have enabled spectacular advances in terms of motor fluidity and physical realism (Alexanderson et al., 2020; Deichler et al., 2023), they still struggle to capture the deep semantics of gestures. A major bottleneck persists: semantic appropriateness. Most current systems effectively generate rhythmic gestures but often fail to produce iconic or metaphoric gestures relevant to the discourse content, as highlighted by the results of the international GENEA challenge (Kucherenko et al., 2025). This limitation stems primarily from the scarcity of semantically annotated data on a large scale. The emergence of Multimodal Large Language Models (MLLM), capable of jointly processing video, audio, and textual streams, now offers a breakthrough opportunity to automate the understanding of human gesture and drive generation (Liu et al., 2024).

spécialité

Informatique

laboratoire

LORIA - Laboratoire Lorrain de Recherche en Informatique et ses Applications

Mots clés

communication multimodale, gestes co-verbaux

Détail de l'offre

Cette thèse vise à développer de nouvelles méthodes pour la génération de gestes co-verbaux sémantiquement pertinents pour des agents conversationnels incarnés (Embodied Conversational Agents). La recherche s'appuiera sur un corpus multimodal unique de 20 heures d'interactions humaines spontanées, acquis dans le cadre du projet ANR SYNCOGEST, combinant parole, vidéo et capture de mouvement haute précision.

La thèse explorera l'utilisation des grands modèles de langage multimodaux (MLLM) comme couche intermédiaire de représentation sémantique entre la communication humaine et la génération de gestes. Elle s'articulera autour de trois axes complémentaires : (1) l'utilisation des MLLM pour annoter automatiquement les gestes selon des taxonomies linguistiques et pragmatiques établies ; (2) l'apprentissage de l'inférence d'interprétations sémantiques de niveau expert à partir de descriptions perceptives des gestes, obtenues à faible coût ; et (3) l'intégration de ces représentations sémantiques dans des modèles génératifs afin de produire des gestes davantage en adéquation avec le sens du discours.

L'objectif global est de dépasser les approches actuelles, qui génèrent principalement des gestes rythmiques ou prosodiques, pour parvenir à une génération de gestes contrôlée sémantiquement, capable de produire, en contexte, des gestes iconiques, métaphoriques et déictiques appropriés au contenu du discours.

Keywords

multimodal communication., co-verbal gesture

Subject details

This PhD aims to develop new methods for generating semantically appropriate co-speech gestures for Embodied Conversational Agents. The research will leverage a unique 20-hour multimodal corpus of spontaneous human interactions, acquired within the ANR SYNCOGEST project, combining speech, video, and high-precision motion capture. The thesis will explore Multimodal Large Language Models (MLLMs) as an intermediate semantic layer between human communication and gesture generation. It will investigate three complementary directions: (1) using MLLMs to automatically annotate gestures according to established linguistic and pragmatic taxonomies; (2) learning to infer expert-level semantic interpretations from low-cost, perceptual gesture descriptions; and (3) incorporating these semantic representations into generative models to produce gestures that are more closely aligned with discourse meaning. The overall objective is to move beyond current approaches that primarily generate prosodic or rhythmic gestures, toward semantically controlled gesture generation, capable of producing appropriate iconic, metaphoric, and deictic gestures in context.

Profil du candidat

Experience with multimodal learning, vision-LLMs or LLMs; Knowledge of
human motion modeling, gesture analysis, or embodied conversational agents ; Familiarity with generative models (e.g., diffusion, sequence-to-sequence) ; Experience with data annotation, evaluation, or human-in-the-loop methods

Candidate profile

Experience with multimodal learning, vision-LLMs or LLMs; Knowledge of
human motion modeling, gesture analysis, or embodied conversational agents ; Familiarity with generative models (e.g., diffusion, sequence-to-sequence) ; Experience with data annotation, evaluation, or human-in-the-loop methods

Référence biblio

1. Alexanderson, S., Henter, G. E., Kemantas, T., & Beskow, J. (2020). Style-controllable speech-driven gesture synthesis using normalising flows. Computer Graphics Forum, 39(2), 487–496.
2. Abel, L (2025). Co-speech gesture synthesis: Towards a controllable and interpretable model using a graph deterministic approach. Thesis.
3. Ao, T., Zhang, Z., & Liu, L. (2023).GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents. ACM Transactions on Graphics (TOG), 42(4), Article 135.
4. Deichler, A., Alexanderson, S., Mehta, S., & Beskow, J. (2023). Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation. Proceedings of the 25th International Conference on Multimodal Interaction (ICMI '23), 32–40.
5. Ghaleb, E., Burenko, I., Rasenberg, M., Pouw, W., ... & Fernández, R. (2024). Co-Speech Gesture Detection through Multi-phase Sequence Labeling. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 4007-4015.
6. Kucherenko, T., Delbosc, A., Nagy, R., Hensel, L. B., Yoon, Y., Celiktutan, O., & Henter, G. E. (2025, October).GENEA Workshop 2025: The 6th Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents. In Proceedings of the 33rd ACM International Conference on Multimedia (pp. 14305-14307).
7. Laurençon, H., Tronchon, L., Cord, M., & Sanh, V. (2024).What matters when building vision-language models?. Advances in Neural Information Processing Systems, 37, 87874-87907.
8. Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, Q., & Yuan, L. (2024). Video-LLaVA: Learning United Visual Representation by Alignment Before Projection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 23681-23691.
9. Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2024). Visual Instruction Tuning. Advances in Neural Information Processing Systems (NeurIPS 2023), 36.
10. Rohrer, P. L., Tütüncübasi, U., Florit-Pons, J., Vilà-Giménez, I., Esteve-Gibert, N., Ren-Mitchell, A.,... & Prieto, P. (2025).Multidimensional Labeling of Gesture in Communication: the M3D Proposal: PL Rohrer et al. Corpus Pragmatics, 1-23.
11. Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., ... & Lin, J. (2024). Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191.
12. Zhang, J., Zhang, Y., Cui, X., Li, S., & Wang, L. (2024). MotionGPT: Human Motion Generation with Large Language Models. arXiv preprint arXiv:2401.03456.
13. Zhang, H., Li, X., & Bing, L. (2023).Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858. Zhu, W., Cao, J., Xie, J., Yang, S., & Pang, Y.
(2024). CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation. IEEE Transactions on Circuits and Systems for Video Technology.
14. Zhu, W., Cao, J., Xie, J., Yang, S., & Pang, Y. (2024). CLIP-VIS: Adapting CLIP for open-vocabulary video instance segmentation. IEEE Transactions on Circuits and Systems for Video Technology.