- Toulouse - 31
- CDD
- Doctorat.Gouv.Fr
📑 Missions du poste
Établissement : Université de Toulouse École doctorale : EDMITT - Ecole Doctorale Mathématiques, Informatique et Télécommunications de Toulouse Laboratoire de recherche : IRIT : Institut de Recherche en Informatique de Toulouse Direction de la thèse : Emmanuelle CLAEYS ORCID 0009000052357632 Début de la thèse : 2027-09-01 Date limite de candidature : 2026-11-16T23:59:59 L'accompagnement par un robot d'une personne en situation de handicape (par exemple ayant une déficience visuelle) dans un espace fréquenté suppose bien plus que l'évitement d'obstacles. Le robot doit rester aux côtés de la personne à une distance confortable, alterner entre une allure lente ou irrégulière, maintenir son attention sans être intrusif, et partager le couloir avec des passants. La difficulté centrale est que les humains réagissent au robot : si celui-ci se décale à gauche, l'humain en face s'adapte aussi. Prédiction et décision sont donc indissociables, on ne peut anticiper le comportement des personnes sans connaître le futur mouvement du robot. Les systèmes actuels contournent ce couplage : certains ne réagissent qu'à l'instant présent, ce qui les rend maladroits dans les espaces contraints ; d'autres prédisent séparément les trajectoires humaines puis planifient contre cette prévision figée. Cette stratégie obtient des blocages en foule ou des trajectoires longeant les murs. Les approches calculant conjointement les futurs du robot et des humains deviennent trop coûteuses dès que le nombre de personnes et l'horizon augmentent.
Cette thèse propose que le robot anticipe avant d'agir. Un World modèle est un réseau qui apprend à répondre à la question « si je fais cette action, à quoi ressemblera la scène dans quelques secondes ? ». Les World models fonctionne comme un simulateur mental, dans lequel plusieurs actions candidates peuvent être déroulées. Les action sont choisies selon un futur plus sûr et plus prévisible. Les World models existants on encore des limitations dans notre contexte d'assistance. Ils compressent généralement la scène en un unique vecteur abstrait, perdant le détail spatial nécessaire pour juger si un passage est franchissable ou une personne trop proche. D'autre part, la littérature reste très limiter lorsqu'il sont couplés avec de l'apprentissage par renforcement (RL) pour des tâches d'assistance.
Le projet traite les deux points. La structure spatiale est préservée en prédisant dans l'espace de représentation d'un extracteur visuel gelé qui encode déjà la géométrie, et la dynamique temporelle est modélisée par un Transformer. À partir d'un unique état interne, le modèle prédit simultanément la géométrie de la scène (profondeur), les trajectoires des personnes alentour, et la réaction de la personne accompagnée (est-elle toujours engagée, le robot est-il trop proche). La politique du robot est ensuite apprise par renforcement avec deux ajouts : le futur prédit est fourni à la politique comme entrée supplémentaire, afin que les décisions reposent sur l'anticipation et non sur la seule image courante ; et la réaction humaine prédite est convertie en pénalité, dissuadant le robot d'actions qui encombreraient, surprendraient ou perdraient la personne (avant même de les exécuter). Parce que le modèle expose l'incertitude de ses prédictions, une faible confiance déclenche un repli conservateur (par exemple ralentir, s'arrêter, demander une clarification) plutôt qu'une décision confiante et erronée.
La plateforme cible est l'humanoïde Pepper (déjà à disposition), qui agit en se déplaçant par le regard et la parole : l'anticipation doit donc couvrir le regard et le langage, pas uniquement le mouvement. Les verrous principaux sont le réalisme de la simulation sociale, l'apprentissage sans épuiser des utilisateurs réels en situation de handicap, le filtrage des comportements socialement inappropriés avant toute exposition à un usager, et l'adaptation à des préférences individuelles très contrastées. L'évaluation combinera métriques de navigation, intégrant la qualité de l'accompagnement. Cette thèse inclue des collaborateurs du milieu hospitalier sur le site de Toulouse et d'Adelaïde. Latent world models [19, 20] and their Transformer-based successors [5] have shown that a learned simulator can support long-term imagination and that policies trained inside it transfer to the real world. Their usual design, however, compresses the entire scene into a single abstract vector, eliminating the spatial detail needed to judge whether a gap is traversable or a person is too close. Self-supervised backbone networks [27, 36] and monocular depth models [35] now provide feature spaces that already encode geometry, making prediction in feature space a credible alternative to a single global latent [36]. This coupling has now been demonstrated for social navigation: NavThinker [12] learns an action-conditioned world model in the patch feature space of Depth Anything V2 with depth and trajectory decoders, injects the imagined future into an on-policy DD-PPO learner via feature fusion and trajectory-based reward shaping, and reports state-of-the-art results on Social-HM3D with transfer to a quadruped platform. Its ablations establish that task-aligned decoders and both foresight injection mechanisms contribute, which, positively, settles the feasibility question that the present proposal would otherwise have had to resolve.
Social navigation has, meanwhile, moved from reactive avoidance to explicitly predictive schemes, combining large-scale reinforcement learning [32, 34] with proxemic and formation-aware control [21, 30, 16], shared evaluation protocols [18] and real human motion datasets [24, 23]; tightly coupled formulations jointly reason over the futures of the robot and humans but do not scale [15]. Yet pedestrians are treated as mobile obstacles whose behavior is independent of the robot, and the assisted traveler, when taken into account at all, is modeled geometrically rather than as an agent whose engagement and comfort react to what the robot does. In assistive robotics, it is known that acceptance depends on legible and individually adapted behavior [17]; platforms such as Pepper act through gaze and speech as much as through motion [28]; safety requirements are normative [22]; and comfort can be measured with validated instruments [2], but interaction and communication remain scripted rather than anticipated, and the gap with a real robot constrained by latency persists even with powerful simulators.
The research gap is therefore not action-conditioned foresight per se, but a joint human-robot world model for assistive navigation. Existing future-aware systems, including [12, 13], are designed for goal-point navigation among pedestrians: the robot has no companion, no communication channel, a discrete action space with four actions, and no notion of when its own predictions should not be followed. Assistance changes all four aspects. The robot must predict how the assisted person reacts (engagement, comfort, tolerated proximity), not only where pedestrians will be; it must anticipate the effect of gaze and speech as much as that of motion; it must know when its predictions are not reliable, since imagination-based decision-making is exposed to cumulative model error [14], and here the cost of a confident but wrong action falls on a vulnerable user; and uncertainty must be able to trigger clarification, slowing down, stopping, or a conservative fallback, while preserving the traveler's authority over the assistance.
Approach
We want the robot to imagine before acting. In machine learning, a world model is a network that learns to answer the question: if I do this, what will the scene look like in a few seconds? It is a learned mental simulator: the robot can pass several candidate actions through it, see which imagined future looks safest and most comfortable, and only then act.
The thesis takes the architecture validated in [12] as a starting point rather than as a proper contribution: prediction in the patch feature space of a frozen geometric backbone [27, 35, 36], temporal modeling by a causal Transformer [5], task-aligned decoders, and foresight injected into an on-policy learner [32, 34] both as anticipated lookahead features and as reward shaping [26]. Adopting a ablated design reduces first-year risk and provides a reproducible baseline. The thesis then extends it along three axes that assistance demands and that no existing system addresses:
A third prediction head for the assisted person. Beyond scene geometry and pedestrian trajectories, the model predicts the assisted traveler's response, are they still engaged, are they comfortable, is the robot too close, so that the shaped penalty is driven by the anticipated reaction of the assisted person, and not only by separation from other pedestrians. This transforms the companion from a geometric constraint into a modeled agent.
Communication as anticipated action. The action space is extended beyond discrete motion to include gaze and speech, through which the target platform [28] acts as much as through motion. The world model must therefore predict the effect of looking at or speaking to the traveler, and the policy must be discouraged from staring at or talking over someone before doing so.
Calibrated uncertainty and a fallback it authorizes. The model estimates when its own predictions are not reliable and routes low confidence toward clarification, slowing down, stopping, or a conservative controller. This addresses the exposure to cumulative model error in imagination-based decision-making [14] and is what makes the approach defensible in front of a vulnerable user.
A fourth cross-cutting requirement is that behavior be adjustable to a user profile from a few interactions, since the assistance one person wants is what another experiences as intrusion.
Main Difficulties
Realistic social simulation. The model must learn attention, engagement, and personal-space behavior well enough that its imagined futures are trustworthy as training signals, and not merely visually plausible.
Learning without exhausting real users. Reinforcement learning is data-hungry, and here every real trial involves a person with a disability. The learned simulator is precisely what offers us this economy.
From simulator to real robot. Pepper's depth camera is coarse, its speed is limited, and its commands arrive with delay. Policies must survive this gap.
Safety and social appropriateness. Inappropriate behaviors (coming too close, staring, talking over someone) must be intercepted by filtering imagined rollouts before the robot is placed in front of a real user [21].
Different people want different things. One user may want the robot close and chatty; another may prefer distance and silence. A fixed behavior cannot serve both, so the model must be adjustable to a user profile from a few interactions [17]. The overarching goal is to develop an uncertainty-aware world model enabling a robot to provide safe, comfortable, and personalized navigation assistance to people with disabilities in busy indoor environments. The target population is defined by the type of assistance required rather than by a single disability: a sensory, motor, or cognitive impairment alters the preferred interpersonal distance, walking pace, and communication channel, and the model is designed to be conditioned on this variation rather than calibrated for a single profile.
Four specific objectives follow from this:
- Learn an action-conditioned model that jointly predicts the future motion of the robot, the assisted traveler, and nearby pedestrians;
- Use these predictions to select socially appropriate actions that maintain a comfortable position alongside the traveler, adapt to their pace, and avoid disturbing other corridor users;
- Quantify predictive uncertainty, so that the robot can slow down, request clarification, stop, or hand over control to a conservative safety controller when its predictions are unreliable;
- Evaluate the approach against reactive navigation, fixed human motion prediction, and joint trajectory planning Methodology
The work will be organised in four overlapping stages, so that user requirements inform the model early and robot constraints inform it before the architecture is frozen. The choices sketched below are indicative: several options remain open at each stage and will be settled as the work progresses.
Stage 1 - User-centred design and requirement elicitation. Working with occupational therapists, mobility instructors, caregivers and users with different disabilities, we will characterise what good assistance means in measurable terms: preferred side and standing distance, tolerated pace, whether guidance should be spoken, shown or purely spatial, and what the traveller expects to remain in control of. These sessions should yield a set of scripted indoor scenarios (corridor crossing, doorway, queue, sudden obstacle), an initial parameterisation of the user-profile vector that will later condition the policy, and the ethics protocol for the subsequent studies. The protocol will be approved by the institutional ethics committee before any trial, with accessible consent, caregiver involvement and an unconditional right to stop.
Stage 2 - World-model development. The predictive core will be trained offline, first on real human-motion datasets [24, 23] and on trajectories logged in an embodied simulator [29], then on data recorded with the robot itself. One natural entry point would be to reproduce [12] on Social-HM3D, giving a verified starting point and a published number to improve on, before adding the traveller-response head and the communication actions incrementally. Ablations will isolate the contribution of spatial structure, of action conditioning and of the human-response head. Predictive uncertainty will be estimated explicitly and calibrated, since it is what later authorises the robot to fall back rather than guess; candidate methods include epistemic uncertainty estimation and open-set recognition developed for robotic perception [8, 9], transposed here to social prediction.
Stage 3 - Policy learning and integration on the robot. Imagined rollouts will be screened against proxemic and normative safety criteria [21, 22] before any behaviour is exposed to a participant, and low model confidence will be routed to a conservative controller: slow, stop or ask. Detecting at run time that the model has left its training distribution may follow the introspection and failure-detection approach developed for perception systems [10, 11], applied here to the world model rather than to a detector. A purely geometric safety layer will remain active at all times and will never depend on the learned model. The policy will then be transferred to Pepper, where coarse depth sensing, limited speed and command latency will be treated as an explicit domain-gap problem, domain randomisation and latency-aware action modelling being the first options to consider. Gaze and speech actions will be integrated at this stage, so that anticipation covers communication and not motion alone.
Stage 4 - Evaluation. Evaluation will proceed from simulation to able-bodied volunteers and finally, under the Stage 1 protocol and with caregiver supervision, to participants with disabilities. Beyond the usual navigation figures, we will measure how well the robot holds a comfortable formation, how often it intrudes on personal space, how much it disturbs other humans and how often it talks over someone, alongside the prediction accuracy and calibration of the model itself. Perceived safety, comfort and control will be collected with established questionnaires [2], and the comparison against the baselines will follow published social-navigation evaluation guidelines [18]. A final study will test personalisation across contrasting user profiles.
Cette thèse propose que le robot anticipe avant d'agir. Un World modèle est un réseau qui apprend à répondre à la question « si je fais cette action, à quoi ressemblera la scène dans quelques secondes ? ». Les World models fonctionne comme un simulateur mental, dans lequel plusieurs actions candidates peuvent être déroulées. Les action sont choisies selon un futur plus sûr et plus prévisible. Les World models existants on encore des limitations dans notre contexte d'assistance. Ils compressent généralement la scène en un unique vecteur abstrait, perdant le détail spatial nécessaire pour juger si un passage est franchissable ou une personne trop proche. D'autre part, la littérature reste très limiter lorsqu'il sont couplés avec de l'apprentissage par renforcement (RL) pour des tâches d'assistance.
Le projet traite les deux points. La structure spatiale est préservée en prédisant dans l'espace de représentation d'un extracteur visuel gelé qui encode déjà la géométrie, et la dynamique temporelle est modélisée par un Transformer. À partir d'un unique état interne, le modèle prédit simultanément la géométrie de la scène (profondeur), les trajectoires des personnes alentour, et la réaction de la personne accompagnée (est-elle toujours engagée, le robot est-il trop proche). La politique du robot est ensuite apprise par renforcement avec deux ajouts : le futur prédit est fourni à la politique comme entrée supplémentaire, afin que les décisions reposent sur l'anticipation et non sur la seule image courante ; et la réaction humaine prédite est convertie en pénalité, dissuadant le robot d'actions qui encombreraient, surprendraient ou perdraient la personne (avant même de les exécuter). Parce que le modèle expose l'incertitude de ses prédictions, une faible confiance déclenche un repli conservateur (par exemple ralentir, s'arrêter, demander une clarification) plutôt qu'une décision confiante et erronée.
La plateforme cible est l'humanoïde Pepper (déjà à disposition), qui agit en se déplaçant par le regard et la parole : l'anticipation doit donc couvrir le regard et le langage, pas uniquement le mouvement. Les verrous principaux sont le réalisme de la simulation sociale, l'apprentissage sans épuiser des utilisateurs réels en situation de handicap, le filtrage des comportements socialement inappropriés avant toute exposition à un usager, et l'adaptation à des préférences individuelles très contrastées. L'évaluation combinera métriques de navigation, intégrant la qualité de l'accompagnement. Cette thèse inclue des collaborateurs du milieu hospitalier sur le site de Toulouse et d'Adelaïde. Latent world models [19, 20] and their Transformer-based successors [5] have shown that a learned simulator can support long-term imagination and that policies trained inside it transfer to the real world. Their usual design, however, compresses the entire scene into a single abstract vector, eliminating the spatial detail needed to judge whether a gap is traversable or a person is too close. Self-supervised backbone networks [27, 36] and monocular depth models [35] now provide feature spaces that already encode geometry, making prediction in feature space a credible alternative to a single global latent [36]. This coupling has now been demonstrated for social navigation: NavThinker [12] learns an action-conditioned world model in the patch feature space of Depth Anything V2 with depth and trajectory decoders, injects the imagined future into an on-policy DD-PPO learner via feature fusion and trajectory-based reward shaping, and reports state-of-the-art results on Social-HM3D with transfer to a quadruped platform. Its ablations establish that task-aligned decoders and both foresight injection mechanisms contribute, which, positively, settles the feasibility question that the present proposal would otherwise have had to resolve.
Social navigation has, meanwhile, moved from reactive avoidance to explicitly predictive schemes, combining large-scale reinforcement learning [32, 34] with proxemic and formation-aware control [21, 30, 16], shared evaluation protocols [18] and real human motion datasets [24, 23]; tightly coupled formulations jointly reason over the futures of the robot and humans but do not scale [15]. Yet pedestrians are treated as mobile obstacles whose behavior is independent of the robot, and the assisted traveler, when taken into account at all, is modeled geometrically rather than as an agent whose engagement and comfort react to what the robot does. In assistive robotics, it is known that acceptance depends on legible and individually adapted behavior [17]; platforms such as Pepper act through gaze and speech as much as through motion [28]; safety requirements are normative [22]; and comfort can be measured with validated instruments [2], but interaction and communication remain scripted rather than anticipated, and the gap with a real robot constrained by latency persists even with powerful simulators.
The research gap is therefore not action-conditioned foresight per se, but a joint human-robot world model for assistive navigation. Existing future-aware systems, including [12, 13], are designed for goal-point navigation among pedestrians: the robot has no companion, no communication channel, a discrete action space with four actions, and no notion of when its own predictions should not be followed. Assistance changes all four aspects. The robot must predict how the assisted person reacts (engagement, comfort, tolerated proximity), not only where pedestrians will be; it must anticipate the effect of gaze and speech as much as that of motion; it must know when its predictions are not reliable, since imagination-based decision-making is exposed to cumulative model error [14], and here the cost of a confident but wrong action falls on a vulnerable user; and uncertainty must be able to trigger clarification, slowing down, stopping, or a conservative fallback, while preserving the traveler's authority over the assistance.
Approach
We want the robot to imagine before acting. In machine learning, a world model is a network that learns to answer the question: if I do this, what will the scene look like in a few seconds? It is a learned mental simulator: the robot can pass several candidate actions through it, see which imagined future looks safest and most comfortable, and only then act.
The thesis takes the architecture validated in [12] as a starting point rather than as a proper contribution: prediction in the patch feature space of a frozen geometric backbone [27, 35, 36], temporal modeling by a causal Transformer [5], task-aligned decoders, and foresight injected into an on-policy learner [32, 34] both as anticipated lookahead features and as reward shaping [26]. Adopting a ablated design reduces first-year risk and provides a reproducible baseline. The thesis then extends it along three axes that assistance demands and that no existing system addresses:
A third prediction head for the assisted person. Beyond scene geometry and pedestrian trajectories, the model predicts the assisted traveler's response, are they still engaged, are they comfortable, is the robot too close, so that the shaped penalty is driven by the anticipated reaction of the assisted person, and not only by separation from other pedestrians. This transforms the companion from a geometric constraint into a modeled agent.
Communication as anticipated action. The action space is extended beyond discrete motion to include gaze and speech, through which the target platform [28] acts as much as through motion. The world model must therefore predict the effect of looking at or speaking to the traveler, and the policy must be discouraged from staring at or talking over someone before doing so.
Calibrated uncertainty and a fallback it authorizes. The model estimates when its own predictions are not reliable and routes low confidence toward clarification, slowing down, stopping, or a conservative controller. This addresses the exposure to cumulative model error in imagination-based decision-making [14] and is what makes the approach defensible in front of a vulnerable user.
A fourth cross-cutting requirement is that behavior be adjustable to a user profile from a few interactions, since the assistance one person wants is what another experiences as intrusion.
Main Difficulties
Realistic social simulation. The model must learn attention, engagement, and personal-space behavior well enough that its imagined futures are trustworthy as training signals, and not merely visually plausible.
Learning without exhausting real users. Reinforcement learning is data-hungry, and here every real trial involves a person with a disability. The learned simulator is precisely what offers us this economy.
From simulator to real robot. Pepper's depth camera is coarse, its speed is limited, and its commands arrive with delay. Policies must survive this gap.
Safety and social appropriateness. Inappropriate behaviors (coming too close, staring, talking over someone) must be intercepted by filtering imagined rollouts before the robot is placed in front of a real user [21].
Different people want different things. One user may want the robot close and chatty; another may prefer distance and silence. A fixed behavior cannot serve both, so the model must be adjustable to a user profile from a few interactions [17]. The overarching goal is to develop an uncertainty-aware world model enabling a robot to provide safe, comfortable, and personalized navigation assistance to people with disabilities in busy indoor environments. The target population is defined by the type of assistance required rather than by a single disability: a sensory, motor, or cognitive impairment alters the preferred interpersonal distance, walking pace, and communication channel, and the model is designed to be conditioned on this variation rather than calibrated for a single profile.
Four specific objectives follow from this:
- Learn an action-conditioned model that jointly predicts the future motion of the robot, the assisted traveler, and nearby pedestrians;
- Use these predictions to select socially appropriate actions that maintain a comfortable position alongside the traveler, adapt to their pace, and avoid disturbing other corridor users;
- Quantify predictive uncertainty, so that the robot can slow down, request clarification, stop, or hand over control to a conservative safety controller when its predictions are unreliable;
- Evaluate the approach against reactive navigation, fixed human motion prediction, and joint trajectory planning Methodology
The work will be organised in four overlapping stages, so that user requirements inform the model early and robot constraints inform it before the architecture is frozen. The choices sketched below are indicative: several options remain open at each stage and will be settled as the work progresses.
Stage 1 - User-centred design and requirement elicitation. Working with occupational therapists, mobility instructors, caregivers and users with different disabilities, we will characterise what good assistance means in measurable terms: preferred side and standing distance, tolerated pace, whether guidance should be spoken, shown or purely spatial, and what the traveller expects to remain in control of. These sessions should yield a set of scripted indoor scenarios (corridor crossing, doorway, queue, sudden obstacle), an initial parameterisation of the user-profile vector that will later condition the policy, and the ethics protocol for the subsequent studies. The protocol will be approved by the institutional ethics committee before any trial, with accessible consent, caregiver involvement and an unconditional right to stop.
Stage 2 - World-model development. The predictive core will be trained offline, first on real human-motion datasets [24, 23] and on trajectories logged in an embodied simulator [29], then on data recorded with the robot itself. One natural entry point would be to reproduce [12] on Social-HM3D, giving a verified starting point and a published number to improve on, before adding the traveller-response head and the communication actions incrementally. Ablations will isolate the contribution of spatial structure, of action conditioning and of the human-response head. Predictive uncertainty will be estimated explicitly and calibrated, since it is what later authorises the robot to fall back rather than guess; candidate methods include epistemic uncertainty estimation and open-set recognition developed for robotic perception [8, 9], transposed here to social prediction.
Stage 3 - Policy learning and integration on the robot. Imagined rollouts will be screened against proxemic and normative safety criteria [21, 22] before any behaviour is exposed to a participant, and low model confidence will be routed to a conservative controller: slow, stop or ask. Detecting at run time that the model has left its training distribution may follow the introspection and failure-detection approach developed for perception systems [10, 11], applied here to the world model rather than to a detector. A purely geometric safety layer will remain active at all times and will never depend on the learned model. The policy will then be transferred to Pepper, where coarse depth sensing, limited speed and command latency will be treated as an explicit domain-gap problem, domain randomisation and latency-aware action modelling being the first options to consider. Gaze and speech actions will be integrated at this stage, so that anticipation covers communication and not motion alone.
Stage 4 - Evaluation. Evaluation will proceed from simulation to able-bodied volunteers and finally, under the Stage 1 protocol and with caregiver supervision, to participants with disabilities. Beyond the usual navigation figures, we will measure how well the robot holds a comfortable formation, how often it intrudes on personal space, how much it disturbs other humans and how often it talks over someone, alongside the prediction accuracy and calibration of the model itself. Perceived safety, comfort and control will be collected with established questionnaires [2], and the comparison against the baselines will follow published social-navigation evaluation guidelines [18]. A final study will test personalisation across contrasting user profiles.
👤 Profil recherché
Connaissances en apprentissage par renforcement, en apprentissage profond pour la vision par ordinateur et la navigation de robots mobiles ; Python et PyTorch ; une expérience avec ROS et les simulateurs incarnés est un atout.