Bayesian reinforcement learning with Dirichlet processes
Apprentissage par renforcement Bayésien avec processus de Dirichlet
- Apprentissage en ligne
- Processus de Dirichlet
- Apprentissage par renforcement (intelligence artificielle)
- Statistique bayésienne non paramétrique
- Problème du bandit manchot
- Algorithmes en ligne
- Reinforcement Learning
- Online learning
- Bayesian nonparametric statistics
- Langue : Anglais
- Discipline : Informatique et applications
- Identifiant : 2026ULILB002
- Type de thèse : Doctorat
- Date de soutenance : 30/01/2026
Résumé en langue originale
L'apprentissage par renforcement (RL) a montré un grand potentiel dans des applications allant des jeux de société à la biologie et au-delà. Cependant, il reste encore à voir le jour dans d'autres domaines difficiles tels que l'agro-écologie, l'inefficacité des algorithmes RL dans de tels domaines essentiellement dus à des distributions de données concrètes compliquées qui ne sont pas conformes aux modèles statistiques standard tels que la famille exponentielle monoparamétrique. Motivée par de tels problèmes du monde réel, cette thèse introduit un nouveau cadre pour RL basé sur les statistiques non paramétriques bayésiennes, en particulier les processus de Dirichlet (DP). La première contribution de la thèse est le Dirichlet Process Posterior Sampling (DPPS), un algorithme non paramétrique bayésien probablement optimal pour les bandits multi-bras basé sur les DP. Essentiellement, DPPS est un algorithme de correspondance de probabilité, et combine la force du bootstrap (bayésien) avec un mécanisme de principe d'incorporation et d'exploitation d'informations antérieures. La thèse propose alors deux nouveaux algorithmes pour l'apprentissage par renforcement dans les processus de décision de Markov. La première est l'apprentissage profond basé sur l'échantillonnage par Thompson, dans lequel, l'exploration est basée sur l'échantillonnage postérieur de fonctions de valeur d'action (représentées par un réseau de neurones profonds). En particulier, l'étape d'échantillonnage postérieure de cet algorithme utilise un nouveau schéma basé sur DP, également introduit dans cette thèse, et qui peut être d'intérêt indépendant pour d'autres applications des réseaux de neurones bayésiens. Le deuxième algorithme est un algorithme RL de distribution qui repose sur une égalité de distribution associée aux DP.
Résumé traduit
Reinforcement learning (RL) has shown great potential in applications ranging from board games to biology and beyond. However, it still remains to see proper light of the day in other challenging domains such as agro-ecology, the inefficiency of RL algorithms in such domains essentially owed to complicated real world data distributions that do not conform to standard statistical models such as single-parametric exponential family. Motivated by such real world problems, this thesis introduces a new framework for RL based on Bayesian nonparametric statistics, specifically Dirichlet Processes (DPs). The first contribution of the thesis is Dirichlet Process Posterior Sampling (DPPS), a provably optimal Bayesian non-parametric algorithm for multi-arm bandits based on DP priors. Essentially, DPPS is a probability-matching algorithm, and combines the strength of (Bayesian) bootstrap with a principled mechanism of incorporating and exploiting prior information. The thesis then proposes two new algorithms for reinforcement learning in Markov decision processes. The first one is Thompson sampling based deep Q learning, wherein, exploration is based on posterior sampling of action-value functions (represented through a deep neural network). In particular, the posterior sampling step in this algorithm utilizes a new DP based scheme, also introduced in this thesis, and which can be of independent interest in other applications of Bayesian neural networks. The second algorithm is a distributional RL algorithm that hinges on a distributional equality associated with the DPs.
- Directeur(s) de thèse : Maillard, Odalric-Ambrym
- Président de jury : Castillo, Ismaël
- Membre(s) de jury : Durand, Audrey
- Rapporteur(s) : Mannor, Shie - Alquier, Pierre
- Laboratoire : Centre Inria de l'Université de Lille - Centre de Recherche en Informatique, Signal et Automatique de Lille
- École doctorale : École graduée Mathématiques, sciences du numérique et de leurs interactions (Lille ; 2021-....)
AUTEUR
- Vashishtha, Sumit



