Explorer les stratégies d'apprentissage par renforcement : exploration aléatoire dans le problème du bandit manchot Ch. 3 | Knoovi