#reinforcementlearning #artificialintelligence
In module 3 we're going to tackle a fundamental problem in reinforcement learning: the explore-exploit dilemma.
Intelligent agent's must balance their desire for short term reward with the prospect of achieving larger rewards in the long run. We explore a few different strategies for resolving this dilemma: optimistic initial values, epsilon-greedy, and off policy learning.
In module 4, we're going to apply all this to the context of solving problems with dynamic programming, so stay tuned.
Learn how to turn deep reinforcement learning papers into code:
Get instant access to all my courses, including the new Prioritized Experience Replay course, with my subscription service. $29 a month gives you instant access to 42 hours of instructional content plus access to future updates, added monthly.
Discounts available for Udemy students (enrolled longer than 30 days). Just send an email to sales@neuralnet.ai
Or, pickup my Udemy courses here:
Deep Q Learning:
Actor Critic Methods:
Curiosity Driven Deep Reinforcement Learning
Natural Language Processing from First Principles:
Reinforcement Learning Fundamentals
Here are some books / courses I recommend (affiliate links):
Come hang out on Discord here:
Need personalized tutoring? Help on a programming project? Shoot me an email! phil@neuralnet.ai