The Baseline That Bites Back
A standard baseline choice in policy-gradient methods is the state-value function V^π(s). When you subtract V^π(s) from the return G_t, the resulting term G_t − V^π(s) is, in expectation, an unbiased estimate of which quantity?
Sign in to answer questions and track your progress
Sign In