White RoomNEW

Two Ways to Score It

Under policy π\pi, the Bellman expectation equation for the state-value function is Vπ(s)=aπ(as)sP(ss,a)[R(s,a,s)+γVπ(s)]V_\pi(s) = \sum_a \pi(a|s) \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_\pi(s')]. Given this, how does the action-value function Qπ(s,a)Q_\pi(s,a) relate to Vπ(s)V_\pi(s)?