TD Temporal-Difference Learning 时序差分法(差分学习)

temporary
英 ['temp(ə)rərɪ]美 [ˈtempəreri]
adj. 临时的,暂时的;短暂的

n. 临时工,临时雇
TD算法是RL的核心算法。TD是DP和MC算法的结合。Like DP, TD methods without waiting for a final outcome (they bootstrap)。

TD(0), or one-step TD

TD Temporal-Difference Learning 时序差分法(差分学习)_第1张图片
TD Temporal-Difference Learning 时序差分法(差分学习)_第2张图片

Advantages of TD Prediction Methods

TD methods update their estimates based in part on other estimates. They learn a guess from a guess,they bootstrap.
TD Temporal-Difference Learning 时序差分法(差分学习)_第3张图片

Q-learning: Off-policy TD Control

在这里插入图片描述
TD Temporal-Difference Learning 时序差分法(差分学习)_第4张图片

你可能感兴趣的:(深度强化学习)