An Actor-Critic Algorithm for Sequence Prediction

Bahdanau, Dzmitry; Brakel, Philemon; Xu, Kelvin; Goyal, Anirudh; Lowe, Ryan; Pineau, Joelle; Courville, Aaron; Bengio, Yoshua

Full-text links:

Download:

(license)

Current browse context:

cs.LG

< prev | next >

new | recent | 1607

Change to browse by:

References & Citations

NASA ADS

Bookmark

(what is this?)

Computer Science > Learning

Title: An Actor-Critic Algorithm for Sequence Prediction

Authors: Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, Yoshua Bengio

(Submitted on 24 Jul 2016 (v1), last revised 26 Jul 2016 (this version, v2))

Abstract: We present an approach to training neural networks to generate sequences using actor-critic methods from reinforcement learning (RL). Current log-likelihood training methods are limited by the discrepancy between their training and testing modes, as models must generate tokens conditioned on their previous guesses rather than the ground-truth tokens. We address this problem by introducing a \textit{critic} network that is trained to predict the value of an output token, given the policy of an \textit{actor} network. This results in a training procedure that is much closer to the test phase, and allows us to directly optimize for a task-specific score such as BLEU. Crucially, since we leverage these techniques in the supervised learning setting rather than the traditional RL setting, we condition the critic network on the ground-truth output. We show that our method leads to improved performance on both a synthetic task, and for German-English machine translation. Our analysis paves the way for such methods to be applied in natural language generation tasks, such as machine translation, caption generation, and dialogue modelling.

Subjects:	Learning (cs.LG)
Cite as:	arXiv:1607.07086 [cs.LG]
	(or arXiv:1607.07086v2 [cs.LG] for this version)

Submission history

From: Dzmitry Bahdanau [view email]
[v1] Sun, 24 Jul 2016 20:05:07 GMT (43kb,D)
[v2] Tue, 26 Jul 2016 16:08:30 GMT (43kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

arXiv.org > cs > arXiv:1607.07086