Policy search in kernel Hilbert space

Bagnell, J. Andrew.; Schneider, Jeff

doi:10.1184/R1/6558323.v1

file.pdf (1.32 MB)

Policy search in kernel Hilbert space

journal contribution

posted on 2003-01-01, 00:00 authored by J. Andrew. Bagnell, Jeff Schneider

Abstract: "Much recent work in reinforcement learning and stochastic optimal control has focused on algorithms that search directly through a space of policies rather than building approximate value functions. Policy search has numerous advantages: it does not rely on the Markov assumption, domain knowledge may be encoded in a policy, the policy may require less representational power than a value-function approximation, and stable and convergent algorithms are well-understood. In contrast with value-function methods, however, existing approaches to policy search have heretofore focused entirely on parametric approaches. This places fundamental limits on the kind of policies that can be represented. In this work, we show how policy search (with or without the additional guidance of value-functions) in a Reproducing Kernel Hilbert Space gives a simple and rigorous extension of the technique to non-parametric settings. In particular, we investigate a new class of algorithms which generalize Reinforce-style likelihood ratio methods to yield both online and batch techniques that perform gradient search in a function space of policies. Further, we describe the computational tools that allow efficient implementation. Finally, we apply our new techniques towards interesting reinforcement learning problems."

History

Publisher Statement

Date

2003-01-01

Usage metrics

Keywords

Hilbert space.Stochastic partial differential equations.Kernel functions.

Licence

In Copyright

Exports

RefWorks

BibTeX

Ref. manager

Endnote

DataCite

NLM

DC

Policy search in kernel Hilbert space

History

Publisher Statement

Date

Usage metrics

Categories

Keywords

Licence

Exports