A Bayesian exploration&#x2212;exploitation approach for optimal online sensing and planning with a visually guided mobile robot

de Freitas, Nando

doi:doi:10.1007/s10514-009-9130-2

A Bayesian exploration−exploitation approach for optimal online sensing and planning with a visually guided mobile robot

Ruben Martinez−Cantin‚ Nando Freitas‚ Eric Brochu‚ José Castellanos and Arnaud Doucet

Abstract

We address the problem of online path planning for optimal sensing with a mobile robot. The objective of the robot is to learn the most about its pose and the environment given time constraints. We use a POMDP with a utility function that depends on the belief state to model the finite horizon planning problem. We replan as the robot progresses throughout the environment. The POMDP is high-dimensional, continuous, non-differentiable, nonlinear, non-Gaussian and must be solved in real-time. Most existing techniques for stochastic planning and reinforcement learning are therefore inapplicable. To solve this extremely complex problem, we propose a Bayesian optimization method that dynamically trades off exploration (minimizing uncertainty in unknown parts of the policy space) and exploitation (capitalizing on the current best solution). We demonstrate our approach with a visually-guide mobile robot. The solution proposed here is also applicable to other closely-related domains, including active vision, sequential experimental design, dynamic sensing and calibration with mobile sensors.

ISSN

0929−5593

Journal

Autonomous Robots

Number

Pages

93–103

Publisher

Springer US

Volume

Year

2009

A Bayesian exploration−exploitation approach for optimal online sensing and planning with a visually guided mobile robot

Abstract

Links

See Also