Heuristics for planning with penalties and rewards formulated in logic and computed through circuits

The automatic derivation of heuristic functions for guiding the search for plans is a fundamental technique in planning. The type of heuristics that have been considered so far, however, deal only with simple planning models where costs are associated with actions but not with states. In this work w...

Full description

Bibliographic Details
Authors: Bonet, Blai, Geffner, Héctor
Format: article
Status:Versión aceptada para publicación
Publication Date:2008
Country:España
Institution:Universitat Pompeu Fabra
Repository:Repositorio Digital de la UPF
OAI Identifier:oai:repositori.upf.edu:10230/36436
Online Access:http://hdl.handle.net/10230/36436
http://dx.doi.org/10.1016/j.artint.2008.03.004
Access Level:Open access
Keyword:Planning
Planning heuristics
Planning with rewards
Knowledge compilation
id ES_23f41f856a4ddbcb376a5c99ed59f0ca
oai_identifier_str oai:repositori.upf.edu:10230/36436
network_acronym_str ES
network_name_str España
repository_id_str
spelling Heuristics for planning with penalties and rewards formulated in logic and computed through circuitsBonet, BlaiGeffner, HéctorPlanningPlanning heuristicsPlanning with rewardsKnowledge compilationThe automatic derivation of heuristic functions for guiding the search for plans is a fundamental technique in planning. The type of heuristics that have been considered so far, however, deal only with simple planning models where costs are associated with actions but not with states. In this work we address this limitation by formulating a more expressive planning model and a corresponding heuristic where preferences in the form of penalties and rewards are associated with fluents as well. The heuristic, that is a generalization of the well-known delete-relaxation heuristic, is admissible, informative, but intractable. Exploiting a correspondence between heuristics and preferred models, and a property of formulas compiled in d-DNNF, we show however that if a suitable relaxation of the domain, expressed as the strong completion of a logic program with no time indices or horizon is compiled into d-DNNF, the heuristic can be computed for any search state in time that is linear in the size of the compiled representation. This representation defines an evaluation network or circuit that maps states into heuristic values in linear-time. While this circuit may have exponential size in the worst case, as for OBDDs, this is not necessarily so. We report empirical results, discuss the application of the framework in settings where there are no goals but just preferences, and illustrate the versatility of the account by developing a new heuristic that overcomes limitations of delete-based relaxations through the use of valid but implicit plan constraints. In particular, for the Traveling Salesman Problem, the new heuristic captures the exact cost while the delete-relaxation heuristic, which is also exponential in the worst case, captures only the Minimum Spanning Tree lower bound.H. Geffner is partially supported by grant TIN2006-15387-C03-03 from MEC/Spain and B. Bonet by a grant from DID/USB/Venezuela. Our planner was built upon a planner by Patrik Haslum. Preliminary experiments were run on the Hermes Computing Resource at the Aragon Institute of Engineering Research (I3A), University of Zaragoza.Elsevier201920192008info:eu-repo/semantics/articleinfo:eu-repo/semantics/acceptedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/36436http://dx.doi.org/10.1016/j.artint.2008.03.004reponame:Repositorio Digital de la UPFinstname:Universitat Pompeu FabraInglésArtificial Intelligence. 2008 Aug;172(12-13):1579-604.info:eu-repo/grantAgreement/ES/2PN/TIN2006-15387-C03-03© Elsevier http://dx.doi.org/10.1016/j.artint.2008.03.004info:eu-repo/semantics/openAccessoai:repositori.upf.edu:10230/364362026-06-12T07:21:37Z
dc.title.none.fl_str_mv Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
title Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
spellingShingle Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
Bonet, Blai
Planning
Planning heuristics
Planning with rewards
Knowledge compilation
title_short Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
title_full Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
title_fullStr Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
title_full_unstemmed Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
title_sort Heuristics for planning with penalties and rewards formulated in logic and computed through circuits
dc.creator.none.fl_str_mv Bonet, Blai
Geffner, Héctor
author Bonet, Blai
author_facet Bonet, Blai
Geffner, Héctor
author_role author
author2 Geffner, Héctor
author2_role author
dc.subject.none.fl_str_mv Planning
Planning heuristics
Planning with rewards
Knowledge compilation
topic Planning
Planning heuristics
Planning with rewards
Knowledge compilation
description The automatic derivation of heuristic functions for guiding the search for plans is a fundamental technique in planning. The type of heuristics that have been considered so far, however, deal only with simple planning models where costs are associated with actions but not with states. In this work we address this limitation by formulating a more expressive planning model and a corresponding heuristic where preferences in the form of penalties and rewards are associated with fluents as well. The heuristic, that is a generalization of the well-known delete-relaxation heuristic, is admissible, informative, but intractable. Exploiting a correspondence between heuristics and preferred models, and a property of formulas compiled in d-DNNF, we show however that if a suitable relaxation of the domain, expressed as the strong completion of a logic program with no time indices or horizon is compiled into d-DNNF, the heuristic can be computed for any search state in time that is linear in the size of the compiled representation. This representation defines an evaluation network or circuit that maps states into heuristic values in linear-time. While this circuit may have exponential size in the worst case, as for OBDDs, this is not necessarily so. We report empirical results, discuss the application of the framework in settings where there are no goals but just preferences, and illustrate the versatility of the account by developing a new heuristic that overcomes limitations of delete-based relaxations through the use of valid but implicit plan constraints. In particular, for the Traveling Salesman Problem, the new heuristic captures the exact cost while the delete-relaxation heuristic, which is also exponential in the worst case, captures only the Minimum Spanning Tree lower bound.
publishDate 2008
dc.date.none.fl_str_mv 2008
2019
2019
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/acceptedVersion
format article
status_str acceptedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10230/36436
http://dx.doi.org/10.1016/j.artint.2008.03.004
url http://hdl.handle.net/10230/36436
http://dx.doi.org/10.1016/j.artint.2008.03.004
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Artificial Intelligence. 2008 Aug;172(12-13):1579-604.
info:eu-repo/grantAgreement/ES/2PN/TIN2006-15387-C03-03
dc.rights.none.fl_str_mv © Elsevier http://dx.doi.org/10.1016/j.artint.2008.03.004
info:eu-repo/semantics/openAccess
rights_invalid_str_mv © Elsevier http://dx.doi.org/10.1016/j.artint.2008.03.004
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Repositorio Digital de la UPF
instname:Universitat Pompeu Fabra
instname_str Universitat Pompeu Fabra
reponame_str Repositorio Digital de la UPF
collection Repositorio Digital de la UPF
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869404671351717888
score 15,812429