Pith. sign in

REVIEW 1 cited by

From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01458 v1 pith:UFCYHPY7 submitted 2024-10-02 cs.AI cs.LG

classification cs.AIcs.LG
keywords q-shapingshapingrewardachievingacrossalternativeefficiencyimprovement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Q-shaping is an extension of Q-value initialization and serves as an alternative to reward shaping for incorporating domain knowledge to accelerate agent training, thereby improving sample efficiency by directly shaping Q-values. This approach is both general and robust across diverse tasks, allowing for immediate impact assessment while guaranteeing optimality. We evaluated Q-shaping across 20 different environments using a large language model (LLM) as the heuristic provider. The results demonstrate that Q-shaping significantly enhances sample efficiency, achieving a \textbf{16.87\%} improvement over the best baseline in each environment and a \textbf{253.80\%} improvement compared to LLM-based reward shaping methods. These findings establish Q-shaping as a superior and unbiased alternative to conventional reward shaping in reinforcement learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    SILIC uses LLM-guided inverse reinforcement learning and Theory of Planned Behavior chain reasoning to infer age, gender, income, and employment from travel trajectories, reportedly beating SVM, XGBoost, CatBoost, and...

Pith tools