PrefBench benchmark shows zero-shot LLMs achieve deal rates above 0.99 but seller profits only slightly above random and far below a simple concession heuristic across 7,500 episodes.
Dynamic Pricing in High-Speed Railways Using Multi- Agent Reinforcement Learning
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
VarLenRec learns variable-length semantic IDs for generative recommendation by allocating longer codes to tail items via popularity-weighted information budget allocation, hyperbolic residual quantization, and a differentiable soft length controller.
An entity-graph MARL framework (RACHE) using R-GCN message passing and attention pooling over train-service nodes outperforms baseline algorithms in railway pricing revenue across two simulated market scenarios.
AIGP combines LLMs with offline RL and DPO to produce interpretable pricing policies that improved GMV by 13.21%, ROI by 7.59%, and milestone achievement by 8.20% in 14-day online tests versus baseline.
citing papers explorer
-
PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations
PrefBench benchmark shows zero-shot LLMs achieve deal rates above 0.99 but seller profits only slightly above random and far below a simple concession heuristic across 7,500 episodes.
-
Learning Variable-Length Tokenization for Generative Recommendation
VarLenRec learns variable-length semantic IDs for generative recommendation by allocating longer codes to tail items via popularity-weighted information budget allocation, hyperbolic residual quantization, and a differentiable soft length controller.
-
Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets
An entity-graph MARL framework (RACHE) using R-GCN message passing and attention pooling over train-service nodes outperforms baseline algorithms in railway pricing revenue across two simulated market scenarios.
-
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
AIGP combines LLMs with offline RL and DPO to produce interpretable pricing policies that improved GMV by 13.21%, ROI by 7.59%, and milestone achievement by 8.20% in 14-day online tests versus baseline.