Pith. sign in

Rewards are dominated by Rparsing and Rexecution (Equation (8) and Table 8), as successful negotiations are rare

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.GT 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

Training Language Models for Bilateral Trade with Private Information

cs.GT · 2026-04-10 · unverdicted · novelty 7.0

Frontier LLMs achieve higher surplus via sequential price discrimination in bilateral trade simulations, while SFT followed by GRPO on Qwen models trades off surplus gains against deal rates and improves consistency across price tiers.

citing papers explorer

Showing 1 of 1 citing paper.

  • Training Language Models for Bilateral Trade with Private Information cs.GT · 2026-04-10 · unverdicted · none · ref 23

    Frontier LLMs achieve higher surplus via sequential price discrimination in bilateral trade simulations, while SFT followed by GRPO on Qwen models trades off surplus gains against deal rates and improves consistency across price tiers.