Pith. sign in

REVIEW 2 cited by

Session-Level Dynamic Ad Load Optimization using Offline Robust Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.05591 v1 pith:WLRRSSWH submitted 2025-01-09 cs.LG

classification cs.LG
keywords offlineapproachdynamicframeworkgainsbiascausalconfounding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Session-level dynamic ad load optimization aims to personalize the density and types of delivered advertisements in real time during a user's online session by dynamically balancing user experience quality and ad monetization. Traditional causal learning-based approaches struggle with key technical challenges, especially in handling confounding bias and distribution shifts. In this paper, we develop an offline deep Q-network (DQN)-based framework that effectively mitigates confounding bias in dynamic systems and demonstrates more than 80% offline gains compared to the best causal learning-based production baseline. Moreover, to improve the framework's robustness against unanticipated distribution shifts, we further enhance our framework with a novel offline robust dueling DQN approach. This approach achieves more stable rewards on multiple OpenAI-Gym datasets as perturbations increase, and provides an additional 5% offline gains on real-world ad delivery data. Deployed across multiple production systems, our approach has achieved outsized topline gains. Post-launch online A/B tests have shown double-digit improvements in the engagement-ad score trade-off efficiency, significantly enhancing our platform's capability to serve both consumers and advertisers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generative Optimization for Incentivized Advertising with Global Level Constraints

    cs.LG 2026-08 conditional novelty 5.0 of 10

    Tokenized autoregressive generation of incentive amounts with a lambda-conditioned policy and constraint-aware alignment (SCPO) reports higher revenue, higher ROI, and fewer ROI violations than six baselines on one in...

  2. Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems

    cs.IR 2025-06 conditional novelty 4.0 of 10

    MTORL jointly learns channel recommendation and budget allocation for online advertising from offline user journeys, and reports better accuracy and reward than prior methods on KuaiRand, Criteo, and a Taobao A/B test.

Pith tools