Pith. sign in

REVIEW 2 cited by

A Multi-Agent Reinforcement Learning Method for Impression Allocation in Online Display Advertising

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.03152 v1 pith:2JGZ4TZV submitted 2018-09-10 cs.AI cs.GTcs.LG

classification cs.AIcs.GTcs.LG
keywords contractsimpressionspublisheradvertisingallocationapproachdisplayenvironment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In online display advertising, guaranteed contracts and real-time bidding (RTB) are two major ways to sell impressions for a publisher. Despite the increasing popularity of RTB, there is still half of online display advertising revenue generated from guaranteed contracts. Therefore, simultaneously selling impressions through both guaranteed contracts and RTB is a straightforward choice for a publisher to maximize its yield. However, deriving the optimal strategy to allocate impressions is not a trivial task, especially when the environment is unstable in real-world applications. In this paper, we formulate the impression allocation problem as an auction problem where each contract can submit virtual bids for individual impressions. With this formulation, we derive the optimal impression allocation strategy by solving the optimal bidding functions for contracts. Since the bids from contracts are decided by the publisher, we propose a multi-agent reinforcement learning (MARL) approach to derive cooperative policies for the publisher to maximize its yield in an unstable environment. The proposed approach also resolves the common challenges in MARL such as input dimension explosion, reward credit assignment, and non-stationary environment. Experimental evaluations on large-scale real datasets demonstrate the effectiveness of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Causal Inference for Advertising in Real-Time Bidding Auctions

    cs.LG 2019-08 conditional novelty 7.0 of 10

    The optimal bid in a second-price ad auction equals the conditional average treatment effect of the ad, and the paper's Thompson-sampling algorithm learns that bid and therefore the ad effect while reducing experiment...

  2. Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Sending fewer ad requests to demand-side platforms, chosen by predicted bid value, lifted net revenue by 4.6% while cutting request volume by 34.2% in a production ad exchange.

Pith tools