Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Contract-Inspired Contest Theory for Controllable Image Generation in Mobile Edge Metaverse

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A fixed-pool contest plus diffusion-guided DRL raises generated image quality and convergence speed in mobile edge Metaverse, reaching average reward 303.67.

desk verdict The compression-sensitivity experiments are worth a look, but the contest-theoretic model collapses because capability is defined as a function of the very effort it is supposed to govern, and the claimed optimality is unsupported. read the letter →

arxiv 2501.09391 v1 pith:GES6BFUF submitted 2025-01-16 cs.NI

classification cs.NI
keywords generativediffusionmodelsdeepreinforcementlearningcontesttheorycontractmobileedgemetaversesemanticcommunicationimagegenerationresourceallocation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that controllable image generation in a mobile edge Metaverse can be made efficient by coupling a contract-inspired contest, where a fixed payment pool is split among edge tasks according to their chosen transmit power, with a diffusion-model-augmented deep reinforcement learner that sets the payment and award parameters. The point of the mechanism is to make edge devices reveal and act on their semantic-transfer capability, so that scarce transmit power goes to semantics whose quality degrades fastest under compression. The authors report that the resulting GDM-DRL solver converges in about 25 steps to an average reward of 303.67, beating average-power allocation by 46.95%, PPO by 26.29%, SAC by 28.39%, and Transformer-based SAC by 14.93%, with visibly smaller fluctuations. If these numbers hold, the framework offers a practical template for allocating resources in dynamic edge networks where quality of generated content is the payoff.

What carries the argument

The load-bearing object is the contest itself: the generation server pays the edge server a fixed pool based on delivered image quality, and the edge server splits that pool among semantic transfer tasks according to their relative capability. Capability is defined as the inverse of the quality produced under the compression forced by the chosen transmit power, cost is $C(a_i,P_i)=P_i/a_i$, and the expected award for a task is the fixed-pool split in Eq. (14), computed from a uniform population distribution of capabilities. On top of this, a diffusion model acts as the policy: its reverse chain turns Gaussian noise into the action vector of payment and award parameters, and a DRL critic scores those actions by total generated image quality, with target networks updated by soft replacement. The result is a closed loop in which the contest sets the incentive structure and the diffusion-driven RL adjusts that structure online.

What would settle it

Recompute the contest best response with capability endogenous, using $a_i = 1/Q_i(Z(D(P_i)))$ inside Eq. (14), and compare the predicted transmit powers with the four reward settings in Fig. 7; if the predicted pattern differs, the contest equilibrium that the incentive mechanism relies on is not the one claimed.

Watch

Extended reading notes

Core claim

The central claim is that a contract-inspired contest can coordinate transmit-power choices for semantic data transfer so as to maximize the quality of diffusion-generated images, and that solving the resulting multi-tier optimization with GDM-DRL is both faster and more stable than PPO, SAC, or Transformer-based SAC. The paper defines each transfer task's capability as the inverse of the image quality it would yield at the compression level forced by its chosen power, then derives an expected-award formula in which the fixed payment pool is split according to relative capability. Running the optimization on a four-task testbed with ControlNet semantics, namely depth, segmentation, pose, and Canny edges, produces an incentive structure where winner-takes-all rewards drive all tasks to maximum power, evenly split rewards drive them to minimum effort, and intermediate decreasing rewards balance effort and quality. The headline quantitative result is the converged average reward of 303.67 for the GDM-DRL solver, which the paper attributes to diffusion-based iterative refinement of the payment and award parameters.

Load-bearing premise

The contest derivation treats each task's capability as a fixed type when computing expected awards, but the paper defines capability as a function of the transmit power the task chooses, so the equilibrium formula may not describe actual choices.

Editorial extensions

If this is right

  • If the framework is correct, an edge image-generation service can use a fixed payment pool to elicit near-optimal transmit power without knowing each device's private capability, because the contest self-selects which tasks transmit at high power.
  • The reported converged reward of 303.67, which is 14.93% above Transformer-based SAC and 26.29% above PPO, implies that diffusion-based action generation can stabilize training in dynamic wireless environments where the reward surface is noisy.
  • The reward-setting experiments imply that an operator can shape power usage by choosing the award vector: winner-takes-all maximizes power, even splits minimize it, and decreasing splits offer a balanced middle path.
  • The depth-map anomaly implies that slight semantic compression can improve generated image quality by removing distracting detail, so the optimal compression level is not always the lowest one.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the capability shortcut in Eq. (10) is taken literally, the expected-award formula in Eq. (14) no longer treats capability as a fixed type, so the reported power allocations in Fig. 7 may not be true contest equilibria; a re-derivation with endogenous capability would be needed before the incentive mechanism's optimality can be taken as established.
  • The same contract-contest construction transfers naturally to other semantic modalities, such as audio, video, or 3D scene descriptors, whenever a central generator pays for transmitted fidelity and each source's quality degrades with compression.
  • A direct testable extension would be to replace the uniform capability distribution in Eq. (13) with measured device statistics and check whether the expected-award split still predicts observed transmit powers, since the paper only demonstrates the uniform case.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a framework for resource allocation and incentive design in mobile edge Metaverse image generation, combining contract-inspired contest theory with a GDM-DRL algorithm. Semantic transfer tasks compete for transmit power and rewards, while a generation server sets a payment plan based on generated image quality. The authors formulate a contest-theoretic optimization problem, propose a diffusion-model-based DRL solver, and evaluate the approach against PPO, SAC, Transformer-based SAC, and average power allocation. Experiments also study how compression level affects image quality for different semantic types. The central claim is that the proposed contract-inspired contest framework with GDM-DRL achieves higher generated image quality, faster convergence, and better stability than traditional DRL methods.

Significance. The problem addressed, namely resource-efficient controllable image generation in mobile edge Metaverse environments, is timely and relevant. The empirical study of the relationship between semantic compression level and generated image quality across depth, segmentation, pose, and Canny edge semantics is a useful contribution, and the use of real ControlNet and image-quality metrics gives the system model practical grounding. If the contest-theoretic derivation were sound, the combination of incentive theory with diffusion-model-based DRL would be a novel design point for AIGC resource allocation. However, the formal contest model contains a load-bearing flaw: capability is defined as a function of the effort variable itself, which invalidates the imported equilibrium formulas and the claimed optimal award scheme. The experimental evaluation also does not isolate the effect of the contest mechanism, since the DRL reward is identical to the optimization objective and no non-DRL or no-contest baseline is provided. As presented, therefore, the central theoretical contribution is not supported.

major comments (4)
  1. [Section IV-B, Eq. (10)] The definition of capability in Eq. (10) as a_i = 1/Q_i(Z(D(P_i))) makes capability a function of the transmit power P_i, which is exactly the effort variable in the contest. The expected award formula M(a_i, r_i) in Eq. (14) is derived under the standard assumption that a_i is an exogenous type drawn from a known distribution P(a_i), independent of effort. The best-response problem in Eq. (15), however, optimizes over P_i with a_i(P_i) substituted into M, so the contestant's own choice of P_i changes the type parameter used to compute the award probability and cost. This breaks the game-theoretic basis of the derivation: Eq. (14) no longer describes a valid best response when types are effort-dependent, and the optimality claim for the award scheme in Section IV-C and the incentive-compatibility constraint (18g) are unsupported. The paper provides no fixed-point or alternative equilibrium argument that repairs this circularity. A concrete fix would be to define a_i as an exogenous parameter independent of P_i (e.g., depending only on semantic type and channel statistics), or to reformulate the problem as a moral-hazard contract model in which effort directly affects quality.
  2. [Section IV-B, Eq. (14)] Equation (14) has the binomial exponents reversed. If P(ai) is the probability that another contestant's capability is larger than ai, then the probability that exactly i-1 of the N_u-1 other contestants have larger capability is C(N_u-1, i-1) P(ai)^{i-1}(1-P(ai))^{N_u-i}. The formula as written, P(ai)^{N_u-i}(1-P(ai))^{i-1}, instead gives the probability that i-1 contestants have smaller capability, i.e., the ranking is reversed. This error propagates into M(a_i, r_i) and therefore into the optimal effort calculation in Eq. (15) and the award scheme in Section IV-C. The summation index i is also reused for both the contestant and the award rank; the expression should sum over ranks j with the corresponding award r_j.
  3. [Section V-B and Section VI-D] The DRL reward in Eq. (32) is defined as R_g = sum_{i=1}^{N_u} Q_i (with a likely typo Q_Nu), which is the same quantity maximized in the overall objective (18a). Consequently, the GDM-DRL agent is simply learning to maximize the image-quality objective directly, without enforcing the contest constraints (18b)-(18g) or using the contest equilibrium in Sections IV-B and IV-C. The experiments in Section VI-D compare GDM-DRL against PPO, SAC, and Transformer-based SAC, which are all DRL algorithms; they do not compare against a direct solution of (16)-(18) by a non-learning optimizer, nor against a variant without the contest mechanism, nor do they report whether the learned parameters satisfy the outage, power-budget, or incentive-compatibility constraints. The central claim that the contract-inspired contest framework improves performance is therefore not tested by these experiments.
  4. [Section IV-B, cost function and capability interpretation] There is a conceptual inconsistency in the interpretation of capability. The text states that higher capability corresponds to 'higher loss of image quality against same level semantic compression,' but the formula a_i = 1/Q_i means high capability corresponds to low Q_i. With the cost function C(a_i, P_i) = P_i / a_i, a high a_i yields a low cost for a given power, meaning tasks with low quality (more sensitive to compression) face lower effort cost. This is the opposite of the usual contest-theoretic interpretation that higher capability means higher ability to exert effort at lower cost, and it also conflicts with the prose's characterization of capability as 'robustness against compression loss.' The direction of the cost function affects the equilibrium effort choices and the optimal award scheme, so this ambiguity should be resolved.
minor comments (6)
  1. [Section V-B, Eq. (32)] Equation (32) uses Q_Nu rather than Q_i; the summation index in the reward definition should be over all tasks.
  2. [Section VI-D, text and Fig. 10(b)] The reported average reward for GDM after convergence differs between the text (303.67) and the value shown in Fig. 10(b) (approximately 323.67), and the claimed improvement of 46.95% over average power allocation is inconsistent with the reported numbers; these discrepancies should be reconciled.
  3. [Section VI-C-1, last-place reward] The sentence 'all tasks reduce their power to the maximum power (constrained by the outage probability)' is self-contradictory; it likely means they reduce to a minimal power consistent with the outage constraint.
  4. [Section V-B, Eq. (33)] Equation (33) uses κ as the learning rate while the surrounding text refers to α; one of these notations should be corrected.
  5. [Section IV-B, Eq. (15)] The notation in Eq. (15) is unclear: 'min(P)' is not defined, and the equality to a minimum of an argmax set is confusing. A conventional best-response equation P_i^* = arg max_{P_i} M(a_i(P_i), r_i) would be clearer.
  6. [Section II-B, reference [21]] Reference [21] is incomplete: the author list gives 'C. L. C.' without a full name, and the page range 'p. 221-221' is likely erroneous.

Circularity Check

2 steps flagged · score 7.0 of 10

The contest-theoretic 'optimal effort' is self-referential: Eq. (10) defines capability a_i as a function of transmit power P_i via image quality Q_i, and Eq. (15) optimizes over P_i inside the type argument, so the optimal award scheme is built from the very objective it claims to derive.

  1. self definitional [Section IV-B, Eq. (10); used in Eq. (15)]
    "The capability of each semantic transfer task is defined by the robustness of its semantics against compression loss. Referring to (1), we model the capability ai as [9]: ai = 1/Qi(Z(D(Pi))). This encourages transfer tasks with semantic that is more sensitive to compression to compete for higher transmit power to maintain quality."

    In contest theory, ai must be exogenous when Pi is chosen: Eq. (14) computes M(ai,ri) from the distribution of ai, and the cost in Eq. (16a) is C(ai,Pi)=Pi/ai with ai fixed. Here Eq. (10) defines ai = 1/Qi(Z(D(Pi))), so ai is a function of the effort variable Pi. Eq. (15) then sets Pi^* = argmax_{Pi} M(ai(Pi), ri), i.e., the 'optimal effort' maximizes an award function whose type argument is a function of Pi. Qi is the quality objective maximized in Eq. (18a) and is the DRL reward Rg = sum_i Qi in Eq. (32), so the contest 'equilibrium' is a self-referential optimization over the same quantity it claims to determine; no fixed-point argument is supplied.

  2. self citation load bearing [Section IV-B, Eq. (10), citation [9]]
    "Referring to (1), we model the capability ai as [9]: ai = 1/Qi(Z(D(Pi)))."

    The definition is not derived in this paper; the only justification is a citation to [9] (J. Wang, H. Du, D. Niyato, J. Kang, X. Shen), whose authors overlap with the present authors. This definition is load-bearing because it is what makes ai a function of Pi and hence breaks the standard expected-award derivation (14)-(15). The citation does not provide independent support—it is not machine-checked or parameter-free with assumptions excluding the target result; it reuses the same modeling convention. Thus the central contest-theoretic optimality claim rests on a self-citation that smuggles in the problematic endogenous-capability assumption.

full rationale

The central formal claim of the paper—that the contract-inspired contest mechanism yields an optimal transmit-power allocation and optimal award scheme—does not follow from contest theory. The load-bearing Eq. (10) defines the contestant 'capability' a_i as 1/Q_i(Z(D(P_i))), where Q_i is precisely the image-quality objective that the overall problem (18a) and the DRL reward (32) maximize, and where P_i is the effort variable that the contest is supposed to determine. Eq. (15) then optimizes M(a_i(P_i), r_i) over P_i, so the 'optimal effort' is a self-referential function of the target quality; no fixed-point or equilibrium existence argument is supplied. The subsequent optimal-award scheme (16) and IC constraint (18g) inherit this collapse. The definition is imported from the authors' own prior work [9], so the central premise is also a load-bearing self-citation. The empirical benchmarks (image-quality vs compression, DRL convergence curves) are self-contained simulations and are not themselves circular, which is why the score is not higher; nevertheless, the paper's advertised contribution—contest-theoretically optimal incentives—reduces, by construction, to optimizing over a quantity derived from the same Q_i it claims to improve.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The framework does not introduce new physical entities; its constructed 'capability' is a modeled quantity, not an entity. The central assumptions are the contest model, the linear cost function, and the uniform type distribution.

free parameters (4)
  • beta = 0.5
    Hand-chosen weight in Eq. (5) balancing ImageReward and SSIM; affects all computed image quality scores.
  • x = 30 fps
    Frame rate requirement in Eq. (4), chosen as a common real-time threshold; directly sets the compression level formula.
  • theta = 0.05
    Outage probability threshold in Table II used in constraint (16e) and (18b).
  • P_total = 100 mW
    Total power budget in Table II, bounds the feasible transmit power allocations.
assumptions (4)
  • standard math Expected award in a multi-prize contest follows the binomial formula in Eq. (14) from contest theory [21].
    The paper imports the standard contest theory result, but the formula as written has swapped exponents.
  • ad hoc to paper The cost function C(a_i, P_i) = P_i / a_i satisfies the conditions in (8).
    This holds only when a_i is constant; with Eq. (10) making a_i a function of P_i, the monotonicity conditions are not guaranteed.
  • ad hoc to paper The capability distribution P(ai) is uniform on [0, aimax].
    The paper states this for tractability and says it can be extended, but the equilibrium formula uses it directly.
  • domain assumption Rayleigh fading with unit-mean exponential channel gain and outage probability formula in Eq. (2).
    Standard wireless communication model imported from [31], [32].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contract-Inspired Contest Theory for Controllable Image Generation in Mobile Edge Metaverse." pith.science (2026). https://pith.science/paper/GES6BFUF

@misc{pith2026250109391,
  author       = {Pith},
  title        = {Pith review of: Contract-Inspired Contest Theory for Controllable Image Generation in Mobile Edge Metaverse},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GES6BFUF}},
  note         = {Machine review of arXiv:2501.09391}
}
read the original abstract

The rapid advancement of immersive technologies has propelled the development of the Metaverse, where the convergence of virtual and physical realities necessitates the generation of high-quality, photorealistic images to enhance user experience. However, generating these images, especially through Generative Diffusion Models (GDMs), in mobile edge computing environments presents significant challenges due to the limited computing resources of edge devices and the dynamic nature of wireless networks. This paper proposes a novel framework that integrates contract-inspired contest theory, Deep Reinforcement Learning (DRL), and GDMs to optimize image generation in these resource-constrained environments. The framework addresses the critical challenges of resource allocation and semantic data transmission quality by incentivizing edge devices to efficiently transmit high-quality semantic data, which is essential for creating realistic and immersive images. The use of contest and contract theory ensures that edge devices are motivated to allocate resources effectively, while DRL dynamically adjusts to network conditions, optimizing the overall image generation process. Experimental results demonstrate that the proposed approach not only improves the quality of generated images but also achieves superior convergence speed and stability compared to traditional methods. This makes the framework particularly effective for optimizing complex resource allocation tasks in mobile edge Metaverse applications, offering enhanced performance and efficiency in creating immersive virtual environments.

Figures

Figures reproduced from arXiv: 2501.09391 by the authors.

Figure 1
Figure 1. System components and data flow for the mobile edge immersive Metaverse image generation. The data flow in the system model is in the following [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the relationship between compression level and [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. System model for the proposed mobile edge immersive Meta [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Illustration of the forward and reverse diffusion processes: The forward [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Image quality as a function of different levels of semantic compression for various types of semantic inputs (depth map, segmentation, pose estimation, [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Detailed analysis of image quality for each prompt using depth map semantics. The anomaly in the depth map results, where a slight compression [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: This figure illustrates how different reward distributions affect [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 9
Figure 9. Figure 9: This figure presents the effect of different total reward levels (ranging [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: (a) illustrate performance in terms of reward over steps. GDM demonstrates superior convergence speed and stability, achieving higher rewards consistently across steps, especially compared to PPO and SAC, which exhibit more fluctuations due to the increased complexity…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    Guiding a pretrained topology-diffusion generator with human-preference reward classifiers is claimed to suppress floating-material and boundary-violation failure modes without retraining the generator.

Reference graph

Works this paper leans on

46 extracted references · 38 canonical work pages · cited by 1 Pith paper

  1. [1]

    A Full Dive into Realizing the Edge- enabled Metaverse: Visions, Enabling Technologies, and Challenges,

    M. Xu, W. C. Ng, W. Y . B. Lim, J. Kang, Z. Xiong, D. Niyato, Q. Yang, X. Shen, and C. Miao, “A Full Dive into Realizing the Edge- enabled Metaverse: Visions, Enabling Technologies, and Challenges,” IEEE Commun. Surveys Tuts. , vol. 25, no. 1, pp. 656–700, 2022

  2. [2]

    Dual- Level Resource Provisioning and Heterogeneous Auction for Mobile Metaverse,

    X. Ren, H. Du, C. Qiu, T. Luo, Z. Liu, X. Wang, and D. Niyato, “Dual- Level Resource Provisioning and Heterogeneous Auction for Mobile Metaverse,” IEEE Trans. Mobile Comput. , pp. 1–15, 2024

  3. [3]

    Spotlighter: Backup Age-Guaranteed Immersive Virtual Vehicle Ser- vice Provisioning in Edge-Enabled Vehicular Metaverse,

    Y . Qiu, M. Chen, H. Huang, W. Liang, J. Liang, Y . Hao, and D. Niyato, “Spotlighter: Backup Age-Guaranteed Immersive Virtual Vehicle Ser- vice Provisioning in Edge-Enabled Vehicular Metaverse,” IEEE Trans. Mobile Comput. , pp. 1–17, 2024

  4. [4]

    MetaSlicing: A Novel Resource Allocation Framework for Metaverse,

    N. H. Chu, D. T. Hoang, D. N. Nguyen, K. T. Phan, E. Dutkiewicz, D. Niyato, and T. Shu, “MetaSlicing: A Novel Resource Allocation Framework for Metaverse,” IEEE Trans. Mobile Comput. , vol. 23, no. 5, pp. 4145–4162, 2024

  5. [5]

    Attention-Aware Resource Allocation and QoE Analysis for Metaverse xURLLC Services,

    H. Du, J. Liu, D. Niyato, J. Kang, Z. Xiong, J. Zhang, and D. I. Kim, “Attention-Aware Resource Allocation and QoE Analysis for Metaverse xURLLC Services,” IEEE J. Select. Areas Commun. , 2023

  6. [6]

    A Unified Framework for Guiding Generative AI with Wireless Perception in Resource Constrained Mobile Edge Networks,

    J. Wang, H. Du, D. Niyato, J. Kang, Z. Xiong, D. Rajan, S. Mao, and X. Shen, “A Unified Framework for Guiding Generative AI with Wireless Perception in Resource Constrained Mobile Edge Networks,” IEEE Trans. Mobile Comput. , 2024

  7. [7]

    Adding Conditional Control to Text-to-Image Diffusion Models,

    L. Zhang, A. Rao, and M. Agrawala, “Adding Conditional Control to Text-to-Image Diffusion Models,” in Proc. IEEE Int. Conf. Comput. Vis., 2023, pp. 3836–3847

  8. [8]

    Semantic Communications for Artificial Intelligence Generated Content (AIGC) Toward Effective Content Creation,

    G. Liu, H. Du, D. Niyato, J. Kang, Z. Xiong, D. I. Kim, and X. Shen, “Semantic Communications for Artificial Intelligence Generated Content (AIGC) Toward Effective Content Creation,” IEEE Netw., pp. 1–1, 2024

Show all 46 references
  1. [9]

    Semantic- Aware Sensing Information Transmission for Metaverse: A Contest Theoretic Approach,

    J. Wang, H. Du, Z. Tian, D. Niyato, J. Kang, and X. Shen, “Semantic- Aware Sensing Information Transmission for Metaverse: A Contest Theoretic Approach,” IEEE Trans. Wireless Commun. , vol. 22, no. 8, pp. 5214–5228, 2023

  2. [10]

    Energy- Efficient Resource Allocation in Generative AI-Aided Secure Semantic Mobile Networks,

    J. Zheng, B. Du, H. Du, J. Kang, D. Niyato, and H. Zhang, “Energy- Efficient Resource Allocation in Generative AI-Aided Secure Semantic Mobile Networks,” IEEE Trans. Mobile Comput. , 2024

  3. [11]

    A Unified Framework for Integrating Semantic Communication and AI-Generated Content in Metaverse,

    Y . Lin, Z. Gao, H. Du, D. Niyato, J. Kang, A. Jamalipour, and X. S. Shen, “A Unified Framework for Integrating Semantic Communication and AI-Generated Content in Metaverse,” IEEE Netw., 2023. 16

  4. [12]

    Blockchain-Aided Secure Semantic Communication for AI-Generated Content in Metaverse,

    Y . Lin, H. Du, D. Niyato, J. Nie, J. Zhang, Y . Cheng, and Z. Yang, “Blockchain-Aided Secure Semantic Communication for AI-Generated Content in Metaverse,” IEEE open j. Comput. Soc. , vol. 4, pp. 72–83, 2023

  5. [13]

    A Contract-Ruled Economic Model for QoS Guarantee in Mobile Peer-to-Peer Streaming Services,

    L. Yang and W. Lou, “A Contract-Ruled Economic Model for QoS Guarantee in Mobile Peer-to-Peer Streaming Services,” IEEE Trans. Mobile Comput. , vol. 15, no. 5, pp. 1047–1061, 2016

  6. [14]

    Vision- Based Semantic Communications for Metaverse Services: A Contest Theoretic Approach,

    G. Liu, H. Du, D. Niyato, J. Kang, Z. Xiong, and B. H. Soong, “Vision- Based Semantic Communications for Metaverse Services: A Contest Theoretic Approach,” in IEEE Global Commun. Conf. IEEE, 2023, pp. 2426–2432

  7. [15]

    High- Resolution Image Synthesis With Latent Diffusion Models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- Resolution Image Synthesis With Latent Diffusion Models,” in Proc. IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 10 684–10 695

  8. [16]

    Disco-Diffusion,

    alembics, “Disco-Diffusion,” 2022. [Online]. Available: https://github. com/alembics/disco-diffusion

  9. [17]

    A Very Preliminary Analysis of DALL-E 2,

    G. Marcus, E. Davis, and S. Aaronson, “A Very Preliminary Analysis of DALL-E 2,” arXiv preprint arXiv:2204.13807 , 2022

  10. [18]

    EdgeFusion: On-Device Text-to-Image Generation,

    T. Castells, H.-K. Song, T. Piao, S. Choi, B.-K. Kim, H. Yim, C. Lee, J. G. Kim, and T.-H. Kim, “EdgeFusion: On-Device Text-to-Image Generation,” arXiv preprint arXiv:2404.11925 , 2024

  11. [19]

    Snapfusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds,

    Y . Li, H. Wang, Q. Jin, J. Hu, P. Chemerys, Y . Fu, Y . Wang, S. Tulyakov, and J. Ren, “Snapfusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds,” Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2024

  12. [20]

    Exploring Collaborative Distributed Diffusion-Based AI-Generated Content (AIGC) in Wireless Networks,

    H. Du, R. Zhang, D. Niyato, J. Kang, Z. Xiong, D. I. Kim, X. S. Shen, and H. V . Poor, “Exploring Collaborative Distributed Diffusion-Based AI-Generated Content (AIGC) in Wireless Networks,” IEEE Netw. , 2023

  13. [21]

    C. L. C. and M. A. Marini, Contest Theory. Edward Elgar Publishing, 2018, vol. II, p. 221–221

  14. [22]

    Incentive Mechanism Design for Heterogeneous Crowdsourcing Using All-Pay Contests,

    T. Luo, S. S. Kanhere, S. K. Das, and H.-P. Tan, “Incentive Mechanism Design for Heterogeneous Crowdsourcing Using All-Pay Contests,” IEEE Trans. Mobile Comput. , vol. 15, no. 9, pp. 2234–2246, 2016

  15. [23]

    A Novel Mobile Data Contract Design with Time Flexibility,

    Y . Wei, J. Yu, T.-M. Lok, and L. Gao, “A Novel Mobile Data Contract Design with Time Flexibility,” IEEE Trans. Mobile Comput. , vol. 18, no. 5, pp. 986–999, 2019

  16. [24]

    Multi- Dimensional Incentive Mechanism in Mobile Crowdsourcing with Moral Hazard,

    Y . Zhang, Y . Gu, M. Pan, N. H. Tran, Z. Dawy, and Z. Han, “Multi- Dimensional Incentive Mechanism in Mobile Crowdsourcing with Moral Hazard,” IEEE Trans. Mobile Comput. , vol. 17, no. 3, pp. 604–616, 2017

  17. [25]

    AI- Generated Incentive Mechanism and Full-Duplex Semantic Communi- cations for Information Sharing,

    H. Du, J. Wang, D. Niyato, J. Kang, Z. Xiong, and D. I. Kim, “AI- Generated Incentive Mechanism and Full-Duplex Semantic Communi- cations for Information Sharing,” IEEE J. Select. Areas Commun. , 2023

  18. [26]

    Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,

    H. Du, R. Zhang, Y . Liu, J. Wang, Y . Lin, Z. Li, D. Niyato, J. Kang, Z. Xiong, S. Cui et al. , “Enhancing Deep Reinforcement Learning: A Tutorial on Generative Diffusion Models in Network Optimization,” IEEE Commun. Surveys Tuts. , 2024

  19. [27]

    Ching and M

    W. Ching and M. Ng, Markov Chains: Models, Algorithms and Applications, ser. International Series in Operations Research & Management Science. Springer, 2006. [Online]. Available: https: //books.google.com.sg/books?id=dr3fm2SAvr4C

  20. [28]

    Generative AI-Enabled Vehicular Networks: Fundamentals, Framework, and Case Study,

    R. Zhang, K. Xiong, H. Du, D. Niyato, J. Kang, X. Shen, and H. V . Poor, “Generative AI-Enabled Vehicular Networks: Fundamentals, Framework, and Case Study,” IEEE Netw., 2024

  21. [29]

    Generative Diffusion Model (GDM) for Optimization of Wi-Fi Networks,

    T. Liu, X. Fang, and R. He, “Generative Diffusion Model (GDM) for Optimization of Wi-Fi Networks,” arXiv preprint arXiv:2404.15684 , 2024

  22. [30]

    Shannon-Theoretic Approach to A Gaussian Cellular Multiple-Access Channel with Fading,

    O. Somekh and S. Shamai, “Shannon-Theoretic Approach to A Gaussian Cellular Multiple-Access Channel with Fading,” IEEE Trans. Inform. Theory, vol. 46, no. 4, pp. 1401–1425, 2000

  23. [31]

    A Lite Distributed Semantic Communication System for Internet of Things,

    H. Xie and Z. Qin, “A Lite Distributed Semantic Communication System for Internet of Things,” IEEE J. Select. Areas Commun. , vol. 39, no. 1, pp. 142–153, 2020

  24. [32]

    A Simple Approach to Evaluate the Er- godic Capacity and Outage Probability of Correlated Rayleigh Diversity Channels With Unequal Signal-to-Noise Ratios,

    M. Moinuddin and I. Naseem, “A Simple Approach to Evaluate the Er- godic Capacity and Outage Probability of Correlated Rayleigh Diversity Channels With Unequal Signal-to-Noise Ratios,” EURASIP J. Wireless Commun., vol. 2013, pp. 1–7, 2013

  25. [33]

    Real Time Video Object Segmentation in Compressed Domain,

    Z. Tan, B. Liu, Q. Chu, H. Zhong, Y . Wu, W. Li, and N. Yu, “Real Time Video Object Segmentation in Compressed Domain,” IEEE Trans. Circuits Syst. Video Technol. , vol. 31, no. 1, pp. 175–188, 2020

  26. [34]

    On the Mathematical Properties of the Structural Similarity Index,

    D. Brunet, E. R. Vrscay, and Z. Wang, “On the Mathematical Properties of the Structural Similarity Index,” IEEE Trans. Image Process., vol. 21, no. 4, pp. 1488–1499, 2011

  27. [35]

    ImageReward: Learning and Evaluating Human Preferences for Text- to-Image Generation,

    J. Xu, X. Liu, Y . Wu, Y . Tong, Q. Li, M. Ding, J. Tang, and Y . Dong, “ImageReward: Learning and Evaluating Human Preferences for Text- to-Image Generation,” Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2024

  28. [36]

    Nested Multiple Instance Learning With Attention Mechanisms,

    S. Fuster, T. Eftestøl, and K. Engan, “Nested Multiple Instance Learning With Attention Mechanisms,” in Proc. Int. Conf. Mach. Learn. IEEE, 2022, pp. 220–225

  29. [37]

    Image Segmentation Using Deep Learning: A Survey,

    S. Minaee, Y . Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image Segmentation Using Deep Learning: A Survey,” IEEE Trans. Pattern Anal. Machine Intell., vol. 44, no. 7, pp. 3523–3542, 2021

  30. [38]

    Holistically-Nested Edge Detection,

    S. Xie and Z. Tu, “Holistically-Nested Edge Detection,” in Proc. IEEE Int. Conf. Comput. Vis. , 2015, pp. 1395–1403

  31. [39]

    Real-Time 2D Multi-Person Pose Estimation on CPU: Lightweight Openpose,

    D. Osokin, “Real-Time 2D Multi-Person Pose Estimation on CPU: Lightweight Openpose,” arXiv preprint arXiv:1811.12004 , 2018

  32. [40]

    Denoising Diffusion Probabilistic Models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” Proc. Adv. Neural Inf. Process. Syst. , vol. 33, pp. 6840–6851, 2020

  33. [41]

    Proxi- mal Policy Optimization Algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proxi- mal Policy Optimization Algorithms,” arXiv preprint arXiv:1707.06347 , 2017

  34. [42]

    Laion- 5b: An Open Large-Scale Dataset for Training Next Generation Image- Text Models,

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsmanet al., “Laion- 5b: An Open Large-Scale Dataset for Training Next Generation Image- Text Models,” Proc. Adv. Neural Inf. Process. Syst. , vol. 35, pp. 25 278– 25...

  35. [43]

    Scene Parsing Through ADE20K Dataset,

    B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene Parsing Through ADE20K Dataset,” in Proc. IEEE Conf. Com- put. Vis. Pattern Recog. , 2017, pp. 633–641

  36. [44]

    A Computational Approach to Edge Detection,

    J. Canny, “A Computational Approach to Edge Detection,” IEEE Trans. Pattern Anal. Machine Intell. , no. 6, pp. 679–698, 1986

  37. [45]

    Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer,

    R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V . Koltun, “Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer,” IEEE Trans. Pattern Anal. Machine Intell. , vol. 44, no. 3, pp. 1623–1637, 2020

  38. [46]

    PromptFix: You Prompt and We Fix the Photo,

    Y . Yu, Z. Zeng, H. Hua, J. Fu, and J. Luo, “PromptFix: You Prompt and We Fix the Photo,” arXiv preprint arXiv:2405.16785 , 2024

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.