Pith. sign in

REVIEW 4 major objections 5 minor 68 references

Reporting language itself, independent of incentives, changes how accurately people report preferences in assignment mechanisms.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 19:40 UTC pith:XLAJSIVQ

load-bearing objection A well-designed pre-registered experiment on reporting language in matching markets, but the abstract's sharpest quantitative claim (a one-third menu-size decomposition) does not appear in the body, and the key SD-CHOICE vs. ACCURACY comparison is confounded. the 4 major comments →

arxiv 2511.22834 v2 pith:XLAJSIVQ submitted 2025-11-28 econ.GN q-fin.EC

Complexity Beyond Incentives: The Critical Role of Reporting Language

classification econ.GN q-fin.EC
keywords reporting languagemessage spaceserial dictatorshipmatching marketspreference complexitymulti-attribute preferencesobviously strategy-proofexperimental economics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether the way applicants are asked to report preferences—full ranking, separate attribute rankings, or one-at-a-time choice—affects accuracy, efficiency, and fairness in multi-attribute matching. Using induced utility formulas over 27 program bundles, it finds that people misreport even when there are no allocation incentives, and errors grow when preferences require trade-offs. Simplified attribute-based interfaces do not beat full-ranking reporting, even when the interface exactly matches the preference structure. Sequential serial dictatorship—choosing one program at each turn—achieves the highest accuracy, and a decomposition attributes about one-third of its edge over full-ranking reporting to smaller menus rather than stronger incentives. The message space is therefore a determinant of mechanism performance, not just a detail of implementation.

Core claim

The central claim, stated as a sympathetic reader would: a strategy-proof allocation mechanism can perform very differently depending on how preferences are elicited. In laboratory markets with three-attribute bundles, participants frequently rank the wrong top option even when paid only for accuracy, and the error rate rises with preference complexity. Lexicographic and weighted-attribute interfaces—modeled on real-world restricted formats—fail to improve accuracy, sometimes hurting it even in the preference domain they are designed for. Sequential serial dictatorship outperforms every direct format, including the no-incentive accuracy benchmark, which implies the gain is not only from obvi

What carries the argument

The experimental engine is a 3×5 design: three preference domains (lexicographic, additively separable, and complementary) crossed with five reporting formats, with preferences induced by utility formulas over university×field×tuition bundles. The load-bearing comparison is SD-CHOICE versus ACCURACY: both use the same induced preferences, but ACCURACY rewards a full 27-item ranking with no allocation, while SD-CHOICE asks for one pick at a time from the remaining set. Because SD-CHOICE beats ACCURACY, the difference cannot be incentives alone; menu-size analysis then apportions part of the gain to interface complexity. The two accuracy measures—choice accuracy and normalized Kendall distance

Load-bearing premise

The paper's cleanest causal number—one-third of the sequential choice gain from smaller menus—rests on comparing two treatments that differ in more than menu size, so if feedback, timing, or payment structure contribute, the interface share changes.

What would settle it

Conduct a treatment with one-at-a-time choices but no feedback about which programs have been taken, with the same 8-minute time limit and accuracy-style payment as ACCURACY; if choice accuracy falls to the ACCURACY level, the smaller-menu explanation fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Direct mechanisms should default to full-list reporting over domain-assuming simplifications, since the simplified formats did not help even when they matched the preference structure.
  • Sequential or staged implementations carry an interface benefit, not just an incentive benefit; mechanisms that shrink menus over time should improve accuracy even when incentives are already strategy-proof.
  • Reporting errors occur in consequential parts of the ranking: lower accuracy coincides with higher efficiency loss and more justified envy, so fixing the reporting task is an allocation-quality intervention.
  • In field applications with far larger choice sets, the 27-item results likely understate the reporting burden, making restricted formats even less attractive at scale.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural next experiment is to separate menu size from feedback: present the same small menus without revealing which programs other participants took; if the accuracy edge persists, the menu-size channel is confirmed rather than the feedback channel.
  • The one-third decomposition may understate the interface channel if timing or payoff differences contributed, so field pilots with staged offers could estimate the real-world share.
  • Because attribute-based interfaces failed even when matched to preferences, adaptive elicitation—asking a few questions and letting a machine infer the rest—becomes more attractive than fixed restricted formats.
  • Designers of sequential mechanisms should weigh operational frictions against the measured accuracy gains; the menu-size result supplies a microfoundation for multi-offer or staged systems that approximate small menus without full serialization.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a laboratory experiment (810 subjects, pre-registered AEARCTR-0013135) on how preference complexity and reporting interfaces affect behavior in a strategy-proof assignment mechanism. Participants have induced utility functions over 27 multi-attribute programs, and the design varies preference complexity (lexicographic, separable, and complementary) within subjects, and reporting interface/mechanism between subjects: full-ranking direct serial dictatorship (SD-DIRECT), attribute-based lexicographic (SD-LEX) and weighted-attribute (SD-WEIGHT) interfaces, sequential serial dictatorship (SD-CHOICE), and a non-allocative accuracy benchmark (ACCURACY). The main findings are: (i) misreporting is substantial even in ACCURACY and increases with preference complexity; (ii) SD-DIRECT generates more misreporting than ACCURACY, interpreted as strategic misperception; (iii) simplified attribute-based interfaces do not improve accuracy and often worsen it relative to full ranking; (iv) SD-CHOICE achieves the highest choice accuracy, lowest efficiency loss, and lowest justified envy. The abstract also claims a decomposition attributes roughly one-third of SD-CHOICE's advantage to smaller menus.

Significance. If the results hold, the paper would provide credible experimental evidence that the message space itself—not just incentives—affects the performance of strategy-proof assignment mechanisms, with direct policy relevance for the design of college admissions interfaces (e.g., the lexicographic format used in China) and for sequential mechanisms. The strengths of the paper include the pre-registered design, the between-subjects interface variation, within-subject preference-complexity variation, the pure-accuracy baseline, and careful use of mixed-effects regressions with clustered standard errors. However, the paper's sharpest quantitative claim—the 'roughly one-third' menu-size decomposition—is absent from the body, and the comparison between SD-CHOICE and ACCURACY is confounded with feedback, timing, and payoff differences. These issues are load-bearing for the central 'reporting language' conclusion and currently weaken the paper's contribution relative to its abstract.

major comments (4)
  1. [Abstract; §3.4; Figure 2] The abstract's central quantitative claim—'a decomposition attributes roughly one-third of its advantage over full-ranking reporting to the smaller menus'—does not appear anywhere in the main text or appendices. There is no estimator, no regression table, no confidence interval, and no definition of the denominator (e.g., the SD-CHOICE vs. SD-DIRECT gap or the SD-CHOICE vs. ACCURACY gap). The only supporting evidence is the menu-size binned scatterplot in Figure 2, which is descriptive and does not identify a share. Because this claim is the paper's most precise statement of the reporting-interface channel, it is load-bearing for the conclusion that message space matters independently of incentives. The authors should either add a proper decomposition with explicit identification assumptions and standard errors, or remove the 'one-third' claim from the abstract and the concluding paragra
  2. [§3.4; §2.1.3; §2.2.2] The comparison SD-CHOICE versus ACCURACY (84.26% vs. 76.72%, Table 2) is used to argue that 'part of the improvement reflects the simpler, one-at-a-time reporting interface rather than incentives alone' (Result 4). However, the two treatments differ on several dimensions beyond menu size: SD-CHOICE runs 12 parallel processes with a 3-minute per-move limit, provides turn-based feedback about the actual remaining set, reveals priority timing, and pays by allocation outcome (rank-based payoffs), whereas ACCURACY uses an 8-minute single-list construction and pays by Kendall distance. Any of these differences—not just the smaller menus—could drive the accuracy gap. Moreover, the Figure 2 menu-size analysis is confounded: in SD-CHOICE, menu size is mechanically determined by priority, so the large-menu bins (25–27) contain only the highest-priority participants. Without a treatment that varies
  3. [§2.2.1; §3.1; Table 2] The COMP preference specification is s = aU + bF − cT + d·U·T with a,b,c∈[30,40] and d∈[−5,5]. Because U∈{200,500,600} and T∈{0,250,500}, the interaction term can be as large as ±1.5 million, while the main effects are at most on the order of 72,000. Realized COMP preferences may therefore be interaction-dominated rather than mildly non-separable, potentially producing non-monotonic or effectively lexicographic orderings. This undermines the intended complexity ordering (SEP vs. COMP) and may explain the null difference between SEP and COMP reported in Result 1. The authors should report the realized distribution of preferences generated by their draws, verify that the intended 'non-separable with complementarities' structure is actually achieved, and consider restricting d so that the interaction is a genuine perturbation. There is also an inconsistency: the Introduction describes the c
  4. [§3.2; §2.1.3] Result 2 attributes the ACCURACY minus SD-DIRECT gap to 'incentive misperception' and states that 'this gap is not attributable to list-length burden, preference complexity or demand effect.' But the payoff functions differ across the two treatments: ACCURACY pays 160×(1−Kendall), penalizing all discordant pairs equally, while SD-DIRECT pays by the rank of the allocated seat, placing extreme weight on the top of the list. The different incentive to exert effort at various margins could by itself produce the observed accuracy gap, independent of strategic misperception. The authors should either control for this difference (e.g., by comparing SD-DIRECT to an allocation-free baseline with the same rank-based payoff for the true assigned seat) or explicitly acknowledge that the gap conflates misperceived incentives with different payoff-induced effort.
minor comments (5)
  1. [Appendix C (SD-DIRECT instructions)] The payoff table lists '13th Preference CNY110' after '12th Preference CNY105'; the pattern implies the 13th rank should be CNY100. This typo in the instructions could confuse participants and should be corrected.
  2. [§3.4] The sentence 'for very large menus (more than 25 options), the accuracy of the top choice in ACCURACY is significantly higher than the accuracy of the chosen option in SD-CHOICE' is confusing because it compares a simulated counterfactual (ACCURACY top choice under random priority) with an actual decision in SD-CHOICE. Please clarify what is being compared and why this comparison supports the menu-size interpretation.
  3. [Figure 2] The notes describe Figure 2 as a scatter plot, but the figure displays binned bar-like data. Either change the description or present the underlying scatterplot.
  4. [Abstract vs §3.4] The abstract says 'a decomposition attributes roughly one-third' while the body only says 'part of the improvement reflects the simpler interface.' The language should be aligned after the decomposition issue (Major Comment 1) is resolved.
  5. [References] Several references have inconsistent diacritics (e.g., 'Sönmez' vs. 'Sönmez') and at least one duplicated entry (Calsamiglia et al. 2010a and 2010b). A careful proofread of the bibliography is needed.

Circularity Check

0 steps flagged

No significant circularity: the experimental claims are tested, not fitted; the abstract's 'one-third' decomposition is unverifiable in the body, but this is a missing-evidence problem, not a circular reduction.

full rationale

The paper's derivation chain is experimental rather than formal. Induced utility formulas, interfaces, and payoff rules are exogenously fixed inputs; accuracy, efficiency loss, and justified envy are measured outcomes. No outcome parameter is estimated and then fed back into the design or into a hypothesis. The central comparisons—ACCURACY vs SD-DIRECT, SD-CHOICE vs ACCURACY, SD-LEX/SD-WEIGHT vs SD-DIRECT—are within-paper treatment contrasts and do not reduce by construction to any fitted quantity. Self-citations appear (e.g., Bó and Hakimov 2024; Hakimov et al. 2023; Guillen and Hakimov 2018; Caspari and Khanna 2025), but they are used as literature context or as replication anchors; the load-bearing evidence for the incentive-misperception and interface effects comes from the paper's own ACCURACY and SD-CHOICE comparisons. Citations of Li (2017) and Pycia and Troyan (2023) are external theory, not the authors' own prior work, and are not invoked to forbid alternative explanations. The abstract's claim that 'a decomposition attributes roughly one-third of its advantage over full-ranking reporting to the smaller menus' is not supported anywhere in the main text or appendices: Section 3.4 reports only aggregate accuracy comparisons and a menu-size scatterplot, with no estimator, regression table, or identification statement. This is an unsupported quantitative assertion and a verification problem, but it is not circular, because the one-third number is not constructed from the paper's own inputs or from a fitted parameter renamed as a prediction. Similarly, the confounds between SD-CHOICE and ACCURACY (parallel play, 3-minute vs 8-minute limits, payoff by allocation vs Kendall score) weaken causal interpretation but do not make the design circular. Overall, no equation is equal to its input by definition, no fitted value is relabeled as a prediction, and no load-bearing premise reduces to a self-citation. Score 2 reflects only the presence of minor, non-load-bearing self-citation.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

No outcome coefficient is estimated and then reused as an input, so the empirical claims do not rest on fitted parameters. However, all interpretations depend on domain assumptions: the induced-value assumption about formula use, the arithmetic claim that LEX ranges enforce lexicographic structure, the maintained behavioral model behind 'strategic misreporting,' and the fidelity of the Random Priority simulation for the ACCURACY baseline. The hand-chosen design parameters (coefficient ranges, attribute values, interface weights, payoff schedules) define the environments compared and are transparent but not varied.

free parameters (5)
  • LEX coefficient ranges (a∈[90,110], b∈[9,11], c∈[0.9,1.1]) = hand-chosen ranges
    Chosen to enforce lexicographic dominance U≻F≻−T on the feasible attribute ranges; a design parameter that defines the simplest preference domain, not fitted to outcomes.
  • SEP/COMP coefficient ranges (a,b,c∈[30,40]; d∈[−5,5]) = hand-chosen ranges
    Chosen to make attributes trade off (SEP) or interact (COMP); determines the difficulty of the task and the structure of the induced rankings.
  • Attribute point values (U∈{200,500,600}, F∈{300,500,700}, T∈{0,250,500}) = fixed values
    Attribute levels entering the utility formulas; non-uniform spacing is what makes SD-WEIGHT partially inexpressive under SEP (acknowledged in §2.2.2), so these values shape the paper's expressiveness conclusions.
  • SD-LEX aggregation weights (100,10,1) and SD-WEIGHT weight bounds [1,100] = 100/10/1; weights 1–100
    Interface design choices defining what restricted reports can express; they determine the representational losses that are central to the comparison between SD-LEX, SD-WEIGHT, and SD-DIRECT.
  • Payoff schedule (CNY160 decreasing by 5 per rank to CNY30) and ACCURACY payoff 160×(1−Kendall) = rank-linear; Kendall-linear
    By construction payoffs depend only on the rank of the assigned program (or on accuracy), making cardinal scores payoff-irrelevant and aligning incentives across treatments.
axioms (4)
  • domain assumption Induced-value assumption: subjects' true preferences equal the ordinal ranking generated by the provided utility formula, which they compute using the on-screen calculator.
    Standard in induced-value experiments; the paper deliberately uses formulas rather than direct rankings (footnote 14) so subjects cannot copy pre-filled tables, but this assumes subjects engage with the formula and interpret scores correctly.
  • standard math Lexicographic dominance holds for LEX draws: coefficient separation guarantees U≻F≻−T over the feasible attribute ranges.
    Invoked in §2.2.1. Checkable arithmetic: min a·ΔU = 90×400 = 36,000 exceeds max (bΔF − cΔT) ≈ 4,950, so the claim holds for the stated ranges.
  • domain assumption Deviations from the formula-induced ranking are interpretable as error or incentive misperception, and the ACCURACY-vs-SD-DIRECT gap is strategic misreporting.
    Maintained behavioral model behind Result 2; it assumes list-length burden, preference complexity, and demand effects are held fixed across the two treatments.
  • domain assumption The 100 simulated Random Priority markets in ACCURACY faithfully reproduce the choice sets participants face in SD-DIRECT and SD-CHOICE.
    Choice accuracy for ACCURACY is computed by simulation (§3.1) because there is no allocation; cross-treatment comparability of the accuracy measure requires this fidelity, and regressions use only the first simulation (Table 3 note).

pith-pipeline@v1.3.0-alltime-deepseek · 4971 in / 5260 out tokens · 196774 ms · 2026-08-03T19:40:20.339069+00:00 · methodology

0 comments
read the original abstract

Mechanisms specify both allocation rules and message spaces. We study how message spaces affect behavior in a laboratory assignment environment in which objects are bundles of three attributes and preferences are induced by utility formulas. We vary preference complexity and compare full-ranking reports, two attribute-based interfaces, and sequential choice under serial dictatorship. Participants make frequent reporting errors even in a treatment that rewards accurate reporting without any allocation, and errors are more frequent when preferences require trade-offs across attributes. Attribute-based interfaces do not improve accuracy: conditional on what they can express, restricted reports track preferences comparatively well, but representational losses---large for lexicographic reports, small for weighted-attribute reports within our preference domains---offset these gains. Sequential choice yields more accurate assignments and lower efficiency loss and less justified envy; a decomposition attributes roughly one-third of its advantage over full-ranking reporting to the smaller menus that participants face. The results show that the message space affects the performance of strategy-proof assignment mechanisms.

Figures

Figures reproduced from arXiv: 2511.22834 by Manshu Khanna, Rustamdjan Hakimov.

Figure 3
Figure 3. Figure 3: Choice Accuracy (CI: 95%) 0 20 40 60 80 100 Choice Accuracy (%) 1−3 4−6 7−9 10−12 LEX 0 20 40 60 80 100 Choice Accuracy (%) 1−3 4−6 7−9 10−12 SEP 0 20 40 60 80 100 Choice Accuracy (%) 1−3 4−6 7−9 10−12 COMP 0 20 40 60 80 100 Choice Accuracy (%) 1−3 4−6 7−9 10−12 Overall ACCURACY SD−DIRECT SD−WEIGHT SD−LEX Notes: X-axis reports rounds. Recall that the experiment design is such that each preference domain is… view at source ↗
Figure 5
Figure 5. Figure 5: Efficiency Loss (CI: 95%) 0 5 10 15 Efficiency Loss (%) 1−3 4−6 7−9 10−12 LEX 0 5 10 15 Efficiency Loss (%) 1−3 4−6 7−9 10−12 SEP 0 5 10 15 Efficiency Loss (%) 1−3 4−6 7−9 10−12 COMP 0 5 10 15 Efficiency Loss (%) 1−3 4−6 7−9 10−12 Overall ACCURACY SD−DIRECT SD−WEIGHT SD−LEX 40 [PITH_FULL_IMAGE:figures/full_fig_p040_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 2 linked inside Pith

  1. [1]

    and S \"o nmez, T

    Abdulkadiro g lu, A. and S \"o nmez, T. (1998). Random serial dictatorship and the core from random endowments in house allocation problems. Econometrica, 66 (3), 689--701

  2. [2]

    --- and S \"o nmez, T. (1999). House allocation with existing tenants. Journal of Economic Theory, 88 (2), 233--260

  3. [3]

    and S \"o nmez, T

    Abdulkadiroglu, A. and S \"o nmez, T. (2003). School choice: A mechanism design approach. The American Economic Review, 93 (3), 729--747

  4. [4]

    and S \"o nmez, T

    Balinski, M. and S \"o nmez, T. (1999). A tale of two mechanisms: student placement. Journal of Economic Theory, 84 (1), 73--94

  5. [5]

    and Hakimov, R

    B \'o , I. and Hakimov, R. (2020). Iterative versus standard deferred acceptance: Experimental evidence. The Economic Journal, 130 (626), 356--392

  6. [6]

    --- and Hakimov, R. (2024). Pick-an-object mechanisms. Management Science, 70 (7), 4693--4721

  7. [7]

    Budish, E. (2011). The combinatorial assignment problem: Approximate competitive equilibrium from equal incomes. Journal of Political Economy, 119 (6), 1061--1103

  8. [8]

    --- and Cantillon, E. (2012). The multi-unit assignment problem: Theory and evidence from course allocation at harvard. American Economic Review, 102 (5), 2237--2271

  9. [9]

    --- and Kessler, J. B. (2022). Can market participants report their preferences accurately (enough)? Management Science, 68 (2), 1107--1130

  10. [10]

    , Haeringer, G

    Calsamiglia, C. , Haeringer, G. and Klijn, F. (2010 a ). Constrained school choice: An experimental study. American Economic Review, 100 (4), 1860--1874

  11. [11]

    Constrained school choice: An experimental study

    --- , --- and --- (2010 b ). Constrained school choice: An experimental study. American Economic Review, 100 (4), 1860--1874

  12. [12]

    and Khanna, M

    Caspari, G. and Khanna, M. (2025). Nonstandard choice in matching markets. International Economic Review, 66 (2), 757--786

  13. [13]

    and Pereyra, J

    Chen, L. and Pereyra, J. S. (2019). Self-selection in school choice. Games and Economic Behavior, 117, 59--81

  14. [14]

    , Katu s c \'a k, P

    Chen, R. , Katu s c \'a k, P. , Kittsteiner, T. and K \"u tter, K. (2024). Does disappointment aversion explain non-truthful reporting in strategy-proof mechanisms? Experimental Economics, 27 (5), 1184--1210

  15. [15]

    and He, Y

    Chen, Y. and He, Y. (2021). Information acquisition and provision in school choice: an experimental study. Journal of Economic Theory, 197, 105345

  16. [16]

    Information acquisition and provision in school choice: a theoretical investigation

    --- and --- (2022). Information acquisition and provision in school choice: a theoretical investigation. Economic Theory, 74 (1), 293--327

  17. [17]

    --- and S \"o nmez, T. (2006). School choice: an experimental study. Journal of Economic theory, 127 (1), 202--231

  18. [18]

    , B \"o ckenholt, U

    Chernev, A. , B \"o ckenholt, U. and Goodman, J. (2015). Choice overload: A conceptual review and meta-analysis. Journal of Consumer Psychology, 25 (2), 333--358

  19. [19]

    and Jehiel, P

    Compte, O. and Jehiel, P. (2004). The wait-and-see option in ascending price auctions. Journal of the European Economic Association, 2 (2-3), 494--503

  20. [20]

    , Glicksohn, O

    Dreyfuss, B. , Glicksohn, O. , Heffetz, O. and Romm, A. (2022). Deferred acceptance with news utility. Tech. rep., National Bureau of Economic Research

  21. [21]

    , Hammond, R

    Dur, U. , Hammond, R. G. and Kesten, O. (2021). Sequential school choice: Theory and evidence from the field and lab. Journal of Economic Theory, 198, 105344

  22. [22]

    and Shapley, L

    Gale, D. and Shapley, L. S. (1962). College admissions and the stability of marriage. The American Mathematical Monthly, 69 (1), 9--15

  23. [23]

    and Kahneman, D

    Gigerenzer, G. and Kahneman, D. (2008). Heuristics and biases: The psychology of intuitive judgment. Cambridge University Press

  24. [24]

    Gonczarowski, Y. A. , Heffetz, O. , Ishai, G. and Thomas, C. (2024). Describing Deferred Acceptance and Strategyproofness to Participants: Experimental Analysis. Tech. rep., National Bureau of Economic Research

  25. [25]

    --- , --- and Thomas, C. (2023). Strategyproofness-exposing mechanism descriptions. Tech. rep., National Bureau of Economic Research

  26. [26]

    and Liang, Y

    Gong, B. and Liang, Y. (2024). A dynamic matching mechanism for college admissions: Theory and experiment. Management Science

  27. [27]

    , Pathak, P

    Greenberg, K. , Pathak, P. A. and S \"o nmez, T. (2024). Redesigning the us army’s branching process: A case study in minimalist market design. American Economic Review, 114 (4), 1070--1106

  28. [28]

    Grenet, J. , He, Y. and K \"u bler, D. (2022). Preference discovery in university admissions: The case for dynamic multioffer mechanisms. Journal of Political Economy, 130 (6), 1427--1476

  29. [29]

    and Hakimov, R

    Guillen, P. and Hakimov, R. (2018). The effectiveness of top-down advice in strategy-proof mechanisms: A field experiment. European Economic Review, 101, 505--511

  30. [30]

    --- and Veszteg, R. F. (2021). Strategy-proofness in experimental matching markets. Experimental Economics, 24, 650--668

  31. [31]

    and Iehl \'e , V

    Haeringer, G. and Iehl \'e , V. (2021). Gradual college admission. Journal of Economic Theory, 198, 105378

  32. [32]

    --- and Klijn, F. (2009). Constrained school choice. Journal of Economic Theory, 144 (5), 1921--1947

  33. [33]

    , Heller, C.-P

    Hakimov, R. , Heller, C.-P. , K \"u bler, D. and Kurino, M. (2021). How to avoid black markets for appointments with online booking systems. American Economic Review, 111 (7), 2127--2151

  34. [34]

    --- and K \"u bler, D. (2021). Experiments on centralized school choice and college admissions: a survey. Experimental Economics, 24 (2), 434--488

  35. [35]

    and Pan, S

    --- , K \"u bler, D. and Pan, S. (2023). Costly information acquisition in centralized matching markets. Quantitative Economics, 14 (4), 1447--1490

  36. [36]

    , Romm, A

    Hassidim, A. , Romm, A. and Shorrer, R. I. (2021). The limits of incentives in economic matching procedures. Management Science, 67 (2), 951--963

  37. [37]

    Hatfield, J. W. , Kominers, S. D. and Westkamp, A. (2020). Stability, strategy-proofness, and cumulative offer mechanisms. The Review of Economic Studies, 88 (3), 1457--1502

  38. [38]

    --- and Milgrom, P. R. (2005). Matching with contracts. American Economic Review, 95 (4), 913--935

  39. [39]

    , Yao, L

    Hu, X. , Yao, L. and Zhang, J. (2025). Reforming china's two-stage college admission: An experimental study. Available at SSRN 5207493

  40. [40]

    and Zhang, J

    Huang, L. and Zhang, J. (2025). Bundled school choice. arXiv preprint arXiv:2501.04241

  41. [41]

    Kahneman, D. (2003). Maps of bounded rationality: Psychology for behavioral economics. American Economic Review, 93 (5), 1449--1475

  42. [42]

    and Kittsteiner, T

    Katu s c \'a k, P. and Kittsteiner, T. (2024). Strategy-proofness made simpler. Management Science

  43. [43]

    , Pais, J

    Klijn, F. , Pais, J. and Vorsatz, M. (2019). Static versus dynamic deferred acceptance in school choice: Theory and experiment. Games and Economic Behavior, 113, 147--163

  44. [44]

    and Troyan, P

    Kloosterman, A. and Troyan, P. (2023). Rankings-dependent preferences: A real goods matching experiment. arXiv preprint arXiv:2305.03644

  45. [45]

    Kwasnica, A. M. , Ledyard, J. O. , Porter, D. and DeMartini, C. (2005). A new and improved design for multiobject iterative auctions. Management Science, 51 (3), 419--434

  46. [46]

    Li, S. (2017). Obviously strategy-proof mechanisms. American Economic Review, 107 (11), 3257--3287

  47. [47]

    Luflade, M. (2017). The value of information in centralized school choice systems (job market paper). Job Market Paper, Duke University

  48. [48]

    and Zhou, Y

    Mackenzie, A. and Zhou, Y. (2022). Menu mechanisms. Journal of Economic Theory, 204, 105511

  49. [49]

    Meisner, V. (2023). Report-dependent utility and strategy-proofness. Management Science, 69 (5), 2733--2745

  50. [50]

    --- and Von Wangenheim, J. (2023). Loss aversion in strategy-proof school-choice mechanisms. Journal of Economic Theory, 207, 105588

  51. [51]

    Milgrom, P. (2009). Assignment messages and exchanges. American Economic Journal: Microeconomics, 1 (2), 95--113

  52. [52]

    Critical issues in the practice of market design

    --- (2011). Critical issues in the practice of market design. Economic Inquiry, 49 (2), 311--320

  53. [53]

    Parkes, D. C. (2005). Auction design with costly preference elicitation. Annals of Mathematics and Artificial Intelligence, 44, 269--302

  54. [54]

    and Troyan, P

    Pycia, M. and Troyan, P. (2023). A theory of simplicity in games and mechanism design. Econometrica, 91 (4), 1495--1526

  55. [55]

    The random priority mechanism is uniquely simple, efficient, and fair

    --- and --- (2024). The random priority mechanism is uniquely simple, efficient, and fair. CEPR Discussion Papers, (19689)

  56. [56]

    Rees-Jones, A. (2018). Suboptimal behavior in strategy-proof mechanisms: Evidence from the residency match. Games and Economic Behavior, 108, 317--330

  57. [57]

    --- and Shorrer, R. (2023). Behavioral economics in education market design: A forward-looking review. Journal of Political Economy Microeconomics, 1 (3), 557--613

  58. [58]

    --- and Skowronek, S. (2018). An experimental investigation of preference misrepresentation in the residency match. Proceedings of the National Academy of Sciences, 115 (45), 11471--11476

  59. [59]

    Roth, A. E. (2002). The economist as engineer: Game theory, experimentation, and computation as tools for design economics. Econometrica, 70 (4), 1341--1378

  60. [60]

    o nmez, T. and \

    --- , S \"o nmez, T. and \"U nver, M. U. (2004). Kidney exchange. The Quarterly journal of economics, 119 (2), 457--488

  61. [61]

    and Boutilier, C

    Sandholm, T. and Boutilier, C. (2005). Preference elicitation in combinatorial auctions. In Combinatorial Auctions, The MIT Press

  62. [62]

    Shorrer, R. I. and S \'o v \'a g \'o , S. (2023). Dominated choices in a strategically simple college admissions environment. Journal of Political Economy Microeconomics, 1 (4), 781--807

  63. [63]

    --- and S \'o v \'a g \'o , S. (2024). Dominated choices under deferred acceptance mechanism: The effect of admission selectivity. Games and Economic Behavior, 144, 167--182

  64. [64]

    S \"o nmez, T. (2013). Bidding for army career specialties: Improving the rotc branching mechanism. Journal of Political Economy, 121 (1), 186--219

  65. [65]

    and Switzer, T

    S \"o nmez, T. and Switzer, T. B. (2013). Matching with (branch-of-choice) contracts at the united states military academy. Econometrica, 81 (2), 451--488

  66. [66]

    --- and \"U nver, M. U. (2010). Course bidding at business schools. International Economic Review, 51 (1), 99--123

  67. [67]

    , Jiang, Y

    Soumalias, E. , Jiang, Y. , Zhu, K. , Curry, M. , Seuken, S. and Parkes, D. C. (2025). LLM -powered preference elicitation in combinatorial assignment. arXiv preprint arXiv:2502.10308

  68. [68]

    Stephenson, D. (2022). Assignment feedback in school choice mechanisms. Experimental Economics, 25 (5), 1467--1491