{"id":"38585693-57c6-498e-8121-cca9b7a29683","arxiv_id":"1908.07962","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Ordinal embedding methods recover one-dimensional monotonic perceptual scales as accurately as MLDS and can additionally recover non-monotonic and multi-dimensional scales from triplet comparisons.","lead":"This paper shows that machine-learning ordinal embedding methods can estimate perceptual scales from triplet comparisons, handling non-monotonic and multi-dimensional perceptions that traditional psychophysical methods struggle with. A generalist reader might care because this offers a practical way to measure subjective experience, such as color or pitch perception, from simple relative judgments.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Multi-dimensional branch of the central claim is under-validated: the only 2D simulation uses an NMDS ground truth, and the real experiment checks only triplet consistency, not whether the recovered coordinates form a meaningful perceptual scale.","rationale":"The 1D monotonic claim is well supported by simulation against known functions, and the real slant experiment shows comparable cross-validated triplet error. The non-monotonic 1D simulation is also convincing because the ground truth is a closed-form function and MLDS is handicapped by its own monotonicity assumption. The multi-dimensional branch, however, is the part that makes the paper more than a methods transfer: it is what justifies replacing MLDS as a default. That branch depends on two unverified links: (i) the ground truth in the only 2D simulation is itself an NMDS Euclidean embedding, so the simulation is partially circular; (ii) the real Eidolon experiment validates only triplet-consistency, not the perceptual meaning of the recovered coordinates. The paper is transparent about both points — Section 3.3 disclaims correctness of the NMDS ground truth, and Section 6.1 lists interpretation as an open issue — which is why this is a conditional-validation gap rather than an internal inconsistency. A simple external-validation analysis against the known stimulus parameters would either defuse the concern or confirm it. If the check fails, the paper's higher-dimensional claims should be downgraded to 'the method produces low-error triplet predictors' rather than 'estimates perceptual scales'; if it passes, the central claim is materially strengthened. Since the current verdict is already CONDITIONAL, this stress-test does not move the verdict.","tokens_in":26835,"tokens_out":11226,"duration_ms":111749,"concrete_test":"Run a single external-validation analysis on the Eidolon data: fit t-STE in d=2, then test whether the recovered Euclidean distances predict a fresh set of human triplet answers better than a baseline logistic model using the known physical parameters (reach, grain, coherence) directly. If the embedding does not beat the physical-parameter baseline, or if the recovered coordinates show no interpretable relation to those parameters, the multi-dimensional 'perceptual scale' claim is not established beyond triplet-consistency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is in the multi-dimensional part of the central claim, not in the 1D monotonic comparison. In Section 3.3 the only multi-dimensional simulation constructs its ground truth by running NMDS on Ekman's color-similarity data: 'we first construct a two-dimensional embedding using NMDS; ... we do not argue that this embedding is \"correct\" in any way; we just use it as a ground truth'. Because NMDS and the tested ordinal embedding methods share the same Euclidean-distance model, this simulation demonstrates only that t-STE can recover an NMDS-style Euclidean configuration from triplets, not that the recovered coordinates are the true perceptual scale. In the real Eidolon experiment, the evidence for a multi-dimensional scale is the drop in cross-validated triplet error from d=1 to d=2 (Figure 9). But cross-validated triplet error is a predictive-consistency measure: the embedding is trained on exactly this kind of triplet answer, so a low validation error shows that the Euclidean embedding summarizes the triplet answers, not that the two recovered dimensions correspond to a perceptually meaningful scale. The paper itself lists 'Interpreting the embedding' as an open issue (Section 6.1) and does not report any relation between the recovered coordinates and the known generation parameters (reach, grain, coherence). Without such an external anchor, the claim that ordinal embedding 'can produce a desirable scaling function' in higher dimensions is underdetermined; the claim could be true, but the current evidence does not separate a valid perceptual scale from a flexible compression of the training signal.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using ordinal embedding methods (STE, t-STE, LOE) to estimate perceptual scaling functions from triplet comparison data, in contrast to the traditional psychophysical methods NMDS and MLDS. It argues that ordinal embedding requires fewer triplets than NMDS, does not impose monotonicity as MLDS does, and can handle multi-dimensional perceptual spaces. The authors support this with simulations covering one-dimensional monotonic and non-monotonic scaling functions and a two-dimensional perceptual space, plus two real experiments (slant-from-texture and Eidolon image distortion). They conclude that in the one-dimensional monotonic case ordinal embedding performs comparably to MLDS, while in non-monotonic and multi-dimensional cases it is the only method among those considered that yields a desirable scaling function.","tokens_in":27032,"tokens_out":6543,"duration_ms":57286,"significance":"If the central claim holds, the paper provides psychophysics with a practical 'default' scaling algorithm based on triplet comparisons, with weaker assumptions than MLDS and lower data requirements than NMDS. The simulation study is extensive and well controlled, covering four noise levels and five triplet fractions, and the real experiments use cross-validated triplet error with a useful human-baseline calibration for the Eidolon data. The comparison with MLDS in the monotonic one-dimensional case is informative, and the non-monotonic one-dimensional simulation makes a clear point. The main limitation is that the multi-dimensional part of the central claim is only validated by a simulation whose ground truth is itself an NMDS Euclidean embedding, and by a real experiment that checks predictive consistency but not the interpretability of the recovered dimensions.","major_comments":[{"comment":"The only multi-dimensional simulation uses as ground truth the two-dimensional NMDS embedding of Ekman's color-similarity data, with the authors explicitly stating that 'we do not argue that this embedding is \"correct\" in any way; we just use it as a ground truth.' Because NMDS and the tested ordinal-embedding methods both model dissimilarity with Euclidean distances, this simulation demonstrates that t-STE can recover an NMDS-style Euclidean configuration from triplet answers, but it does not validate that the recovered coordinates equal the true perceptual scale. The abstract's claim that in higher dimensions 'only our ordinal embedding methods can produce a desirable scaling function' therefore rests on this simulation plus the Eidolon experiment, and the latter provides only predictive consistency, not perceptual interpretability.","section":"Section 3.3, Figure 6(a)"},{"comment":"The evidence that the Eidolon perceptual space is two-dimensional is the decrease in cross-validated triplet error from d=1 to d=2. Cross-validated triplet error measures agreement with the same type of triplet answers used for training; an embedding that fits triplets well need not have coordinate axes that correspond to any meaningful perceptual attribute. The paper itself lists 'Interpreting the embedding' as an open issue in Section 6.1, and no analysis is reported relating the recovered coordinates to reach, grain, or coherence. To support the multi-dimensional scaling claim, the authors should either provide an external validation (for example, correlation of embedding coordinates with the generation parameters, or a separate judgment task) or explicitly limit the claim to predictive consistency of triplet answers.","section":"Section 4.1, Figure 9"}],"minor_comments":[{"comment":"The text says the average MSE and triplet errors for the monotonic simulation are depicted in 'Figure 5 (d) and (e)', but the relevant panels are in Figure 4; this cross-reference should be corrected.","section":"Section 3.2.1"},{"comment":"The caption describes the comparison as 'the ordinal embedding methods (MLDS, STE and TSTE)' but MLDS is not an ordinal embedding method; the caption should list LOE, STE, and t-STE.","section":"Figure 6 caption"},{"comment":"There is a typo in the text: 'called chroma and height by Shapard' should read 'by Shepard'.","section":"Section 2.3.2"},{"comment":"The side experiment on triplet difficulty should clarify the relationship between the measured percentage of inconsistent repeated answers (9.2%, 9.8%, 11%) and the statement that 'we would expect at least 0.10 triplet error'; as written, the step from the observed inconsistency rate to the 0.10 error floor is not immediate.","section":"Section 4.1"},{"comment":"The description of choosing the best of 10 restarts by 'the least triplet error' should specify whether this is the training-set triplet error, since the definition in Equation (5) allows both training and validation variants.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The one-dimensional and non-monotonic results are solid and the paper is well within the journal's scope. The main weakness is the multi-dimensional branch of the central claim; if the authors can add an external validation of the Eidolon embedding or soften the claim to predictive consistency, the paper would be a strong contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThis is a methods-transfer paper that mostly works. The authors take t-STE, STE, and LOE from the ML ordinal-embedding literature and make a credible case that they should become a practical default for psychophysical scaling from triplet comparisons. The headline result—matching MLDS in the 1D monotonic case while also recovering non-monotonic functions—is supported by sane simulations: four scaling functions, four noise levels, five triplet fractions, and both MSE and triplet error as metrics. The slant-from-texture experiment is a nice real-data check, and the practical rules of thumb in Section 5 are genuinely useful for a psychophysics audience.\n\nThe soft spots are real but do not sink the 1D core. First, the abstract promises an R implementation that the preprint does not provide; Section 5.5 says it will come “upon acceptance of the paper.” For a paper whose pitch is “use this as your default tool,” that is an actionable reproducibility gap. Second, the best-of-10 restart selection uses training triplet error, which mildly favors methods with more capacity. Not fatal—the real experiments use cross-validation—but it deserves a line of caveat.\n\nThe bigger issue is the multi-dimensional part of the central claim. The stress-test note gets this right. In Section 3.3 the only 2D simulation builds ground truth by running NMDS on Ekman’s color data, and the authors explicitly say they do not argue the embedding is “correct.” Because NMDS and the ordinal embedding methods share the same Euclidean-distance model, that simulation shows only that t-STE can recover a Euclidean configuration from triplets—not that the recovered coordinates are the true perceptual scale. In the Eidolon experiment, the evidence for a second dimension is a drop in cross-validated triplet error from d=1 to d=2. But cross-validated triplet error measures consistency with the same triplet-response model the embedding was optimized to fit. It does not tell you that the two dimensions correspond to reach, grain, or coherence; the paper never checks against those known generation parameters, and it lists “interpreting the embedding” as an open issue. So the claim that ordinal embedding “can produce a desirable scaling function” in higher dimensions is underdetermined. The claim may be true, but the current evidence does not separate a valid perceptual scale from a flexible compression of the training signal.\n\nWho gets value? Psychophysicists and vision scientists who use MLDS or NMDS and want a modern alternative for non-monotonic or multi-dimensional cases. The 1D material is solid enough to serve as a practical tutorial. The multi-D claim should be either softened or backed by an external anchor, such as testing whether recovered dimensions align with known stimulus parameters.\n\nRecommendation: send it out. The 1D contribution is sound and the paper has enough new material to justify referee time. But the multi-D claim needs a revision that either qualifies it or adds external validation, and the R-implementation promise should be fixed.","headline":"Solid methods-transfer paper: the 1D monotonic and non-monotonic results hold up, but the multi-dimensional 'desirable scaling function' claim is under-validated and should be tempered or better anchored.","tokens_in":27630,"tokens_out":2815,"would_cite":true,"duration_ms":27540,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ordinal embedding from machine learning, applied to triplet comparisons, estimates perceptual scaling functions as accurately as MLDS in the standard one-dimensional monotonic case and also handles non-monotonic and multi-dimensional…","keywords":["perceptual scaling","ordinal embedding","triplet comparisons","method of triads","maximum likelihood difference scaling","stochastic triplet embedding","psychophysics","multi-dimensional scaling"],"falsifier":"Conduct a triplet experiment on a stimulus set known to violate the Euclidean metric (e.g., similarity judgments that fail the triangle inequality); if t-STE's cross-validated triplet error remains far above the repeat-answer baseline while a non-Euclidean model fits well, the Euclidean assumption is the load-bearing weakness.","tokens_in":26583,"feed_emoji":"👁️","tokens_out":9159,"duration_ms":79063,"temperature":0.7,"pith_summary":"The paper proposes that ordinal embedding methods from machine learning—especially t-STE—can serve as a general tool for estimating psychophysical scaling functions from triplet comparisons of the form 'is stimulus i more similar to j or to k?'. Across simulations and two real psychophysics experiments, the methods match the accuracy of maximum likelihood difference scaling (MLDS) when the perceptual scale is one-dimensional and monotonic. They also recover non-monotonic scaling functions and multi-dimensional perceptual spaces, cases MLDS cannot handle by construction, and they need only on the order of $d n \\log n$ triplet judgments instead of the full dissimilarity ordering required by NMDS. If the paper is right, psychophysics gains a default scaling algorithm that keeps the strengths of MLDS while dropping its monotonicity and one-dimensionality restrictions.","feed_headline":"Machine-learning triplets recover perceptual scales MLDS cannot","feed_subtitle":"Matches the psychophysics gold standard on simple scales, and extends to non-monotonic and multidimensional perception.","key_machinery":"The central object is the ordinal embedding problem: given triplet answers, find points $y_1,\\dots,y_n$ in $d$-dimensional Euclidean space whose distances are consistent with the answers. The workhorse is stochastic triplet embedding (STE), which models the probability that stimulus $i$ is judged closer to $j$ than to $k$ as $p_{ijk} = \\frac{\\exp(-\\lVert y_i - y_j \\rVert^2)}{\\exp(-\\lVert y_i - y_j \\rVert^2) + \\exp(-\\lVert y_i - y_k \\rVert^2)}$, then maximizes the likelihood of the observed answers; the t-STE variant replaces the Gaussian kernel with a heavy-tailed Student-t kernel for robustness to noise. This machinery carries the argument because it imposes neither monotonicity nor a one-dimensional target, and it is paired with the cross-validated triplet error as a ground-truth-free way to judge whether a recovered scale is any good.","core_discovery":"The discovery is that the classic scaling problem can be re-cast as ordinal embedding: treat the stimuli as abstract items, the triplet answers as ordinal constraints, and the recovered Euclidean coordinates as the perceptual scale. The paper shows that stochastic triplet embedding, in particular the t-distributed variant t-STE, produces scaling functions whose cross-validated triplet error is on par with MLDS in the one-dimensional monotonic setting, while MLDS fails on non-monotonic functions and cannot produce multi-dimensional scales at all. In a simulation built from color-similarity data, the embedding methods reconstruct the two-dimensional color circle from a fraction of the triplets, where NMDS loses the structure to noise. In a real slant-from-texture experiment with eight observers, the ordinal embedding methods match MLDS's triplet error and, unlike MLDS, reveal non-monotonic perceptual patterns in two observers. The paper presents this as evidence that ordinal embedding methods are promising default scaling algorithms.","pith_inferences":["Editorial: if the Euclidean model is wrong, a t-STE embedding could fit triplet answers well while misrepresenting the true geometry, so testing on known non-Euclidean similarity spaces would delimit when the method is trustworthy.","Editorial: because embeddings are only defined up to rotation, reflection, and scaling, comparing observers requires an alignment step; adding such an alignment could make the method a tool for studying inter-observer agreement in perceptual geometry.","Editorial: the $d\\,n\\log n$ sample bound invites adaptive triplet selection that targets comparisons near estimated equality boundaries, potentially reducing data requirements below random sampling.","Editorial: the same triplet-embedding formulation should transfer to conjoint measurement and multimodal stimuli, offering a way to map combined perceptual spaces without additivity or independence assumptions."],"forward_implications":["If the claim holds, t-STE can serve as a default scaling method in the standard one-dimensional monotonic setting, matching MLDS accuracy when enough triplets are collected.","Non-monotonic perceptual functions such as U-shaped or sinusoidal scales become estimable from triplet data, a case where MLDS systematically produces the wrong function.","Multi-dimensional perceptual spaces such as color or pitch become recoverable from triplet judgments, extending scaling to settings MLDS cannot address.","Because about $d\\,n\\log n$ randomly sampled triplets suffice, experiments are feasible where collecting a full dissimilarity ordering would be impossible.","The cross-validated triplet error supplies a ground-truth-free criterion for judging when an embedding is good enough, usable for MLDS and NMDS as well."],"supporting_citations":[{"why":"Proves the finite-sample bound that O(d n log n) triplets suffice for recovering an embedding up to small error, used to justify triplet budgets and rules of thumb.","marker":"(Jain et al., 2016)"},{"why":"Introduces stochastic triplet embedding (STE) and t-STE, the probabilistic model and algorithms that are the paper's main methods.","marker":"(Van Der Maaten and Weinberger, 2012)"},{"why":"Introduces MLDS, the parametric difference-scaling baseline the paper compares against in the one-dimensional monotonic case.","marker":"(Maloney and Yang, 2003)"},{"why":"Provides the MLDS R implementation used in the simulations and experiments.","marker":"(Knoblauch and Maloney, 2010)"},{"why":"Supplies the slant-from-texture triplet dataset used for the real one-dimensional scaling experiment.","marker":"(Aguilar et al., 2017)"},{"why":"Supplies the color similarity data whose NMDS embedding is used as ground truth for the two-dimensional simulation.","marker":"(Ekman, 1954)"},{"why":"Foundational NMDS stress formulation and algorithm used as a baseline and for constructing the ground-truth color circle.","marker":"(Kruskal, 1964b)"}],"fun_headline_variants":["Ordinal embedding recovers scales MLDS misses","Triplet embedding matches MLDS, goes beyond","Scaling without monotonicity: ordinal embedding wins","Multidimensional perceptual scales via ordinal embedding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes perceived dissimilarity is Euclidean and low-dimensional, so if the true perceptual geometry is not Euclidean the recovered embedding can be a best-fit artifact rather than the actual perception.","fun_headline_variants_meta":{"raw":{"variants":["Ordinal embedding recovers scales MLDS misses","Triplet embedding matches MLDS, goes beyond","Scaling without monotonicity: ordinal embedding wins","Multidimensional perceptual scales via ordinal embedding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000522,"raw_usage":{"total_tokens":2548,"prompt_tokens":989,"completion_tokens":1559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":1501}},"tokens_in":605,"tokens_out":1559,"duration_ms":13075,"temperature":1.0,"reasoning_tokens":1501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:52:26.614863+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a triplet experiment on a stimulus set known to violate the Euclidean metric (e.g., similarity judgments that fail the triangle inequality); if t-STE's cross-validated triplet error remains far above the repeat-answer baseline while a non-Euclidean model fits well, the Euclidean assumption is the load-bearing weakness.","supporting_citations":[{"cited_title":"G., and Nowak, R","cited_arxiv_id":null,"evidence_quote":"Proves the finite-sample bound that O(d n log n) triplets suffice for recovering an embedding up to small error, used to justify triplet budgets and rules of thumb."},{"cited_title":"and Weinberger, K","cited_arxiv_id":null,"evidence_quote":"Introduces stochastic triplet embedding (STE) and t-STE, the probabilistic model and algorithms that are the paper's main methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces MLDS, the parametric difference-scaling baseline the paper compares against in the one-dimensional monotonic case."},{"cited_title":"and Maloney, L","cited_arxiv_id":null,"evidence_quote":"Provides the MLDS R implementation used in the simulations and experiments."},{"cited_title":"A., and Maertens, M","cited_arxiv_id":null,"evidence_quote":"Supplies the slant-from-texture triplet dataset used for the real one-dimensional scaling experiment."},{"cited_title":"Dimensions of color vision","cited_arxiv_id":null,"evidence_quote":"Supplies the color similarity data whose NMDS embedding is used as ground truth for the two-dimensional simulation."}],"review_version":1}