{"id":"bf921ebf-420a-4dc6-8ef8-926394bb2df5","arxiv_id":"2505.01178","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Gaussian-process-based mixture model reveals that swap errors in visual working memory can depend on report-dimension (location) distance, pointing to a possible encoding-stage mechanism.","lead":"This paper introduces a Bayesian nonparametric model that lets the probability of a memory 'swap error' depend flexibly on any feature of the items being remembered. Fitted to a motion-location task, the model finds that swaps also depend on the spatial distance between items, suggesting some errors may start at encoding rather than at retrieval.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 7's parametric null assumes the probe-only model is correctly specified; Figure 2C shows spurious report-dimension modulation on probe-only data, so the RDK ΔBIC* could reflect misspecification rather than true report-dimension dependence.","rationale":"The reader's weakest assumption—that the Figure 7 null distribution requires the probe-only model to faithfully reproduce the real data's statistical structure—is exactly the load-bearing concern. I agree with this assessment. The central claim of a report-dimension swap dependence is established through a model-based null that is sensitive to any misspecification of the probe-only model. The paper's own text flags the issue: Figure 2C shows spurious report-dimension modulation in the both model on probe-only data, and the Mixed dependence section explicitly worries that this 'may cast doubt' on pruning. The BIC approximation is also explicitly heuristic and omits variational parameters, as stated in the Model description section. These are not external criticisms; they are limitations the authors acknowledge but do not resolve. The concern is internal validity, not disagreement with the field. The proposed non-parametric conditional-independence test would settle whether the report-dimension effect is real without relying on the GP/BIC null. Because this concern is addressable and the empirical claim may survive, the conditional verdict is appropriate; my read does not change the reader's verdict.","tokens_in":11999,"tokens_out":7281,"duration_ms":75209,"concrete_test":"Perform a model-free conditional-independence test on the direction-cued RDK dataset. For each trial, compute the probe-distance and report-distance of every distractor to the cued item. Stratify trials by binned probe distance (e.g., quartiles) and, within each stratum, fit a non-parametric regression of swap indicator on report distance (e.g., logistic GAM or isotonic regression). Test the report-distance term against a permutation null obtained by randomly permuting the report-feature values among distractors within each trial, preserving probe distances and stimulus sampling. If the report-distance effect is not significant within at least one probe-distance stratum after multiple-comparison correction, the BNS report-dimension claim is an artifact of the GP/BIC machinery. This test avoids the parametric null assumption and directly probes the reported dependence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—report-dimension dependence in the direction-cued RDK dataset—rests on the Figure 7 validation. The null distribution for ΔBIC*(both;probe) is generated by fitting the probe-only model to the real data, sampling synthetic datasets from that fitted model, and refitting candidate models to each synthetic dataset. This is a parametric bootstrap and is only valid if the probe-only model captures all statistical structure of the real data except the report-dimension dependence. The paper provides no evidence for this. On the contrary, Figure 2C shows that when the both model is fitted to probe-only synthetic data at realistic trial counts, it retains spurious modulation in the report dimension; the null distribution is meant to calibrate this, but it inherits any misspecification of the probe model. The probe model uses a restrictive Weinland kernel, pools all subjects into one swap function, approximates the posterior with SVGP, and evaluates BIC with a Monte Carlo marginal likelihood that omits variational parameters. Any of these could make the synthetic null fail to reproduce the real data's structure, inflating ΔBIC* even without a true report-dimension swap error. The test merely shows that real data differ from synthetic probe-only data in some way the both model captures; it does not isolate the report dimension as the cause. The paper itself acknowledges this risk in the Mixed dependence section: 'the existence of any modulation when fitting to 1D synthetic data may cast doubt on the ability of BNS to prune out dependence in the report dimension.' This is the load-bearing soft spot because the encoding-error interpretation in the Discussion depends on the report-dimension dependence being real.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BNS, a Bayesian non-parametric mixture model of swap errors in visual working memory. The response distribution is a mixture over remembered items, with component weights determined by a Gaussian-process swap function of the circular displacement between distractor and cue in the probe and/or report feature dimensions. Sparse variational GP inference is used, and model comparison is based on a per-subject BIC approximation. Fitted to several existing datasets, the model recovers probe-similarity dependence in orientation-cued and Schneegans & Bays data, and reveals a non-monotonic probe dependence together with a report-location dependence in an RDK direction-cued, location-report dataset. The authors interpret the report-dimension dependence as evidence for an encoding-stage error mechanism, challenging retrieval-based accounts of swap errors.","tokens_in":12390,"tokens_out":3655,"duration_ms":39390,"significance":"If the report-dimension dependence were firmly established, the paper would make an important empirical contribution by challenging the dominance of retrieval-based explanations of swap errors, and the BNS model itself would be a useful flexible descriptive tool. The clear generative specification and the synthetic-data recovery experiments in Figures 2 and 3 are genuine strengths. However, the central validation test for the report-dimension effect is fragile, and the current evidence is not sufficient to support the strong mechanistic conclusion drawn in the Discussion.","major_comments":[{"comment":"The null distribution for ΔBIC* is generated by fitting the probe-only model to the real data, sampling synthetic datasets from that fitted model, and refitting candidate models. This parametric bootstrap is only valid if the probe-only model captures all statistical structure of the real data except the report-dimension dependence. The paper itself shows in Figure 2C that the both model retains spurious report-dimension modulation when fitted to probe-only synthetic data at realistic trial counts, so the null distribution is expected to contain such artifacts. The test therefore demonstrates only that real data differ from synthetic probe-only data in some respect that the both model captures; it does not isolate the report dimension as the cause. A control using null data generated from a both model with the report-dimension component set to zero, or a posterior predictive check on the report-dimension marginal, is needed to support the claim.","section":"Results validation and model recovery; Figure 7 and Eq. (6)"},{"comment":"This conclusion overstates what the model can show. Earlier in the paper, in the section 'Relation to mechanisms underlying swap errors', the authors correctly note that 'we have not ruled out the possibility of a report dimension dependence also arising from retrieval error'. The modeling result establishes a statistical dependence in the report dimension, not a processing stage. The conclusion should be phrased as consistency with an encoding contribution rather than as a direct implication.","section":"Discussion ('this implies swap errors in this dataset were made due to an error in encoding at stimulus presentation…"},{"comment":"The disaggregated BIC omits the dimensionality of the variational parameters ψ and is described in the text as 'largely heuristic'. This is load-bearing because model recovery in Figure 3 shows that a true 'both' model is not recovered at realistic trial counts, which motivates the alternative test in Figure 7. The authors should assess the sensitivity of the Figure 7 conclusions to the BIC penalty, for example by also reporting cross-validated log-likelihood or a penalty based on the effective number of parameters.","section":"Model comparison; Eq. (4)"}],"minor_comments":[{"comment":"The denominator contains 'ee ˜πn', which appears to be a typo for e^{\\tilde{\\pi}_n}; similarly, 'over likelihood function' should read 'overall likelihood function'.","section":"Equation (3) and surrounding text"},{"comment":"The caption contains the typo 'miniumum' and the text has 'indiviaul'; these should be corrected.","section":"Figure 2 caption"},{"comment":"The term 'Bayesian non-parametric' is used although the GP prior has kernel hyperparameters; clarifying that 'non-parametric' refers to the swap function f, not to the absence of all parameters, would help readers.","section":"Abstract and terminology"},{"comment":"The sentence 'Figure 5C shows a similar degree as modulation for M = report as for probe' is grammatically unclear and should be reworded.","section":"Results, Mixed dependence"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper introduces a flexible Bayesian non-parametric model for swap errors that allows swap probability to depend on both probe and report feature distances. The key new empirical claim is a report-dimension dependence in a direction-cued, location-report RDK dataset, which would support encoding-stage misbinding rather than retrieval errors. The model itself is a real improvement: it relaxes strong parametric assumptions, the synthetic recovery experiments are thorough, and the non-monotonic probe dependence is a new observation. The central soft spot is the validation of the report-dimension claim. The Figure 7 test builds a parametric null from the probe-only model fit to the real data. That null is only valid if the probe-only model is correctly specified, and the paper's own Figure 2C shows that a both-model retains spurious report-dimension modulation on probe-only data. So the significant delta-BIC* could reflect model misspecification rather than a true report-dimension effect. The authors acknowledge this risk but still lean on the claim. Also, parameters are pooled across subjects, the BIC approximation omits variational parameters, and there is no code. None of these are fatal, but they warrant caution. The paper deserves serious review: the method is reusable and the validation strategy is nearly state-of-the-art. I would ask for code and a direct test of the null, ideally nonparametric or cross-validated. If the report-dimension effect survives that, it is a meaningful challenge to the retrieval account.","headline":"A flexible Bayesian non-parametric swap-error model that is well-built and honestly validated, but its central claim of report-dimension dependence rests on a self-referential null that deserves a cautious referee.","tokens_in":725,"tokens_out":1889,"would_cite":true,"duration_ms":32783,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Swap errors in a direction-cued, location-reported working-memory task depend on distance in the report dimension, implying that some misbindings happen at encoding rather than retrieval.","keywords":["visual working memory","swap errors","Bayesian nonparametric mixture model","Gaussian process","feature binding","delayed estimation","model comparison"],"falsifier":"A targeted experiment would vary the spatial separation between items in the direction-cued, location-report task while holding the distribution of direction distances fixed; if swap rates do not rise when locations are closer, the inferred report-dimension dependence is not a genuine behavioural effect. Alternatively, refitting the Figure 7 null distribution with synthetic data generated from a probe-only model that matches the real dataset's trial counts, set sizes, and stimulus layouts, and showing that ΔBIC* against the both model no longer falls outside the null, would falsify the encoding-error conclusion.","tokens_in":11818,"feed_emoji":"🧠","tokens_out":8263,"duration_ms":79947,"temperature":0.7,"pith_summary":"This paper aims to determine where swap errors in visual working memory originate—whether they come from failures to encode the right feature bindings, from noise in storage, or from errors at retrieval when the cue is shown. To separate these, it introduces a Bayesian non-parametric mixture model, BNS, that lets the probability of swapping to each distractor depend freely on its distance to the target in both the probed and the reported feature dimensions, with a Gaussian-process prior over that dependence. Fitting BNS to several published delayed-estimation datasets reproduces the well-known rise in swap errors with cue similarity. In one direction-cued, location-reported dataset, however, BNS finds an additional non-monotonic modulation by distance in the reported location dimension. Because the cue is unknown when the array is encoded, the paper concludes that these particular swaps reflect misbinding at stimulus presentation, challenging the prevailing retrieval-only explanation.","feed_headline":"Swap errors in memory can start at encoding, not retrieval","feed_subtitle":"A flexible Bayesian model shows direction-cued memory errors also depend on reported location, challenging retrieval-only explanations.","key_machinery":"The load-bearing object is the swap function f, a Gaussian-process function on the displacement vector between each distractor and the cued item, x = (x_p, x_r), whose evaluations at distractor displacements become the logits of the mixture components in a circular response model, with the function equipped with a Weinland (periodic) kernel. Choosing which coordinates enter x gives model classes M ∈ {none, probe, report, both}, so comparing models amounts to asking whether swapping depends on cue-feature proximity, report-feature proximity, neither, or both. A sparse variational Gaussian-process approximation (SVGP) makes the intractable posterior over f tractable, and the model's interpretable structure lets the inferred f be read directly as evidence about the mechanism: probe-only dependence points to retrieval failure, report-dimension dependence in a location-report task points to encoding failure.","core_discovery":"The paper's central discovery is a report-dimension dependence in swaps for a random-dot-motion direction-cued, location-reported task, in which the probability of swapping to a distractor increases when that distractor is spatially close to the target, with a sharp non-monotonic structure in the direction dimension as well. The authors show that a model conditioning swaps only on cue-feature distance cannot reproduce the real data's statistical structure, while a model conditioning on both probe and report distances can; the report-dimension modulation survives a null-distribution test built on BIC differences between real and synthetic data. They interpret this as evidence that some swap errors in that dataset are caused by an encoding error at stimulus presentation time—a misbinding of cue and report features before the cue is known—rather than by a failure to retrieve the correctly bound item.","pith_inferences":["A direct testable extension would vary spatial separation between items in the direction-cued, location-report task while holding direction-distance distributions fixed; if swap rates do not rise with spatial proximity, the inferred encoding-stage effect is not genuine.","Applied to other two-feature delayed-estimation tasks, BNS may uncover report-dimension dependencies that one-dimensional mixture models would misattribute to cueing errors, potentially reassigning some known swap-error effects to encoding.","Because the swap function is pooled across subjects, the report-dimension effect could be driven by a subset of participants or by low-coherence trials; subject-level or hierarchical fits with more data would test this.","The validity of the Figure 7 null distribution rests on the probe-only model being a faithful generative description whenever no true report-dimension effect exists; checking the null model's adequacy directly is needed before the encoding-error conclusion is fully secured."],"forward_implications":["Swap variability in the direction-cued RDK dataset is not exhausted by probe-feature similarity, so models built only on retrieval errors will leave systematic structure unexplained.","The inferred report-dimension dependence implies at least some swap errors originate before the cue is presented, during encoding of feature bindings.","The non-monotonic probe dependence for RDK directions is compatible with directions being stored as orientation-like representations; BNS discovered this without needing a predefined metric on the direction circle.","The BIC-based null-distribution test provides a reusable template for deciding when an added dependency in swap behaviour is real rather than a complexity artifact."],"supporting_citations":[{"why":"Provides the original three-component mixture model (target, uniform guesses, swap components) that BNS generalises and uses as a baseline.","marker":"Bays et al., 2009"},{"why":"Supplies the non-parametric approach to swap errors without a uniform component, which BNS extends to full feature dependencies.","marker":"P. M. Bays, 2016"},{"why":"Supplies the direction-cued RDK and orientation-cued ellipse datasets and the prior claim that swap errors are fully explained by cue-feature variability.","marker":"McMaster et al., 2022"},{"why":"Presents the neural population model attributing swap errors to retrieval failure at cuing time, the account the paper's encoding-error finding challenges.","marker":"Schneegans & Bays, 2017"},{"why":"Gives the Gaussian-process prior used for the swap function.","marker":"Rasmussen & Williams, 2005"},{"why":"Provides the sparse variational Gaussian-process inference and ELBO objective used to approximate the swap-function posterior.","marker":"Leibfried et al., 2020"},{"why":"Shown to be equivalent to the BNS 'none' model after marginalisation, linking BNS to existing mixture models used on neural data.","marker":"Alleman et al., 2024"}],"fun_headline_variants":["Encoding, not retrieval, drives some memory swap errors","Bayesian model finds swap errors depend on reported location","Direction-cued memory swaps also hinge on report features","Non-monotonic encoding effects appear in memory swaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the probe-only model, when fitted to real data and used to generate synthetic data, faithfully reproduces the real data's statistical structure whenever swap errors truly depend only on the probe dimension; if that generative null model is misspecified, the observed report-dimension effect could arise without any true report-dimension dependence.","fun_headline_variants_meta":{"raw":{"variants":["Encoding, not retrieval, drives some memory swap errors","Bayesian model finds swap errors depend on reported location","Direction-cued memory swaps also hinge on report features","Non-monotonic encoding effects appear in memory swaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000838,"raw_usage":{"total_tokens":3694,"prompt_tokens":1026,"completion_tokens":2668,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":2605}},"tokens_in":642,"tokens_out":2668,"duration_ms":19313,"temperature":1.0,"reasoning_tokens":2605,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:24:31.415716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A targeted experiment would vary the spatial separation between items in the direction-cued, location-report task while holding the distribution of direction distances fixed; if swap rates do not rise when locations are closer, the inferred report-dimension dependence is not a genuine behavioural effect. Alternatively, refitting the Figure 7 null distribution with synthetic data generated from a probe-only model that matches the real dataset's trial counts, set sizes, and stimulus layouts, and showing that ΔBIC* against the both model no longer falls outside the null, would falsify the encoding-error conclusion.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the original three-component mixture model (target, uniform guesses, swap components) that BNS generalises and uses as a baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the non-parametric approach to swap errors without a uniform component, which BNS extends to full feature dependencies."},{"cited_title":"M., Tomi ´c, I., Schneegans, S., & Bays, P","cited_arxiv_id":null,"evidence_quote":"Supplies the direction-cued RDK and orientation-cued ellipse datasets and the prior claim that swap errors are fully explained by cue-feature variability."},{"cited_title":"J., & Johnston, W","cited_arxiv_id":null,"evidence_quote":"Shown to be equivalent to the BNS 'none' model after marginalisation, linking BNS to existing mixture models used on neural data."}],"review_version":1}