{"id":"3431ce6a-ffeb-47b3-8a3c-0a2a0d0608b2","arxiv_id":"2411.13203","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"PAM couples Hierarchical Gaussian Filter beliefs to the Diffusion Decision Model, Lognormal Race Model, and Racing Diffusion Model, with parameter recovery simulations supporting its accuracy.","lead":"This paper introduces PAM, a modeling framework that connects Bayesian prediction (the Hierarchical Gaussian Filter) with three classic evidence-accumulation models of decision-making. It validates the framework with simulated parameter recovery and demonstrates a real-data tutorial, aiming to give cognitive scientists a practical tool for studying how expectations shape choices and response times.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LNR/RDM validation rests on Ter=0, an assumption likely violated in real data; nonzero non-decision time could bias recovered belief slopes, undercutting the claim of highly accurate recovery for all models.","rationale":"The reader's weakest assumption identified the Ter=0 default as load-bearing, and my independent read converges on the same point. The authors' own report of |ρ| > .9 between estimated Ter and all LNR parameters is a direct admission of identifiability trouble, and fixing Ter to zero in both simulations and the tutorial makes the headline recovery claim inapplicable to real RT data unless one assumes Ter is negligible. This is not a stylistic or presentational issue: it bears directly on whether the model can estimate the key belief-modulation slopes without systematic bias. I therefore see no reason to change the reader's CONDITIONAL verdict. The proposed simulation-based test is decisive: if the LNR/RDM recovery remains accurate with Ter = 150 ms, the concern is resolved; if not, the main text must qualify the 'all scenarios' claim and provide a sensitivity analysis or a default that estimates Ter with appropriate priors.","tokens_in":26703,"tokens_out":5045,"duration_ms":55334,"concrete_test":"Simulate 100 participants per LNR/RDM scenario (Tables 5–7) with Ter = 0.15 s (matching the DDM simulations), then fit with the tutorial defaults (Ter fixed to 0). Compare median recovered belief slopes b/bv and intercepts to the true values. If the absolute median deviation for b or bv exceeds the paper's reported maxima (15% for LNR b, 13% for RDM slopes) in more than a small fraction of scenarios, the default Ter=0 confounds the central recovery claim for realistic non-decision times.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'In all scenarios and with all the proposed decision models, the results showed highly accurate parameter recovery' (Final Remarks)—is conditional on an assumption the paper itself shows to be problematic. For the LNR and RDM, the Model fitting section fixes Ter to 0 because, as stated in Parameter Recovery (LNR), 'the estimated Ter parameter was highly correlated with all the LNR parameters (|ρ| > .9).' Real response times necessarily include encoding and motor components (Ter > 0), as the DDM simulations acknowledge by adding Ter = 0.15 s. If Ter is nonzero in real data, an LNR/RDM that omits it will absorb the constant shift into the lognormal location parameters (θ, or thresholds/drifts in the RDM), and because HGF belief trajectories are correlated with the trial sequence, the belief-modulation slopes b/bv can be biased rather than merely shifted. The tutorial then sets Ter = 0 as the default for LNR and RDM, so the framework's quantitative estimates of predictive effects from real data inherit this unvalidated assumption. The main text only points to Supplementary Information 1 for Ter inclusion; no main-text sensitivity analysis is reported. The validation therefore does not establish that the framework recovers belief-modulation parameters accurately under conditions that are standard in the EAM literature.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces PAM, a computational framework that couples a Hierarchical Gaussian Filter perceptual model with three evidence accumulation models (DDM, LNR, RDM), using trial-by-trial beliefs and belief precision to linearly modulate EAM parameters. The paper validates PAM through parameter recovery simulations across different accuracy and response-time scenarios, reports median recovery and confidence intervals for each model, and provides a step-by-step MATLAB tutorial applied to a random-dot kinematogram dataset, including Bayesian model selection among model variants. The central claim is that the framework yields highly accurate parameter recovery in all tested scenarios and with all three decision models, while remaining computationally efficient.","tokens_in":27035,"tokens_out":5217,"duration_ms":52420,"significance":"If the validation claims hold, PAM is a useful and timely contribution: it provides a concrete, open-source implementation bridging predictive-coding style belief models and standard EAMs, with a tutorial that lowers the barrier to adoption. The paper's strengths include the availability of code and data, the use of standard parameter-recovery methodology, and the demonstration of computational efficiency. However, the validation is an internal-consistency check (simulate from the same model, then recover), not a test against independent data, and the headline claim of highly accurate recovery is conditional on a strong auxiliary assumption about non-decision time for the LNR and RDM. The empirical tutorial is illustrative rather than confirmatory, which the authors acknowledge, but the main-text framing of the simulation results should be reconciled with the reported recovery accuracy.","major_comments":[{"comment":"The central claim that 'in all scenarios and with all the proposed decision models, the results showed highly accurate parameter recovery' is not supported for the LNR and RDM because Ter is fixed to zero in both the simulations and the tutorial default, and the manuscript itself reports that estimating Ter leads to |ρ| > .9 correlations with LNR parameters. Real response times include non-decisional encoding and motor components, as the DDM simulations acknowledge by adding Ter = 0.15 s. If real data contain nonzero Ter, an LNR or RDM that omits it can absorb the shift into location parameters, and because belief trajectories are correlated with the trial sequence, the belief-modulation slopes b and bv can be biased rather than merely shifted. The main text defers the nonzero-Ter analysis to Supplementary Information 1 but does not report its results; this is load-bearing for the quantitative claims of the framework. I request a main-text sensitivity analysis with realistic nonzero Ter values, or a clear restriction of the validation claim to the Ter = 0 case.","section":"Parameter Recovery (LNR and RDM); Model fitting; Tutorial"},{"comment":"The reported recovery for the LNR slope parameter b contradicts the text's statement that recovery 'deviated by at most 15%.' For example, in the first row of Table 5, the simulated value is b = -1.04 and the median recovered value is listed as -0.63, a deviation of roughly 40% of the true value. Several subsequent rows appear misaligned (e.g., the second row lists an estimated a of 0.28 against a simulated a of -0.53, with what look like true values displaced into parentheses). This makes the LNR recovery results unreliable as reported and undermines the accuracy claim for that model. Please re-run or re-present the table with clearly aligned columns and report the actual deviation values.","section":"Results – LNR, Table 5"},{"comment":"The paper simultaneously claims highly accurate recovery and reports that the perceptual parameter ω2 deviates by up to 28% in the reduced model and 14% in the full model, with wide interquartile ranges (e.g., Table 4, first row, reduced model: median -3.14, IQR 1.17, for true ω2 = -4). Because ω2 drives the belief trajectories that modulate the EAM parameters, this imprecision is directly relevant to the joint-model recovery claim. The final remarks should either temper the 'highly accurate' wording to reflect the observed accuracy of the slope and perceptual parameters or provide additional diagnostics (e.g., bias, correlation, or root mean squared error) that justify the characterization.","section":"Results – DDM, paragraph on omega2 recovery"}],"minor_comments":[{"comment":"The sigmoid function s() used in the precision modulation of the boundary separation is not defined in the text; please define it explicitly (e.g., s(x) = 1/(1+exp(-x))).","section":"Model Specifications, Eq. (4)"},{"comment":"Table 7's simulated-parameter row appears to contain four values while the header lists three (av, bval, bv), and the first value '-4' is inconsistent with the recovered aa shown as 2.00; please check the column alignment and the table caption (the caption says 'boundary (v)' but v is the drift rate, not the boundary).","section":"Parameter Recovery, Table 7"},{"comment":"The text states that 'on each dataset (excluding LNR), we carried out two analyses using full and reduced model configurations'; please clarify why the LNR was excluded from the full/reduced comparison and whether full-model LNR results exist in the supplement.","section":"Model fitting"},{"comment":"In the sentence 'here we consider the DDR, LNR, and RDM', 'DDR' should be 'DDM'.","section":"Introduction"},{"comment":"The sentence 'The HGF, it is a generic computational model' contains a grammatical error; please remove 'it'.","section":"Perceptual model"}],"recommendation":"major_revision","confidential_remarks":"The main risk to the paper's central claim is the Ter = 0 assumption for LNR and RDM. If the supplementary analysis with nonzero Ter convincingly shows that belief-modulation slopes remain unbiased, the revision could be minor; otherwise the validation claim must be substantially qualified. I also recommend checking the integrity of Tables 5 and 7 in the production version, as the current rendering is internally inconsistent. The paper is within the journal's scope and the open-code/tutorial contributions are genuinely useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a genuinely useful integration: for each trial, HGF belief trajectories linearly modulate the parameters of three standard EAMs, and the authors ship working MATLAB code plus a tutorial on real RDK data. The modeling field has been circling this combination for a while, but I don't know of a public toolbox that does it this cleanly. That alone makes the manuscript worth taking seriously.\n\nWhat is actually new is the specific set of modulation equations for DDM, LNR, and RDM, the systematic parameter recovery across fast/slow and high/low accuracy scenarios, and the open implementation. The recovery design is standard internal consistency checking, and the DDM results look good: intercepts recover within a couple of percent, Ter recovers to within 10 ms, and the false-positive rates on zero slopes look acceptable. The paper is also honest about the places where recovery is worse, though the prose sometimes oversells it: the slope parameters, especially DDM bv and LNR b, deviate by 10-15% with wide interquartile ranges, and that deserves clearer wording.\n\nThe main soft spot is the Ter=0 assumption for LNR and RDM. The paper reports that Ter correlates with all LNR parameters at |r| > .9, so they fix it to zero and point to Supplementary Information 1 for the full treatment. That is an honest acknowledgment, but the central validation claim—highly accurate recovery 'in all scenarios and with all the proposed decision models'—is conditional on Ter=0. Real speeded responses include encoding and motor time; the DDM simulations themselves use 150 ms. If a constant shift that large is present in real data, an LNR or RDM without a shift parameter will absorb it into location or threshold parameters, and because beliefs are correlated with the block structure, the belief slopes can be biased. A sensitivity analysis in the main text, simulating data with Ter=150 ms and refitting with Ter=0, would show whether this matters. Without it, the quantitative estimates from the tutorial are on shaky ground.\n\nTwo secondary issues. The tutorial's Bayesian model selection compares PAM variants but includes no standard EAM without the HGF component, so it cannot tell readers whether the predictive machinery actually improves explanation. And the tables in the version I saw are badly formatted—the LNR table in particular is garbled—which is minor but needs fixing before publication.\n\nOverall, this is a useful, well-scoped methods contribution for computational cognitive modelers who want to connect hierarchical Bayesian learning to sequential sampling. It is not circular, and the code and data are real. I would send it to peer review, with the Ter sensitivity analysis as a required revision.","headline":"Useful modular toolbox joining HGF to EAMs, with open code and honest recovery checks; the LNR/RDM validation rides on a Ter=0 assumption that real data will violate.","tokens_in":27546,"tokens_out":3531,"would_cite":true,"duration_ms":38827,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PAM couples Bayesian beliefs to decision models and recovers parameters accurately.","keywords":["Evidence accumulation models","Predictive processing","Hierarchical Gaussian Filter","Drift diffusion model","Parameter recovery","Bayesian inference","Decision-making","Computational modeling"],"falsifier":"Simulate LNR and RDM datasets with a known nonzero non-decision time, for example 150 ms, and fit them with the framework's default Ter = 0; if the recovered belief slopes or drift and boundary intercepts shift systematically away from their true values, the central recovery claim fails for realistic data.","tokens_in":26500,"feed_emoji":"🧠","tokens_out":8769,"duration_ms":80147,"temperature":0.7,"pith_summary":"PAM is a computational framework for joining predictive, Bayesian accounts of learning with standard evidence-accumulation models of speeded choice. The paper's target is to show that the two modeling traditions can be coupled in a way that is both valid and practical: the trial-by-trial beliefs produced by a Hierarchical Gaussian Filter are used to modulate decision parameters of the Drift Diffusion Model, the Lognormal Race Model, and the Racing Diffusion Model. Parameter recovery simulations across fast and slow, easy and hard scenarios show highly accurate recovery of both learning and decision parameters. A tutorial with real random-dot-motion data illustrates the workflow and shows that models in which prior belief modulates drift rate outperform models in which it modulates starting point or boundary. If correct, PAM gives researchers a tool to estimate how predictions shape decisions from choice and response-time data alone.","feed_headline":"Belief-coupled decision models recover their parameters accurately","feed_subtitle":"The PAM framework estimates Bayesian learning and decision-making parameters jointly from choices and response times.","key_machinery":"The load-bearing object is the trial-wise coupling equation between the perceptual model's inferred states and the EAM's parameters: predicted belief (and its precision) enters each decision parameter through centered linear terms, e.g. $v(t) = a_v + b_v(\\hat{\\mu}(t) - 0.5)$ for drift, with analogous expressions for start point, boundary, lognormal means, and racing-diffusion thresholds. The Hierarchical Gaussian Filter supplies the trial-wise belief trajectory and precision from the stimulus sequence; the EAM supplies the likelihood of the observed response time and choice. These pieces are optimized jointly with a quasi-Newton maximum-a-posteriori routine, which is what lets the framework estimate learning and decision parameters at the same time.","core_discovery":"The central claim is that predictive processes can be integrated into evidence accumulation models without losing identifiability. In PAM, predicted belief about the stimulus, $\\hat{\\mu}(t)$, and the precision of that belief are inserted as linear modulators of the decision model's parameters: for the DDM they shift start point, boundary, and drift; for the LNR they shift the lognormal means of the two accumulators; for the RDM they shift both accumulator thresholds and drift rates. The perceptual and decision parameters are then estimated jointly by maximum a posteriori optimization. The paper reports that in every simulated scenario and with every decision model the recovery was highly accurate, with the belief-modulation slopes being the least precise but still within confidence intervals, and that the winning model on real data modulated drift rather than starting point.","pith_inferences":["Beyond the paper's claims, if the recovery results generalize, PAM could be used to re-examine existing expectation effects in diffusion-model studies, asking whether effects previously attributed to start-point shifts are better described as drift-rate modulations.","A testable extension the authors do not report is whether fixing non-decision time at zero for the LNR and RDM biases recovered belief slopes when real participants have non-negligible encoding or motor time; simulating nonzero non-decision time would settle this.","The centered linear coupling could be replaced by nonlinear or precision-weighted forms, and Bayesian model comparison across such variants would show whether the linear assumption is a real constraint or a harmless convenience."],"forward_implications":["Researchers can estimate a participant's Bayesian learning rate and decision parameters in a single fit from choice and response-time data.","The framework makes it possible to test competing hypotheses about where predictions act, since models that modulate start point, boundary, or drift can be compared with Bayesian model selection.","Because recovery was accurate across fast and slow, high- and low-accuracy scenarios, PAM can be applied to a wide range of two-choice speeded decision tasks.","The modular setup extends beyond the three demonstrated EAMs and beyond the Hierarchical Gaussian Filter, with other perceptual models already supported."],"supporting_citations":[{"why":"Supplies the Hierarchical Gaussian Filter whose trial-wise belief trajectories drive the decision model.","marker":"Mathys et al., 2011, 2014"},{"why":"Establishes the observing-the-observer meta-Bayesian approach that PAM implements.","marker":"Daunizeau et al., 2010"},{"why":"Defines the diffusion decision model used as one of PAM's decision components.","marker":"Ratcliff, 1978"},{"why":"Defines the lognormal race model adapted as a PAM decision component.","marker":"Rouder et al., 2015"},{"why":"Defines the racing diffusion model adapted as a PAM decision component.","marker":"Tillman et al., 2020"},{"why":"Provides the fast first-passage-time density computation used for the DDM likelihood.","marker":"Navarro & Fuss, 2009"},{"why":"Provides the HGF toolbox fitting procedure and priors on which PAM's joint optimization builds.","marker":"Frässle et al., 2021"},{"why":"Motivates the parameter-recovery validation design used to assess PAM.","marker":"Wilson & Collins, 2019"}],"fun_headline_variants":["Predictive processes now fit into evidence accumulation models","New PAM framework links Bayesian inference to choice models","Decision models get a predictive upgrade with PAM","PAM bridges predictive brain and evidence accumulation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recovery results for the LNR and RDM assume that non-decision time can be fixed at zero without bias, even though the model's own estimates show non-decision time highly correlated with LNR parameters, so a real participant's meaningful encoding or motor time could confound the recovered belief effects.","fun_headline_variants_meta":{"raw":{"variants":["Predictive processes now fit into evidence accumulation models","New PAM framework links Bayesian inference to choice models","Decision models get a predictive upgrade with PAM","PAM bridges predictive brain and evidence accumulation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2752,"prompt_tokens":879,"completion_tokens":1873,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1814}},"tokens_in":495,"tokens_out":1873,"duration_ms":13229,"temperature":1.0,"reasoning_tokens":1814,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:42:24.603690+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate LNR and RDM datasets with a known nonzero non-decision time, for example 150 ms, and fit them with the framework's default Ter = 0; if the recovered belief slopes or drift and boundary intercepts shift systematically away from their true values, the central recovery claim fails for realistic data.","supporting_citations":[],"review_version":1}