{"id":"eb8d877f-8341-4e4b-896e-8dd3192c2aca","arxiv_id":"2607.07574","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"A frozen LSTM backbone with FiLM-based few-shot context adaptation estimates tip-level contact forces for deformable swabbing tools across nine surface-tool regimes using only wrist-mounted proprioception.","lead":"This paper trains a compact LSTM to estimate contact forces at the tip of a deformable swabbing tool using only wrist-mounted sensors, then adapts it to new surfaces with a few-shot context vector. It matters because disposable, sterile robotic swabbing tools cannot easily embed sensors at the contact point, and this approach offers a practical path to force-aware control without that hardware.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"§V-D.1 says few-shot adaptation updates \"context vector and readout,\" but §IV-B and Eq. 4 state the regression head ϕ is permanently frozen. If ϕ is co-adapted, the parameter-isolation claim is overstated.","rationale":"The reader correctly identified that the adaptation mechanism is under-specified and that the dimensionality/role of z_k is unclear. However, the reader focused on representational capacity of z_k (d_z dimensionality, 5-trial sufficiency), which is a speculative generalization concern. The more immediately load-bearing issue is the direct textual contradiction between §IV-B (\"regression head parameters ϕ are permanently frozen\") and §V-D.1 (\"updating only context vector and readout\"). This is not speculative — it is a concrete discrepancy in the paper's own description of its method versus its experiments. If the regression head is being co-adapted, the parameter-isolation claim — which is the paper's central contribution (contribution ii) — is overstated, and the improvements may be partly attributable to readout re-fitting rather than context conditioning. That said, the verdict remains CONDITIONAL for the same reasons the reader gave: the experimental design is solid (9 regimes, 5 seeds, ablations, forgetting check), and this is a clarification/verification issue rather than a fundamental flaw. If the authors confirm that only z_k was optimized (matching Eq. 4) and correct the §V-D.1 wording, the paper moves toward ACCEPT. If the regression head was co-adapted, the contribution needs to be re-scoped. Either way, the verdict stays CONDITIONAL — the concern sharpens the specific condition but does not change the overall assessment. The reader's other points (ADC units, cherry-picked 63%, no code release) are valid but secondary.","tokens_in":12307,"tokens_out":3123,"duration_ms":169885,"concrete_test":"Re-run the few-shot adaptation on all nine regimes optimizing strictly only z_k (as specified in Eq. 4, with θ_enc, FiLM projectors γ/β, and regression head ϕ all frozen at pretrained values). Compare the resulting RMSE to Table IV's few-shot column. If the values match within seed-level standard deviation, the §V-D.1 wording is purely terminological and the parameter-isolation claim holds. If RMSE degrades by more than ~10% in any regime, the regression head co-adaptation is contributing substantially and the central claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on parameter-isolated adaptation: only the low-dimensional context vector z_k is optimized while the LSTM backbone, FiLM projectors, and regression head remain frozen (§IV-B, Eq. 4). The framework explicitly states: \"the recurrent encoder parameters θ_enc, FiLM projector parameters, and regression head parameters ϕ are permanently frozen.\" Eq. 4 optimizes only z_k. However, the experimental description in §V-D.1 says adaptation involves \"5 trials updating only context vector and readout.\" This is a direct contradiction. The regression head g_ϕ is a two-layer MLP (128→32→1, roughly 4,100+ parameters). If g_ϕ is co-adapted during few-shot adaptation, then: (1) the adapted parameter set is far larger than d_z, undermining the \"parameter-isolated\" and \"low-dimensional conditioning\" narrative; (2) a substantial portion of the 18–63% RMSE reduction may come from readout re-fitting rather than context-vector modulation; (3) the analogy to VPT/L2P (which freeze everything except prompts) becomes inaccurate. The paper never clarifies whether \"readout\" in §V-D.1 refers to the FiLM modulation output (loosely, the \"readout\" of z_k) or to the regression head g_ϕ. If the former, the language is misleading but the claim holds; if the latter, the core contribution of minimal parameter adaptation is weakened because the effective adapted parameter count is an order of magnitude larger than implied.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The paper presents a context-modulated LSTM framework for estimating tip-level contact forces during robotic swabbing with deformable tools. The core idea is to train a recurrent backbone on a baseline surface and then adapt to new surface/tool combinations by optimizing only a low-dimensional context vector via FiLM modulation, while keeping the backbone, FiLM projectors, and regression head frozen. Experiments span nine interaction regimes (three tools × three surfaces) with five-seed evaluation, trial-level data splits, input modality ablations, and a forgetting check. The approach is motivated by a real practical need: wrist-mounted force/torque sensing is decoupled from true contact forces by viscoelastic tool dynamics, and tool-integrated sensors are impractical for sterile/disposable swabbing.","tokens_in":13199,"tokens_out":1138,"duration_ms":178107,"significance":"The problem is well-motivated and the experimental design is thorough relative to the scope: nine interaction regimes, five seeds, trial-level splits, modality ablation, and an explicit catastrophic-forgetting check. The parameter-isolated adaptation idea (FiLM-modulated context vector, frozen backbone) is a reasonable application of prompt-tuning-style ideas to robotic force estimation. The FSR calibration procedure using a NIST-traceable texture analyzer and the sub-millisecond latency measurement add practical value. However, a central inconsistency in the experimental description (see Major Comment 1) must be resolved before the contribution can be properly assessed.","major_comments":[{"comment":"§V-D.1 states that few-shot adaptation involves '5 trials updating only context vector and readout,' but §IV-B and Eq. (4) explicitly state that the regression head g_ϕ is permanently frozen and only z_k is optimized. This is a direct contradiction. If g_ϕ (a two-layer MLP, 128→32→1, ~4,100+ parameters) is co-adapted during adaptation, the parameter-isolation claim is substantially weakened and the reported improvements may partly reflect readout re-fitting rather than context-vector modulation. The authors must clarify whether 'readout' in §V-D.1 refers loosely to the FiLM modulation output or to g_ϕ itself. If g_ϕ is indeed frozen (consistent with Eq. 4), the wording in §V-D.1 should be corrected. If g_ϕ is co-adapted, the central claim of minimal-parameter adaptation needs to be re-scoped accordingly.","section":null},{"comment":"The dimensionality d_z of the context vector z_k is never reported, despite being a load-bearing parameter for the 'low-dimensional conditioning' claim (§IV-B, Eq. 4). The paper states d_z ≪ d_h = 128 but does not give the actual value. Without knowing d_z, the reader cannot assess whether the adapted parameter count is genuinely small relative to the backbone, nor whether the analogy to VPT/L2P is apt. Please report d_z and justify the choice.","section":null},{"comment":"Table IV and the abstract headline the 'up to 63%' improvement figure, which is cherry-picked from the best-performing regime (Viscoelastic Composite, Standard tool: 61.8%; Hard tool: 62.2%). The improvement range across all nine regimes is 17.7–62.2%. The abstract and conclusion should report the full range rather than only the maximum, to give a representative picture.","section":null}],"minor_comments":[{"comment":"The abstract and conclusion should report the full improvement range (17.7–62.2%) rather than only 'up to 63%' to give a representative picture.","section":null},{"comment":"Table II shows RMSE increasing from 43.22 (10% data) to 50.23 (15% data) before decreasing. The text attributes this to insufficient variance in the initial 5 trials, but this explanation is not tested. A brief note on whether this pattern is stable across seeds would help.","section":null},{"comment":"§IV-B: The context bank Z = {z_0, z_1, ..., z_K} is mentioned but it is unclear how K is determined or how many domains were used in total. Please clarify.","section":null},{"comment":"Fig. 3 caption mentions 'N=14' but the figure itself is not clearly labeled with axis units. Please add explicit axis labels (ADC units for the x-axis, count/density for the y-axis).","section":null},{"comment":"The Acknowledgment section mentions use of Grammarly and Adobe Photoshop AI for figure editing. This is fine but should be verified against the journal's policies on AI tool disclosure.","section":null},{"comment":"Reference [28] cites a YouTube video from 'ABC Research Laboratories' for environmental swabbing protocol. A more authoritative or peer-reviewed source for the swabbing protocol would strengthen the methodology.","section":null},{"comment":"§III: The temporal window length T is mentioned as a model parameter but its value is not reported in the main text (only that the model captures 'up to 2 s' of history). Please state T explicitly.","section":null}],"recommendation":"major_revision","confidential_remarks":"The §V-D.1 vs. §IV-B contradiction is the key issue. If the authors confirm that only z_k is adapted (consistent with Eq. 4) and the word 'readout' was a loose description of the FiLM output, this could drop to minor revision. If g_ϕ was actually co-adapted, the experiments may need to be re-run with g_ϕ truly frozen to support the parameter-isolation claim. I recommend asking the authors to clarify this point first before requesting additional experiments."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful reading and for identifying a genuine inconsistency in the experimental description. All three major comments are well-taken. The contradiction regarding whether the regression head g_ϕ is co-adapted during few-shot adaptation arises from imprecise wording in §V-D.1; the actual experimental protocol is consistent with Eq. (4)—only z_k is optimized, and g_ϕ remains frozen. We will correct the wording accordingly. We will also report d_z explicitly and revise the abstract/conclusion to present the full improvement range rather than only the maximum.","responses":[{"response":"The referee is correct that this is a contradiction, and we acknowledge it as an error in the manuscript text. The actual experimental protocol is consistent with §IV-B and Eq. (4): during few-shot adaptation, only the context vector z_k is optimized, while the LSTM backbone θ_enc, the FiLM projectors γ(·) and β(·), and the regression head g_ϕ are all permanently frozen. The phrase 'context vector and readout' in §V-D.1 was intended to refer loosely to the FiLM-modulated output that is then passed through the frozen readout, but this wording is imprecise and misleading. We will revise §V-D.1 to state unambiguously that '5 trials updating only the context vector z_k, with all other parameters—including the regression head g_ϕ—frozen.' This correction makes the experimental description fully consistent with Eq. (4) and the parameter-isolation claim. No re-scoping of the central claim is needed because the experiments were conducted as described in §IV-B; only the prose in §V-D.1 was erroneous.","revision_made":"yes","referee_comment":"§V-D.1 states that few-shot adaptation involves '5 trials updating only context vector and readout,' but §IV-B and Eq. (4) explicitly state that the regression head g_ϕ is permanently frozen and only z_k is optimized. This is a direct contradiction."},{"response":"We agree that d_z should be explicitly reported. In our experiments, d_z = 8, yielding a total of 8 adapted parameters per interaction regime (the FiLM projectors γ and β are frozen after pretraining, so they are not part of the adapted parameter count). This represents approximately 0.013% of the backbone parameters (60,545). We chose d_z = 8 based on preliminary experiments showing that values in the range 4–16 performed comparably, while d_z = 8 provided a good balance between expressiveness and minimal adaptation footprint. We will add this value and justification to §IV-B, and will also include the explicit adapted parameter count in the text so readers can directly assess the parameter-isolation claim relative to the backbone.","revision_made":"yes","referee_comment":"The dimensionality d_z of the context vector z_k is never reported, despite being a load-bearing parameter for the 'low-dimensional conditioning' claim."},{"response":"This is a fair point. Reporting only the maximum improvement does not give a representative picture of performance across the nine regimes. We will revise the abstract to state 'reducing zero-shot estimation error by 17.7–62.2% across nine interaction regimes' and will make a corresponding change in the conclusion. The body text in §V-D already reports the full range ('18–63%'), but we will ensure consistency by using the precise range throughout.","revision_made":"yes","referee_comment":"Table IV and the abstract headline the 'up to 63%' improvement figure, which is cherry-picked from the best-performing regime. The improvement range across all nine regimes is 17.7–62.2%. The abstract and conclusion should report the full range rather than only the maximum."}],"tokens_in":12137,"tokens_out":813,"duration_ms":119566,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Quick take on this paper: it applies parameter-isolated adaptation (FiLM-modulated context vectors on a frozen LSTM backbone) to contact force estimation in robotic swabbing with deformable tools. The experiments are well-designed — 9 surface-tool regimes, 5 seeds, trial-level splits, ablations, forgetting check. The core idea is sound and the application is real: environmental swabbing needs force control, disposable tools can't carry embedded sensors, and wrist-mounted sensing is decoupled from tip forces by viscoelastic hysteresis. The paper earns credit for a clean experimental matrix and honest reporting of per-regime improvements (17.7% to 63%, not just the headline number). The forgetting check showing baseline performance is preserved after all adaptations is the right experiment to run. The ablation confirming wrist F/T measurements are essential (removing them doubles error) is useful for anyone building similar systems. The LSTM comparison against Transformer and TCN is fair, with a Wilcoxon test on the key comparison. Now the soft spots. The stress-test flag about a contradiction between §IV-B (which says the regression head ϕ is permanently frozen and only z_k is optimized) and §V-D.1 (which says adaptation involves 'context vector and readout') is a real concern. Eq. 4 clearly optimizes only z_k. The regression head g_ϕ is a two-layer MLP (~4,100 parameters). If 'readout' in §V-D.1 means g_ϕ is co-adapted, the parameter-isolation claim is overstated and a chunk of the improvement may come from readout re-fitting rather than context modulation. If 'readout' loosely refers to the FiLM modulation output, the claim holds but the language is sloppy. The authors need to clarify this — it's a one-sentence fix but it matters for the central contribution. The reader's other points are valid but minor: d_z is never reported (reproducibility gap), RMSE is in ADC units with only an approximate Newton conversion (obscures practical significance), and no code or data is released. The 'up to 63%' framing is cherry-picked but the full range is reported in the table, so it's a presentation issue not a substance issue. The claim that z_k captures 'effective material interaction properties' is untested — no analysis of what the embedding encodes — but this is an aspirational interpretation, not a load-bearing claim. The experimental results stand on their own without it. This is a competent systems paper with a genuine application and above-average experimental rigor for the subfield. The ambiguity about what gets adapted during few-shot learning is the one thing that needs resolution before publication. Recommend serious peer review — a good referee should ask for the d_z value, the ADC-to-Newton conversion done properly, code release, and most importantly a clear statement of exactly which parameters are updated during adaptation.","headline":"Parameter-isolated FiLM adaptation for force estimation in deformable tool swabbing — solid experiments, one real ambiguity about what gets adapted","tokens_in":13346,"tokens_out":668,"would_cite":false,"duration_ms":86147,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Five swabs teach a robot the feel of a new surface","keywords":[],"falsifier":"A surface or tool whose viscoelastic response cannot be captured by the context vector's modulation of the frozen backbone, such that five-trial adaptation fails to reduce error below the zero-shot baseline, would falsify the claim that the backbone encodes transferable dynamics and that the context vector is a sufficient adaptation mechanism.","tokens_in":12366,"feed_emoji":"","tokens_out":1181,"duration_ms":173737,"temperature":0.7,"pith_summary":"The paper addresses a practical problem: when a robot uses a soft, deformable tool (like a sponge swab) to interact with surfaces, the forces measured at the robot's wrist are a distorted, delayed version of the forces actually applied at the tool tip due to viscoelastic hysteresis in the tool material. The authors propose that this force-estimation problem can be split into two parts: (1) a general recurrent dynamics model (an LSTM) that learns the history-dependent deformation behavior of the tool, and (2) a small, surface-specific context vector that adjusts the model's internal representation to match the material properties (stiffness, friction, damping) of whatever surface or tool is currently in use. The LSTM backbone is frozen after initial training; adaptation to a new surface requires optimizing only the low-dimensional context vector from as few as five swabbing trials. This parameter isolation means that learning a new surface does not degrade performance on previously learned surfaces, eliminating catastrophic forgetting. The core mechanism is feature-wise linear modulation (FiLM), which applies a learned scale and shift to the LSTM's hidden state conditioned on the context vector, effectively re-tuning the deformation model without retraining it.","feed_headline":"","feed_subtitle":"","key_machinery":"Frozen LSTM backbone (two layers, 64 hidden units each) + FiLM-modulated context vector z_k (low-dimensional, dimension d_z not specified) + lightweight MLP projectors mapping z_k to scale (gamma) and bias (beta) vectors that element-wise modulate the LSTM hidden state h_t. The regression head is also frozen after pretraining. Only z_k is optimized during adaptation.","core_discovery":"The central finding is that a frozen LSTM backbone, when modulated by a surface-specific context vector learned from only five trials, can recover accurate contact-force estimation across nine distinct tool-surface interaction regimes (three tools of varying stiffness crossed with three surfaces of varying material properties), reducing zero-shot estimation error by 18-63% depending on the regime, while the baseline domain's performance remains statistically unchanged after all adaptation cycles. The largest improvements occur on the viscoelastic composite surface, where damping-induced phase lag and hysteresis cause the most severe zero-shot degradation (error increases over 200% relative),","pith_inferences":["The paper does not report the dimensionality d_z of the context vector or analyze what z_k actually encodes. If d_z is very small (e.g., 2-4 dimensions), the context vector may be capturing only a coarse stiffness scalar rather than the full viscoelastic parameter space, which would limit generalization to surfaces with novel combinations of stiffness, damping, and friction.","The 5-trial support set may be sufficient for the nine tested regimes but may not be a lower bound. Surfaces with more complex viscoelastic behavior (e.g., rate-dependent materials, anisotropic textures, moisture-varying conditions) could require more trials or a higher-dimensional context vector.","The context vector is currently selected manually during inference (the operator must know which surface is being swabbed). Automatic context inference from proprioceptive signals alone, mentioned as future work, would be necessary for truly autonomous deployment and is a non-trivial open problem.","The approach assumes the LSTM backbone, trained on a single baseline surface, captures transferable deformation dynamics. If the baseline surface is atypical, the backbone may encode surface-specific rather than general dynamics, and the context vector would need to compensate for a larger domain gap."],"forward_implications":["If the separation of shared deformation dynamics from domain-specific conditioning is valid, then any robotic task involving compliant tool-environment contact (polishing, brushing, surgical swabbing, wiping) could benefit from the same frozen-backbone-plus-context-vector architecture, reducing per-surface calibration from hundreds of trials to five.","The approach could extend to real-time surface identification: if the context vector z_k encodes effective material properties, clustering or classifying z_k values across surfaces could serve as an implicit material classifier without explicit material sensing.","The parameter-isolation property means a robot operating in a non-stationary environment could maintain a growing bank of context vectors for every surface it encounters, with constant inference cost and no retraining, enabling lifelong accumulation of surface-specific calibration.","The finding that joint torque commands are noise rather than signal for this task suggests that simpler, cheaper robot platforms without torque sensing could achieve comparable force-estimation performance using only kinematics and wrist force-torque data."],"fun_headline_variants":["Few-Shot LSTM Adapts Robotic Swabbing Force Across Surfaces","Context-Aware LSTM Recovers Contact Force in Robotic Swabbing","Frozen LSTM with Few-Shot Tuning Improves Robotic Swab Force","Robotic Swab Force Estimation via Few-Shot LSTM Adaptation","Few-Shot Context Modulation Reduces Robotic Swab Force Error"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The claim that a single low-dimensional context vector optimized on only five trials fully captures the viscoelastic property shift between surfaces rests on the assumption that the baseline-trained LSTM backbone already encodes transferable deformation dynamics, so only a lightweight modulation is needed to adapt. If the backbone has over-specialized to the training surface, or if the context vector's dimensionality is too small to represent the relevant material parameters,","fun_headline_variants_meta":{"raw":{"variants":["Few-Shot LSTM Adapts Robotic Swabbing Force Across Surfaces","Context-Aware LSTM Recovers Contact Force in Robotic Swabbing","Frozen LSTM with Few-Shot Tuning Improves Robotic Swab Force","Robotic Swab Force Estimation via Few-Shot LSTM Adaptation","Few-Shot Context Modulation Reduces Robotic Swab Force Error"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1111,"prompt_tokens":526,"completion_tokens":585,"prompt_tokens_details":null},"tokens_in":526,"tokens_out":585,"duration_ms":27337,"temperature":1.0,"reasoning_tokens":483,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T06:24:18.319275+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"A surface or tool whose viscoelastic response cannot be captured by the context vector's modulation of the frozen backbone, such that five-trial adaptation fails to reduce error below the zero-shot baseline, would falsify the claim that the backbone encodes transferable dynamics and that the context vector is a sufficient adaptation mechanism.","supporting_citations":[],"review_version":1}