{"id":"c0f2fe6f-b401-425a-ac02-ba819f9aa499","arxiv_id":"2501.12810","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A dual-channel V1-MT-style model trained on non-Lambertian materials acquires human-like second-order motion perception.","lead":"A neural network modeled on the brain's V1-MT motion pathway grows a human-like ability to see 'second-order' motion, such as textures and reflections, after training on videos of shiny and transparent objects. The result suggests a new evolutionary reason for this visual ability, but the paper's main human comparison uses training data as the test set.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Most load-bearing risk: the second-order benchmark is contaminated by first-order cues (water waves, swirls, and random flow fields are pixel warps, not Fourier-balanced), so the non-Lambertian training advantage may reflect improved first-order feature tracking rather than emergent second-order…","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the second-order benchmark modulations are not verified to be free of first-order luminance cues. My stress-test confirms this is the decisive issue for the strongest claim, not merely a secondary caveat. The paper itself states that water wave, swirl, and random flow field modulations warp pixels, which is a direct description of first-order luminance displacement. Because the headline result averages over seven modulation types, a model trained to estimate object motion despite optical turbulence on non-Lambertian surfaces could plausibly improve on these pixel-warp conditions by better first-order feature matching, without ever developing a genuine second-order pathway. Restricting evaluation to truly Fourier-balanced stimuli such as drift-balanced motion would settle the question. I do not think the Sintel train/test overlap, though a real problem for the human-alignment comparisons, is the more fundamental threat to the paper's central scientific claim, because the second-order emergence experiment uses separately rendered datasets. The benchmark purity issue is more load-bearing, and the reader already flagged it; hence the verdict of rejection should stand unchanged. A single training run and lack of significance testing would matter only after the measurement itself is validated, so the concrete test above is the most direct way to resolve the concern.","tokens_in":19928,"tokens_out":5312,"duration_ms":62241,"concrete_test":"Compute an Adelson-Bergen first-order motion-energy balance for every stimulus in the Section 4.3.2 benchmark, and retain only the subset (at minimum drift-balanced motion) with zero net directional first-order energy. Then recompute the model-vs-human correlations shown in Fig. 5-C for the diffuse-trained and non-diffuse-trained models on that subset. If the non-diffuse training advantage disappears or falls well below r = 0.902, the central emergence claim is an artifact of first-order cues in the water wave, swirl, and random-flow modulations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that training on non-Lambertian materials causes the dual-channel model to acquire human-comparable second-order motion perception. That claim depends on the Section 4.3.2 benchmark actually isolating second-order motion. It does not. Only drift-balanced motion is explicitly Fourier-balanced; the other six modulations are admitted in the text to be near-indiscernible in Fourier space rather than verified to be free of first-order information. More seriously, the generation procedure states that water wave, swirl, and random flow field modulations warp pixels using specific flow fields. Warping pixels physically displaces luminance, so a conventional first-order motion-energy sensor has a directional signal to lock onto. The natural-image background being static does not remove this cue; the warp itself creates local luminance displacement. Thus the high average correlation with human responses (r = 0.902) across all seven modulations can be achieved by a first-order motion estimator that is robust to appearance changes, which is exactly what training on glossy and transparent renderings would teach. The second-order channel and the evolutionary-functional conclusion are therefore not identified by the experiment unless the evaluation is restricted to genuinely balanced conditions. The paper's own phrasing flags this limitation but does not resolve it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a two-stage, biologically inspired model of human visual motion processing. Stage I contains a trainable motion-energy sensor bank (first-order channel) plus a 3D CNN preprocessing channel intended to extract higher-order features, and Stage II uses a self-attention-based motion graph with recurrent GRU integration to produce dense optical flow and, via graph cuts, motion segmentation. The authors train the model on naturalistic datasets and on newly rendered datasets containing diffuse or non-diffuse (glossy, transparent, metallic) objects, and they claim that training on non-Lambertian materials endows the model with human-comparable second-order motion perception. The paper reports first-order motion benchmarks against human perceived flow on Sintel, a seven-modulation second-order benchmark with human psychophysical data, comparisons with several optical-flow models, and qualitative motion-segmentation results.","tokens_in":20217,"tokens_out":3967,"duration_ms":43200,"significance":"If the central claim were fully supported, the paper would make an important interdisciplinary contribution: it would provide a computational account of how second-order motion perception could arise from natural statistics, and it would offer a human-aligned flow model that also handles non-Lambertian optical turbulence. The manuscript has notable strengths: it ships code and data-generation pipelines, it uses human perceived flow with partial correlations to control for ground-truth confounds, and it attempts in-silico neurophysiological validation with component/pattern cell analyses. However, the load-bearing evidence for the evolutionary hypothesis is weakened by two methodological problems: the second-order benchmark does not isolate second-order motion for six of its seven modulations, and the model is trained and evaluated on overlapping Sintel data. The paper's significance is therefore conditional on resolving these issues.","major_comments":[{"comment":"The second-order benchmark is contaminated by first-order luminance cues. The text states that water wave, swirl, and random flow field modulations 'warp pixels using specific flow fields'; pixel warping creates ordinary luminance displacement, which a first-order motion-energy sensor can lock onto. Only drift-balanced motion is explicitly designed to be Fourier-balanced, and the claim that the other modulations are 'near-indiscernible in Fourier space' is not supported by any spectral analysis or control condition. The high average correlation with human responses (r = 0.902) and the non-diffuse training advantage could therefore reflect better first-order feature tracking rather than a genuinely emergent second-order mechanism. The authors should either restrict the claim to the drift-balanced condition or provide quantitative evidence that each of the seven modulations contains no usable first-order motion information; a simple test is to show that a first-order-only baseline cannot recover the modulation direction above chance.","section":"Section 4.3.2 and Fig. 5-C"},{"comment":"The model is trained on MPI-Sintel and Sintel-Slow as part of Dataset A (Section 4.2.1) and then validated on the Sintel slow benchmark with human-perceived flows (Section 2.2, Table 1), with no reported train/validation split. This creates a risk of circularity: the high partial correlations with human responses on first-order natural scenes could be inflated by direct supervision on the same benchmark's ground truth. The authors should clarify whether the Sintel-Slow evaluation sequences were excluded from training, or retrain and evaluate on non-overlapping splits.","section":"Section 4.2.1 vs. Section 2.2 and Table 1"},{"comment":"The central diffuse-versus-non-diffuse ablation is presented without training-seed variance or inferential statistics. The text says the results 'indicate that both the dataset material properties and the model architecture significantly influence' second-order motion perception, but no significance test is reported, and the per-modulation comparisons in Fig. 5-C appear to be single runs. The authors should report multiple training runs, error bars across seeds, and a statistical comparison focused on the drift-balanced condition, which is the only modulation that unambiguously requires second-order processing.","section":"Section 4.2.1 and Fig. 5-C"}],"minor_comments":[{"comment":"For the 'random noise' and 'Gaussian blur' modulations, it is unclear whether the carrier itself moves or whether the modulation pattern moves over a static carrier; please clarify the generation procedure.","section":"Section 4.3.2"},{"comment":"The phrase 'near-indiscernible in Fourier space' is vague; please report a quantitative measure, such as the ratio of first-order motion energy in the balanced and unbalanced conditions.","section":"Section 2.3"},{"comment":"There is a typographical error in the definition of Lon: the term 'S * Im[G]) * Im[T]' has an unbalanced parenthesis.","section":"Equation (2)"},{"comment":"The labels 'Type-I' and 'Type-II' for diffuse and non-diffuse training are not explicitly mapped to Datasets D and E; please make this mapping explicit.","section":"Section 4.2.1"},{"comment":"The repository and project links contain the placeholder 'anoymized'; please update them to the final publicly accessible URLs.","section":"Code Availability"}],"recommendation":"major_revision","confidential_remarks":"The paper's core hypothesis is interesting and the modeling effort is substantial, but the main empirical claim is not yet identified by the experiments as reported. The benchmark contamination and the train/test overlap are central issues that should be resolved before publication. I would want the revision to include a re-analysis restricted to validated second-order stimuli, a clear data split, and per-seed statistics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has a genuinely new scientific idea — that second-order motion perception is acquired because non-Lambertian materials (gloss, transparency) make pure luminance tracking unreliable — and the controlled diffuse-vs-non-diffuse training experiment is the right way to test it. The evidence in hand is weaker than the abstract claims, and the two weakest spots are the benchmark and the statistics.\n\nWhat's actually new: the material-controlled training paradigm. Rendering matched scenes with Lambertian vs specular/glossy/transparent objects and showing that only the non-diffuse-trained dual-channel model acquires human-comparable performance on second-order motion stimuli. I don't know of prior work posing the ecological question this way. The dual-channel architecture itself (trainable motion-energy bank + 3D CNN nonlinear preprocessing + graph integration) is assembled from known parts, but the assembly is sensible, and the model does replicate a wide range of textbook findings (component/pattern cell populations, aperture problem, adaptive pooling). They also ship a lot: datasets, code, and human psychophysics on the second-order benchmark.\n\nWhere it gets soft. First, Section 4.2.1 lists MPI-Sintel and Sintel-Slow in the training set, and Section 2.2 validates on the Sintel slow benchmark, with no stated split. The partial-correlation human-alignment numbers in Table 1 and Fig. 3 could be in-distribution artifacts. That's a concrete fix, but it needs to be stated and re-run on held-out data. Second, the second-order benchmark: of the seven modulations, only drift-balanced is genuinely Fourier-balanced. Section 2.3 admits the water-wave/swirl class is \"not pure second-order motion but near-indiscernible in Fourier space,\" and Section 4.3.2 says those conditions literally warp pixels. Warping displaces luminance; a first-order energy sensor with good appearance invariance can lock onto residual cues. \"Near-indiscernible\" is asserted, not verified by spectral analysis. The RAFT control (r=0.102) is suggestive but not clean — RAFT isn't trained on these modulations. The clean control is the paper's own first-order channel, and we need the per-modulation breakdown to see whether the non-diffuse advantage survives on drift-balanced alone. Third, the headline comparison rests on one training run per condition; the Fig. 5-C error bars are across scenes, not seeds. No significance tests anywhere. The limitations section is honest about interpretability but doesn't flag any of this.\n\nNone of this kills the idea. The experiment is the right design, and keeping the architecture fixed while varying only material properties is a solid comparison. The claim is less \"the channel emerged from scratch\" than \"the environment makes the second-order channel useful,\" which is still a real finding if the controls hold. This is major-revision territory, not rejection: add held-out evaluation, seed statistics, spectral checks of the modulations, and the drift-balanced-only comparison.\n\nRecommendation: send it to review. A serious referee will ask for exactly these controls, and if they hold, the ecological-function claim is a real advance.","headline":"A genuinely new ecological hypothesis for second-order motion perception, tested with the right controlled experiment, but the current evidence is single-seed, partly in-distribution, and built on a second-order benchmark the authors themselves admit is not purely second-order.","tokens_in":20772,"tokens_out":7627,"would_cite":true,"duration_ms":78837,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dual-pathway neural network trained to estimate the motion of glossy, transparent objects acquires human-like second-order motion perception without ever seeing explicit second-order training stimuli.","keywords":["visual motion perception","second-order motion","optical flow","motion energy model","graph neural network","V1-MT pathway","motion segmentation","non-Lambertian materials"],"falsifier":"One decisive check is to run a standard intensity-conservation optical-flow algorithm on the seven modulation movies and measure whether the ground-truth motion direction is recoverable above chance from luminance alone; if any modulation passes, re-score the model on a strictly Fourier-balanced subset—the emergence claim would collapse if the dual-channel model's advantage over the first-order channel disappears once contaminating modulations are removed.","tokens_in":19691,"feed_emoji":"👁️","tokens_out":7706,"duration_ms":77497,"temperature":0.7,"pith_summary":"This paper tries to show that a machine can learn to perceive visual motion the way humans do, including second-order motion—motion carried by contrast or texture modulation rather than by moving luminance boundaries. The authors build a two-stage model that mirrors the cortical V1-to-MT pathway: a bank of trainable motion-energy sensors feeds a recurrent graph network that integrates motion globally. To that they add a second channel whose 3D convolutional preprocessing extracts nonlinear spatiotemporal features before motion-energy sensing. Their central claim is that training this dual-channel model to estimate the motion of non-Lambertian objects—glossy, transparent, and metallic surfaces that generate moving highlights and refractions—spontaneously endows it with human-comparable second-order motion perception. If true, this would explain why biological vision evolved a separate non-Fourier motion system: it lets an observer track object motion through the optical turbulence that real materials create.","feed_headline":"Glossy-object training gives AI human-like second-order motion perception","feed_subtitle":"A V1-MT-style network learns contrast-modulated motion from shiny, transparent objects—matching human vision where standard flow models…","key_machinery":"The load-bearing machinery is a dual-channel V1-MT-style network. Stage I's first-order channel holds 256 trainable quadrature Gabor filters (spatiotemporal motion-energy sensors) arranged in a multiscale pyramid, with preferred speeds and directions learned during training. A second channel stacks five 3D convolutional layers with residual ReLU connections before the same energy computation, realizing the filter-rectify-filter preprocessing thought to underlie second-order motion. Stage II turns every spatial location into a node of a fully connected graph whose adjacency is a cosine-similarity (self-attention) matrix; a convolutional gated recurrent unit repeatedly mixes motion signals across the graph to integrate global motion, and normalized cuts on that graph give object segmentation with no extra training. The datasets that trigger the emergence are generated by a physics-based rendering pipeline, with diffuse (matte) and non-diffuse (specular, glossy, transparent, anisotropic) versions of the same scenes, so the only controlled difference is material.","core_discovery":"The discovery is that a task not obviously related to second-order motion—estimating the object-level motion of non-Lambertian materials—is sufficient to produce a human-like second-order motion system. On a benchmark with seven contrast- and texture-modulation types (drift-balanced motion, Gaussian blur, water waves, swirls, noise, and Fourier and pixel shuffles), human observers matched the physical ground truth with mean correlation 0.983; the dual-channel model trained on non-diffuse rendered materials averaged 0.902 with human responses, whereas a representative state-of-the-art optical-flow model scored 0.102. The higher-order channel's units become direction-tuned to drifting second-order gratings, and this tuning is sharpened by non-diffuse training, while the first-order channel remains tuned to luminance motion. The same model also yields training-free object segmentation from the motion graph and reproduces known component- and pattern-cell distributions across the V1-to-MT stages.","pith_inferences":[],"forward_implications":["Second-order motion perception in the model is not a fixed architectural property: it appears only after training on non-Lambertian motion, so the learning signal is the material-driven optical turbulence rather than the presence of explicit second-order labels.","A single trained model covers both major human motion phenomena: first-order flow estimation that correlates with human perception beyond ground truth, and second-order perception that reaches an average correlation of 0.902 with human judgments.","The motion graph encodes object structure implicitly, so object segmentation from motion—including segments defined only by drift-balanced motion—falls out of the same trained network without task-specific supervision.","The dual-channel design stabilizes flow estimation in naturally noisy scenes such as transparent containers with moving liquid, where luminance-based flow models become unstable.","The model replicates the physiological split between component cells and pattern cells across stages, suggesting the V1-MT pathway and the graph-integration mechanism are functionally equivalent at the level of population responses.","A testable extension follows from the paper's logic: controlled-rearing experiments with animals in environments rich in specular surfaces (water, glossy foliage) should produce stronger second-order motion sensitivity than matte-only environments, matching the model's learning trajectory.","The result predicts that adding other sources of optical turbulence during training—caustics, subsurface scattering, heat shimmer—should further improve generalization of the higher-order channel; the paper does not run these ablations.","The same learning principle could be transferred to engineering: inserting an end-to-end nonlinear preprocessing stream into intensity-conservation optical-flow models might make them robust to non-Lambertian scenes, but the paper leaves that engineering transfer implicit."],"supporting_citations":[{"why":"Defines drift-balanced motion, the canonical second-order stimulus used in the benchmark.","marker":"[16]"},{"why":"Establishes the prior single-channel version of the model that provided the first-order baseline lacking second-order capacity.","marker":"[15]"},{"why":"Supplies the human-perceived optical flow dataset and the partial-correlation method used to show human alignment.","marker":"[7]"},{"why":"Provides the recurrent optical-flow architecture used as a state-of-the-art comparison and the guide for the separable GRU design.","marker":"[33]"},{"why":"A biologically inspired dorsal-stream model used as a baseline that does not capture second-order motion.","marker":"[11]"},{"why":"The rendering and physics pipeline used to generate the diffuse and non-diffuse material-controlled training datasets.","marker":"[82]"},{"why":"Provides normalized cuts, the training-free graph-partition method used for motion-based segmentation.","marker":"[14]"},{"why":"Defines the plaid-stimulus component- and pattern-cell analysis and partial-correlation classification replicated by the model.","marker":"[27]"},{"why":"Supplies the filter-rectify-filter conceptual basis for the nonlinear 3D CNN preprocessing channel.","marker":"[26]"}],"fun_headline_variants":["Training on shiny objects makes AI see motion like humans do","AI learns human-like motion perception from glossy objects","Glossy objects teach AI to see second-order motion","Shiny-object training gives AI human-like motion perception","Non-Lambertian training yields human-like second-order motion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's seven second-order modulations are assumed to contain no usable first-order luminance cues, but only drift-balanced motion is explicitly designed to be Fourier-balanced; water waves, swirls, noise, blur, and shuffles are called near-indiscernible rather than verified cue-free, so part of the model's measured second-order advantage could be better first-order feature extraction.","fun_headline_variants_meta":{"raw":{"variants":["Training on shiny objects makes AI see motion like humans do","AI learns human-like motion perception from glossy objects","Glossy objects teach AI to see second-order motion","Shiny-object training gives AI human-like motion perception","Non-Lambertian training yields human-like second-order motion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2661,"prompt_tokens":1031,"completion_tokens":1630,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1551}},"tokens_in":647,"tokens_out":1630,"duration_ms":11769,"temperature":1.0,"reasoning_tokens":1551,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:46:22.401593+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One decisive check is to run a standard intensity-conservation optical-flow algorithm on the seven modulation movies and measure whether the ground-truth motion direction is recoverable above chance from luminance alone; if any modulation passes, re-score the model on a strictly Fourier-balanced subset—the emergence claim would collapse if the dual-channel model's advantage over the first-order channel disappears once contaminating modulations are removed.","supporting_citations":[{"cited_title":"& Sperling, G","cited_arxiv_id":null,"evidence_quote":"Defines drift-balanced motion, the canonical second-order stimulus used in the benchmark."},{"cited_title":"& Nishida, S","cited_arxiv_id":null,"evidence_quote":"Establishes the prior single-channel version of the model that provided the first-order baseline lacking second-order capacity."},{"cited_title":"& Deng, J","cited_arxiv_id":null,"evidence_quote":"Provides the recurrent optical-flow architecture used as a state-of-the-art comparison and the guide for the separable GRU design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The rendering and physics pipeline used to generate the diffuse and non-diffuse material-controlled training datasets."},{"cited_title":"& Malik, J","cited_arxiv_id":null,"evidence_quote":"Provides normalized cuts, the training-free graph-partition method used for motion-based segmentation."},{"cited_title":"& New- some, W","cited_arxiv_id":null,"evidence_quote":"Defines the plaid-stimulus component- and pattern-cell analysis and partial-correlation classification replicated by the model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the filter-rectify-filter conceptual basis for the nonlinear 3D CNN preprocessing channel."}],"review_version":1}