{"id":"d9c776b7-60b8-4fc7-bf76-13507e062c44","arxiv_id":"1908.07446","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A neural network trained on human-like priors about hidden coins does not always get fooled by sleight of hand, revealing which magic tricks depend on human-specific biases.","lead":"This paper used a deep-learning tracking tool to follow coins during magic tricks and trained it to guess where hidden coins are, as a human would. It shows examples where the machine's 'belief' differs from a human's, and argues this contrast exposes our own perceptual biases.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The machine is not an independent spectator: its occluded-coin inferences are trained on the same annotator's guesses, so the human-AI asymmetry in Fig. 1F may reproduce human priors rather than reveal them.","rationale":"The reader's weakest assumption identifies exactly the load-bearing flaw: the network's occluded-coin inference was trained on human guesses from the same videos, so the human-machine asymmetry cannot independently reveal human cognitive biases. The manuscript itself flags this in the second stated limitation, and also acknowledges in the first limitation that the quadrant labels are the authors' interpretation rather than a measured fooled/not-fooled classification. I agree that the appropriate disposition is conditional acceptance: the 'artificial spectator' framing is a promising proof of concept, but the central claim is not yet supported by the evidence as presented. The proposed retraining test is concrete and would distinguish between a machine that genuinely forms its own occluded-location inferences and one that merely echoes the annotator's priors. No further objection is needed; the paper's positive contribution as an exploratory demonstration should be credited, but its strong interpretive claim requires the additional validation described above. Therefore the reader's conditional verdict should remain unchanged.","tokens_in":3852,"tokens_out":2827,"duration_ms":29531,"concrete_test":"Retrain the network on the same five videos with all occluded-frame labels removed, following the standard DeepLabCut protocol of labeling only visible coins, and record the inferred coin position during each occlusion interval. If the retrained network no longer places the coin in the closed fist during tricks #1 and #2, or if its inferred trajectories differ substantially from the reported ones, then the machine's 'beliefs' are inherited from the annotator rather than emergent. To test generalizability and the human side, apply the visible-only network to a held-out performance of the same tricks by the same or a different magician, and compare its occlusion inferences to forced-choice location judgments from naive participants; the claimed human-machine asymmetry requires the machine to err differently from humans without having been trained on their errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires the network's hidden-coin inference to be an independent observer whose errors can be contrasted with human errors. But the Methods state that for frames where coins are not visible, 'the human annotating frames clicked where she thought the coin was (i.e. in the left fist; under the right palm). The machine learned those human priors.' Training and evaluation use the same videos, as the authors' second stated limitation concedes. Therefore the machine's 'fooled' state in tricks #1 and #2, inferring the coin in the closed fist when it is not, is literally the supervised target supplied by a human who was also fooled. The quadrant classification in Fig. 1F is not a comparison between human and machine cognition; it is a comparison between one human's guesses and the authors' interpretation of the true coin location. In addition, 'human was fooled' is asserted rather than measured: no participant data are reported, and the first stated limitation notes that the quadrants reflect 'what the human from the perspective of the machine would claim.' The result may stand as a demonstration that DeepLabCut can be trained to output occluded locations, but it cannot support the claim that differences between human and machine errors expose specific cognitive biases.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes using DeepLabCut, a deep-learning pose-estimation tool, as an \"artificial spectator\" for coin magic. The authors record a professional magician performing seven brief motor tricks, train DeepLabCut to track visible coins and to infer occluded coin positions from labels placed by a human annotator, and then classify each trick by whether it would fool a human, the machine, both, or neither (Figure 1F). The central claim is that this setup reveals human cognitive biases through human-machine asymmetries in which tricks succeed.","tokens_in":4076,"tokens_out":4525,"duration_ms":42294,"significance":"The idea of repurposing a deep-learning tracker as a model spectator for magic is creative and potentially useful; if the human-machine comparison were valid, it could open a new way to study cognitive illusions. The authors are commendably explicit about limitations, and they share a supplementary video and data. However, the reported evidence does not support the central interpretive claim: the machine's occluded-coin inference is trained on human guesses from the same videos, and \"human fooled\" is asserted rather than measured. At present the paper is a demonstrative proof-of-concept that DeepLabCut can be trained to extrapolate occluded locations, not a demonstration that machine errors reveal human cognitive biases.","major_comments":[{"comment":"The central claim requires the machine's occluded-coin inference to be an independent observer whose errors can be contrasted with human errors. However, the Methods state that when coins are not visible, \"the human annotating frames clicked where she thought the coin was (i.e. in the left fist; under the right palm). The machine learned those human priors.\" Because the network is trained on the same annotator's guesses from the same videos used for evaluation (as the second stated limitation concedes), the machine's \"fooled\" state in tricks #1 and #2 is the supervised target supplied by a human who was also fooled. The human-machine asymmetry in Figure 1F is therefore a comparison between one human annotator's priors and the authors' interpretation of the true coin location, not between independent human and machine cognitive processes.","section":"Methods (training the network on occluded coin locations)"},{"comment":"The quadrant classification in Figure 1F labels some tricks as \"only fool a human,\" but no participant data are reported and the first stated limitation explicitly says that the quadrants \"reflect what the human from the perspective of the machine would claim with respect to each trick.\" The textual descriptions of tricks #3, #4, and #5 (perceptual overload, misdirection, Gestalt symmetry) are plausible narrative interpretations, not measurements. Without independent human responses to the same video stimuli, the orange quadrant (\"human fooled, machine not\") is unsupported.","section":"Figure 1F and first stated limitation"},{"comment":"The paper provides no statistical test, no error bars, no independent evaluation set, and no objective definition of \"fooled.\" The seven tricks were selected after the fact to illustrate the desired quadrants, and the abstract's broad conclusion that \"magic from the perspective of the machine reveals our own cognitive biases\" goes beyond what a small, post hoc, qualitative set can establish. At minimum, the claim would require pre-registered stimuli, naive human participants, and a quantitative criterion for being fooled.","section":"Overall evaluation (seven selected tricks)"}],"minor_comments":[{"comment":"The Methods report that labeling was \"typically of the order of 5% of the total frame number\" with \"at least 200K iters per video,\" but do not report the total number of frames, the fraction of frames used for validation, or the effect of training iterations; this is insufficient for reproducibility.","section":"Methods"},{"comment":"The claim that the machine \"analyzes the whole image pixel by pixel\" is imprecise: DeepLabCut uses convolutional neural networks with local receptive fields, not exhaustive pixel-by-pixel matching. Please rephrase to describe the actual architecture.","section":"Main text, trick descriptions"},{"comment":"The Supplementary Video is cited, but no timestamps or frame numbers are given for each of the seven tricks, making it difficult for a reader to verify the quadrant assignments in Figure 1F.","section":"Supplementary Video"},{"comment":"There are typographical and spacing issues in the abstract (for example, \"a p p e a r a n d d i s a p p e a r\"), and the text would benefit from proofreading.","section":"Abstract"},{"comment":"The Methods cite \"DeepLabCut 2.0\" but reference [6] describes the original DeepLabCut method; please specify the exact software version and, if applicable, the version-specific documentation.","section":"References"}],"recommendation":"reject","confidential_remarks":"The four stated limitations are candid, but they are not peripheral caveats: the first limitation concedes that the quadrant labels are the authors' interpretation rather than measured human responses, and the second concedes that training and evaluation share the same videos. Together with the fact that the occluded-coin labels are the human annotator's guesses, these limitations mean the central empirical claim cannot be saved by local revision. The paper might be more appropriate as a perspective piece or as a proof-of-concept note if reframed accordingly. I also note that no analysis code or labeled datasets are provided beyond a video link, which further limits verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know this paper is a creative proof of concept, not a rigorous demonstration. The new idea is the 'artificial spectator': instead of using DeepLabCut purely for tracking, the authors train it to guess hidden coin locations the way a human would, then ask where the machine is and isn't fooled by the same motor maneuvers. That framing is fresh, and repurposing a mainstream pose-estimation tool to study perceptual inference in a naturalistic magic setting is genuinely useful. The authors are also candid about the limitations.\n\nThe main soft spot is exactly what the stress-test note says. The machine's occluded-coin inferences are trained on one human annotator's guesses from the same videos. So when the machine 'is fooled' in tricks #1 and #2, inferring the coin in the closed fist, it is reproducing the annotator's prior, not independently producing a human-like error. The 'human fooled' side is asserted from the trick design, not measured. The Figure 1F quadrants are therefore closer to a comparison between the annotator's guesses and the authors' interpretation of the true location than to a clean human-machine dissociation. The authors acknowledge this in the first limitation, which is to their credit, but it means the abstract's claim outruns the evidence.\n\nThe more interesting cases are the ones where the machine is not fooled because it lacks attention or Gestalt priors—tricks #3-#5. Those differences are architectural, not label artifacts. But they are qualitative, with no formal metric or independent validation. The paper is best read as an invitation for a research program, not a result.\n\nThe audience is cognitive scientists working on magic, human-AI comparison, or perception. The paper deserves a serious referee because the idea is original and the prose is honest. But the bar should be high: the authors need held-out validation, actual participant data, and a clear definition of 'fooled' before the central claim is accepted.\n\nMy recommendation: engage with it, send it to review, and insist on those additions.","headline":"Original 'artificial spectator' idea, but the headline asymmetry is largely trained into the machine; worthy of peer review as a proof of concept, not as a validated result.","tokens_in":4625,"tokens_out":3792,"would_cite":false,"duration_ms":35835,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a deep network trained to guess hidden coin locations can serve as an \"artificial spectator\" whose disagreements with humans expose the cognitive biases that make magic work.","keywords":["magic","deep learning","perceptual inference","cognitive biases","misdirection","pose estimation","artificial spectator","sleight of hand"],"falsifier":"Retrain the network on the same tricks using occlusion guesses from several different annotators, and also train a version with no occlusion labels at all. If the \"machine not fooled\" classifications change from one annotator to another, or disappear when no human priors are supplied, the reported asymmetries are an artifact of the training choice rather than a stable feature of human cognition.","tokens_in":3660,"feed_emoji":"🪄","tokens_out":6570,"duration_ms":60447,"temperature":0.7,"pith_summary":"Magic works by exploiting cognitive processes, and this paper proposes that a deep neural network can be trained to watch a magic trick the way a person does: following visible coins, and guessing where the coins are when they disappear. Using a professional magician's brief, purely motor coin tricks, the authors constructed an \"artificial spectator\" whose inferences were then compared with what a human would believe. The result is a classification of tricks into cases where both human and machine are fooled, cases where only the human is fooled, and one case where only the machine is fooled. If this comparison holds, magic tricks become a controlled probe: the places where machine and human disagree mark the specific perceptual, attentional, and Gestalt priors that let deception succeed.","feed_headline":"A neural network watches magic tricks and shows which ones fool people","feed_subtitle":"Trained to guess hidden coins like a human, the AI flags where our biases do the deceiving.","key_machinery":"The load-bearing object is the \"artificial spectator\": a deep neural network trained not only to track visible coins but to infer hidden coin locations from the same human guesses a spectator would make. The mechanism is a two-by-two classification of each trick by whether the human is fooled and whether the machine is fooled, turning a magic performance into a confusion matrix. The inference training is the crucial part: because the network is taught to label occluded coins with human priors, its errors are not random tracking failures but structured departures from those priors. These departures, along with the tricks where the network and human agree, are what carry the argument.","core_discovery":"On the paper's own terms, the central discovery is that the difference between a human and a machine watching the same sleight of hand is not noise but information: it separates the motor mechanics of a trick from the cognitive priors that make it magical. The authors trained a deep-learning pose-estimation network to label coin positions even when coins were occluded, using a human annotator's guesses as the training target; the network thereby learned human-style expectations, such as assuming a coin remains in the closed fist. In tricks #1 and #2, human and machine were both fooled, showing that some cognitive illusions transfer to a network. In tricks #3, #4, and #5, the human was fooled while the machine was not, because the machine never suffered perceptual overload, attentional misdirection, or Gestalt symmetry assumptions. In trick #6, performed badly on purpose, the machine was fooled where a human would not be, because the flying coin looked like a rod rather than a circle. The central claim is that this human-machine confusion diagram is a readable map of human cognitive biases, made visible from the perspective of the machine.","pith_inferences":["Editorial inference: since training labels came from a single annotator, training the same network on many people's occlusion guesses would test whether the machine's \"not fooled\" verdicts are stable; if they are, the method measures individual differences in susceptibility to magic.","Editorial inference: the authors' observation that the machine treats forward and backward video identically suggests a direct experiment—a \"palindrome\" trick—that would fool the machine in both directions while fooling humans in only one; a positive result would confirm the asymmetry comes from temporal expectation.","Editorial inference: the results point to a design loop the authors only sketch: generate candidate movements, have a human-like network watch them, and use its misinferences to predict which routines will deceive a live audience."],"forward_implications":["Magic can be decomposed into motor maneuvers whose perceptual effect is independently measurable: a network's inference errors tell which moments carry the deception.","Training neural networks on occluded, not just visible, targets turns them from trackers into models of expectation, opening a general method for probing perception beyond magic.","Tricks that fool both human and machine point to priors that are shared and learnable, while tricks that fool only humans point to attentional and Gestalt mechanisms that current models lack.","The same comparison could guide the design of adversarial displays: routines would be iterated against a human-like network before being shown to people."],"supporting_citations":[{"why":"Supplies the deep-learning pose-estimation method used to track fingers and coins and to infer coin locations.","marker":"[6]"},{"why":"Establishes the scientific study of magic as a window onto attention and awareness, which the paper builds on.","marker":"[3]"},{"why":"Provides the eye-tracking background and the cognitive-neuroscience framing of magic that motivates the artificial-spectator idea.","marker":"[5]"},{"why":"Supplies the adversarial-image analogy against which the authors position their adversarial cognitive tricks.","marker":"[7]"},{"why":"Motivates the premise that machines trained with human-like priors can be used to study human cognition.","marker":"[9]"}],"fun_headline_variants":["AI exposes the cognitive tricks behind magic","Magic tricks show AI how humans fool themselves","Neural network learns magic, reveals human bias","Why AI resists magic that fools humans","Machine vs human: magic reveals our biases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the network's hidden-coin guesses are not just a stored copy of the human annotator's guesses used to train it; if the network only echoes that one person's prior, the human-machine differences reveal nothing new about human perception in general.","fun_headline_variants_meta":{"raw":{"variants":["AI exposes the cognitive tricks behind magic","Magic tricks show AI how humans fool themselves","Neural network learns magic, reveals human bias","Why AI resists magic that fools humans","Machine vs human: magic reveals our biases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1356,"prompt_tokens":927,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":543,"tokens_out":429,"duration_ms":4764,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:17:44.216123+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the network on the same tricks using occlusion guesses from several different annotators, and also train a version with no occlusion labels at all. If the \"machine not fooled\" classifications change from one annotator to another, or disappear when no human priors are supplied, the reported asymmetries are an artifact of the training choice rather than a stable feature of human cognition.","supporting_citations":[{"cited_title":"Nature Neuroscience 21: 1281–1289","cited_arxiv_id":null,"evidence_quote":"Supplies the deep-learning pose-estimation method used to track fingers and coins and to infer coin locations."},{"cited_title":"Nature Reviews Neuroscience 9:871-9","cited_arxiv_id":null,"evidence_quote":"Establishes the scientific study of magic as a window onto attention and awareness, which the paper builds on."},{"cited_title":"Current Biology 26(10):R387-407","cited_arxiv_id":null,"evidence_quote":"Provides the eye-tracking background and the cognitive-neuroscience framing of magic that motivates the artificial-spectator idea."},{"cited_title":"artificial illusionism","cited_arxiv_id":null,"evidence_quote":"Supplies the adversarial-image analogy against which the authors position their adversarial cognitive tricks."},{"cited_title":"Basic Books, NY","cited_arxiv_id":null,"evidence_quote":"Motivates the premise that machines trained with human-like priors can be used to study human cognition."}],"review_version":1}