REVIEW 5 major objections 6 minor 189 references
What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective
T0 review · 5 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper argues that an integrated framework combining visual, auditory, and audiovisual selective attention can bridge human behavioral and neural findings with computational simulation.
desk verdict A competent, useful review of selective attention with a smart audiovisual frame, but its Section 5.2 overstates what computational models have actually demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is crossmodal integration and conflict resolution: the brain binds sights and sounds that plausibly come from one source, resolves mismatches by weighting the modality with higher reliability, and uses attention to gate the coupling between sensory areas. On the computational side, the load-bearing mechanisms are saliency maps with winner-take-all selection, locally excitatory and globally inhibitory oscillator networks for auditory stream segregation, and deep attention architectures with top-down prediction and conflict-monitoring modules, all of which the paper connects to neural mechanisms such as gamma-band enhancement and alpha-band inhibition.
What would settle it
Take an audiovisual model built on the paper's integrated mechanisms and compare it with a simple fusion baseline in a naturalistic incongruent scene, such as a speaker's voice arriving from a different location than the visible lip movements; if the human-derived crossmodal mechanisms never improve localization or conflict resolution beyond the baseline, the central bridge claim is not supported.
Extended reading notes
Core claim
The central claim is that the similarities and differences among selective attention mechanisms across modalities are exactly what an integrated framework needs to capture, and that such a framework can bridge human behavioral and neural patterns with intelligent system simulation. The paper assembles evidence that visual attention, auditory attention, and audiovisual integration all involve a common weight-allocation logic, but express it through different routes: saliency maps and winner-take-all selection in vision, stream segregation and top-down prediction in hearing, and modality-appropriate weighting and conflict monitoring when sights and sounds compete. It then maps these findings onto computational models, from classical saliency and oscillator networks to deep learning attention mechanisms, and argues that crossmodal modeling is not an optional extra but a necessary step for real-world robotics and human-robot interaction.
Load-bearing premise
The review's roadmap depends on the assumption that laboratory effects such as pop-out visual search, listening tasks with different sounds in each ear, and ventriloquism survive in rich real-world scenes and in autonomous agents; if those effects are task-specific, the proposed bridge between human attention and computational modeling weakens.
Editorial extensions
If this is right
- If the integrated framework is correct, an audiovisual model should reproduce the ventriloquism effect: visual lip movements should bias sound localization more strongly when the auditory cue is unreliable.
- Cocktail-party findings imply that top-down prediction of a target voice, not just bottom-up stream separation, should improve speech separation in multi-speaker noise.
- Neural oscillation findings suggest that computational models could benefit from gating mechanisms analogous to alpha suppression of task-irrelevant inputs and gamma enhancement of task-relevant inputs.
- A priority map that includes selection history and semantic meaning should outperform maps based only on physical salience when predicting where humans look in real scenes.
- Co-saliency and meaning-map approaches could be combined to improve image and video interpretation by prioritizing the most informative content for humans.
- Robots that resolve crossmodal conflicts by choosing the more reliable modality should localize sounds and recognize events more accurately in naturalistic environments.
Reading between the lines
- Going beyond the paper, the reviewed findings suggest that modern self-attention architectures could be made more human-like by adding explicit modality-reliability weighting, so that auditory temporal precision or visual spatial precision wins depending on the situation.
- A testable extension would be to train a deep network on audiovisual conflict data with and without a conflict-monitoring module, and to compare its error patterns directly with human performance on incongruent speaker-lip movement scenes.
- The paper's emphasis on conflict-driven curiosity implies a developmental-robotics prediction: agents that treat crossmodal mismatches as learning signals, rather than errors to discard, should acquire new multimodal concepts faster.
- An implicit consequence for experimental psychology is that computational models could serve as falsifiable implementations of attention theories, making theoretical disagreements such as stimulus-driven versus goal-driven capture testable in simulation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This review article surveys selective attention from psychology, neuroscience, and computational modeling, focusing on visual 'pop-out', auditory 'cocktail party', and audiovisual crossmodal integration/conflict resolution. It reviews theories, behavioral and neural mechanisms, and computational models for each, then discusses gaps and future directions. The paper's stated aims are to integrate unimodal and crossmodal selective attention findings and to bridge human behavioral/neural patterns with intelligent-system simulation, particularly in robotics.
Significance. If fully substantiated, the review would provide a useful interdisciplinary map and a roadmap for transferring psychological findings into computational models. As a compilation it is strong: it covers classic theories (top-down/bottom-up, priority maps, neural oscillations, free-energy), summarizes a broad literature, and organizes it around three well-chosen representative effects. It also explicitly acknowledges several limitations of the human literature (e.g., correlational neural evidence, Section 2.3). The central gap is the computational side of the claimed bridge: the key evidence in Section 5.2 is self-cited and not quantitatively validated against human data, so the roadmap is better described as a research program than a demonstrated bridge. The paper's explicit, falsifiable roadmap and clear admission of loose connections between psychology and computer science are assets; the missing cross-validation is the main obstacle to accepting the strongest claims.
major comments (5)
- [5.2] The paragraph beginning 'Many studies focus on multimodal fusion' reports human behavioral experiments (Parisi et al. 2017, 2018; Fu et al. 2018) and then states that 'human-like responses were modelled' and that the 'work above shows that DL can simulate humans' selective attention and conflict resolution.' However, no quantitative comparison is provided between model outputs and the human psychophysical results described in the same paragraph: there are no effect sizes for the congruence effect in the model, no statistical test against human performance, and no ablation isolating the contribution of the attention/crossmodal layer. Since the paper's central claim is that computational models can bridge human behavioral/neural patterns, this evidence gap is load-bearing. Please either add the quantitative comparison from the cited papers or soften the claim to 'a proof-of-concept simulation' and explicitly state the missing validation as a limitation.
- [5.2] The computational crossmodal review is based almost entirely on the authors' own prior work (Parisi et al. 2017, 2018; Fu et al. 2018; Barros et al. 2018). No independent work implementing audiovisual selective attention with quantitative validation is cited. This is not a fatal flaw for a review, but it means the general roadmap rests on a single line of evidence. Please either survey independent models (e.g., audiovisual saliency or multisensory integration models outside the authors' group) or explicitly frame the section as a case study from the authors' lab rather than a representative field-wide survey.
- [1, 5.1, 6] The Introduction and Section 5.1 set up a bridge between laboratory selective-attention effects (pop-out, cocktail party, ventriloquism) and autonomous agents, and Section 6 makes robotics recommendations based on this bridge. However, the review never addresses whether these simplified laboratory effects survive real-world complexity (e.g., moving cameras, reverberant audio, task-relevant semantic context). Section 2.5 itself concedes that 'the connection between computer science models and psychology is still loose and broad.' Please add a discussion of ecological validity for each of the three representative effects, or explicitly state that the transfer is an open research question.
- [2.3 vs 5.1] Section 2.3 correctly concedes that the neural-oscillation evidence is 'mainly correlations and descriptive results rather than causal relationships.' Yet Section 5.1 later states that the gamma-alpha oscillation pattern 'is proposed to be the information gating mechanism,' and the summary attributes this mechanism without repeating the correlational caveat. Please carry the same epistemic qualifier through to Section 5.1, or clearly distinguish correlational findings from mechanistic claims.
- [Abstract, Introduction, 3, 4, 5] The Abstract and Introduction promise an 'integrated framework' that 'combine[s] and compare[s] selective attention mechanisms from different modalities.' In practice, Sections 3 and 4 are parallel reviews with separate computational-model subsections, and Section 5 is largely independent; cross-cutting comparisons appear only in each section's closing paragraphs. A reader looking for an explicit side-by-side comparison of visual and auditory attention mechanisms (e.g., a table of shared and differing computational principles) will not find it. Please add such a comparison or temper the claim of integration.
minor comments (6)
- [2.5] The abbreviation 'LTSM' should be 'LSTM' (Long Short-Term Memory).
- [3.1] The phrase 'super colliculus' should be 'superior colliculus'.
- [6] The text reads 'One the one hand' and should read 'On the one hand'.
- [6] The phrase 'the -state-of-the-art approaches' contains a stray hyphen and should be 'the state-of-the-art approaches'.
- [Figures] Figures 1, 2, and 4 are explicitly labeled 'adapted from' previous publications; please verify that all required permissions for reuse have been obtained for the final version.
- [5.2] The iCub robot is mentioned without a citation; please add a reference to the iCub platform (e.g., Metta et al., 2008) or clarify the source.
Circularity Check
No circularity: the review is descriptive and rests on independent empirical literature; its concentrated self-citations in Section 5.2 raise validation rather than circularity concerns.
full rationale
This article is a narrative review rather than a derivation, so the circularity patterns do not apply. The psychological and neural sections are grounded in independent experimental literature (e.g., Cherry 1953; Desimone and Duncan 1995; Corbetta and Shulman 2002; Talsma et al. 2010), and the paper explicitly acknowledges open problems rather than claiming a closed formal result. Section 5.2 does rely heavily on the authors' own prior computational work (Parisi et al. 2017, 2018; Fu et al. 2018; Barros et al. 2018) to support the statement that deep learning can simulate crossmodal conflict resolution. That is a self-citation cluster and a validation/overclaim concern: the review does not provide quantitative model-human comparisons, and the concluding sentence in Section 5.2 exceeds what is demonstrated. However, it is not circular: the cited experiments are empirical studies with human participants and models, not definitions of the conclusion in terms of the premises. No equation is fitted and renamed a prediction, no uniqueness theorem or ansatz is imported from the authors' prior work, and no target quantity is defined in terms of an input quantity. The paper even concedes that 'research about selective attention and conflict resolution in computer science is limited' and that the connection between computer science models and psychology is 'still loose and broad.' Therefore no step in the paper reduces by construction to its own input; the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption Laboratory findings such as pop-out, dichotic listening, and ventriloquism generalize to real-world scenarios and to computational agents.
- domain assumption Neural correlates of attention, such as ERP components, alpha oscillations, and fMRI activations, provide mechanistic evidence rather than mere correlation.
Cite this review
Pith. "Pith review of What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective." pith.science (2026). https://pith.science/paper/VUC3GKVU
@misc{pith2026190905654,
author = {Pith},
title = {Pith review of: What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/VUC3GKVU}},
note = {Machine review of arXiv:1909.05654}
}
read the original abstract
Selective attention plays an essential role in information acquisition and utilization from the environment. In the past 50 years, research on selective attention has been a central topic in cognitive science. Compared with unimodal studies, crossmodal studies are more complex but necessary to solve real-world challenges in both human experiments and computational modeling. Although an increasing number of findings on crossmodal selective attention have shed light on humans' behavioral patterns and neural underpinnings, a much better understanding is still necessary to yield the same benefit for computational intelligent agents. This article reviews studies of selective attention in unimodal visual and auditory and crossmodal audiovisual setups from the multidisciplinary perspectives of psychology and cognitive neuroscience, and evaluates different ways to simulate analogous mechanisms in computational models and robotics. We discuss the gaps between these fields in this interdisciplinary review and provide insights about how to use psychological findings and theories in artificial intelligence from different perspectives.
Figures
Reference graph
Works this paper leans on
-
[1]
what” and “where
Ahveninen, J., Ja¨a¨skela¨inen, I. P., Raij, T., Bonmassar, G., Devore, S., Ha¨ma¨la¨inen, M., et al. (2006). Task-modulated “what” and “where” pathways in human auditory cortex. Proceedings of the National Academy of Sciences 103, 14608–14613
2006
-
[2]
What” and “where
Alain, C., Arnott, S. R., Hevenor, S., Graham, S., and Grady, C. L. (2001). “What” and “where” in the human auditory system. Proceedings of the National Academy of Sciences 98, 12301–12306
2001
-
[3]
and Burr, D
Alais, D. and Burr, D. (2004). The ventriloquist effect results from near-optimal bimodal integration. Current Biology 14, 257–262
2004
-
[4]
A., Laurent, P
Anderson, B. A., Laurent, P. A., and Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences 108, 10367–10371
2011
-
[5]
V., and Theeuwes, J
Awh, E., Belopolsky, A. V., and Theeuwes, J. (2012). Top-down versus bottom-up attentional control: A failed theoretical dichotomy. Trends in Cognitive Sciences 16, 437–443
2012
-
[6]
Aytar, Y., Castrejon, L., Vondrick, C., Pirsiavash, H., and Torralba, A. (2017). Cross-modal scene networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 40, 2303–2314
2017
-
[7]
Ba, J., Mnih, V., and Kavukcuoglu, K. (2014). Multiple object recognition with visual attention. In International Conference on Learning Representations
2014
-
[8]
Bacon, W. F. and Egeth, H. E. (1994). Overriding stimulus-driven attentional capture. Perception & Psychophysics 55, 485–496
1994
Show all 189 references
-
[9]
Baddeley, A., Hitch, G., and Bower, G. (1974). Recent advances in learning and motivation. Working Memory 8, 647–667
1974
-
[10]
Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations
2014
-
[11]
and An, S
Bai, S. and An, S. (2018). A survey on automatic image caption generation. Neurocomputing 311, 291–304
2018
-
[12]
Barbey, A. K. (2018). Network neuroscience theory of human intelligence. Trends in Cognitive Sciences 22, 8–20 Fu et al. Selective attention mechanisms and modeling 21
2018
-
[13]
I., Fu, D., Liu, X., and Wermter, S
Barros, P., Parisi, G. I., Fu, D., Liu, X., and Wermter, S. (2018). Expectation learning and crossmodal modulation with a deep adversarial network. In 2018 International Joint Conference on Neural Networks (IJCNN) (IEEE), 1–8
2018
-
[14]
Bee, M. A. and Micheyl, C. (2008). The cocktail party problem: what is it? how can it be solved? and why should animal behaviorists study it? Journal of Comparative Psychology 122, 235–251
2008
-
[15]
Benes, F. M. (2000). Emerging principles of altered neural circuitry in schizophrenia. Brain Research Reviews 31, 251–269
2000
-
[16]
Bizley, J. K. and Cohen, Y. E. (2013). The what, where and how of auditory -object perception. Nature Reviews Neuroscience 14, 693–707
2013
-
[17]
and Jensen, O
Bonnefond, M. and Jensen, O. (2015). Gamma activity coupled to alpha phase as a mechanism for top - down controlled gating. PloS One 10, e0128667
2015
-
[18]
and Itti, L
Borji, A. and Itti, L. (2012). State -of-the-art in visual attention modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence 35, 185–207
2012
-
[19]
M., Braver, T
Botvinick, M. M., Braver, T. S., Barch, D. M., Carter, C. S., and Cohen, J. D. (2001). Conflict monitoring and cognitive control. Psychological Review 108, 624–652
2001
-
[20]
Bregman, A. S. (1994). Auditory scene analysis: The perceptual organization of sound (MIT press)
1994
-
[21]
Broadbent, D. E. (2013). Perception and Communication (Elsevier)
2013
-
[22]
Brungart, D. S. (2001). Informational and energetic masking effects in the perception of two simultaneous talkers. The Journal of the Acoustical Society of America 109, 1101–1109
2001
-
[23]
and Sporns, O
Bullmore, E. and Sporns, O. (2012). The economy of brain network organization. Nature Reviews Neuroscience 13, 336–349
2012
-
[24]
Calvert, G. A. (2001). Crossmodal processing in the human brain: insights from functional neuroimaging studies. Cerebral Cortex 11, 1110–1123
2001
-
[25]
S., Langer, A., and Kaiser, J
Chan, J. S., Langer, A., and Kaiser, J. (2016). Temporal integration of multisensory stimuli in autism spectrum disorder: a predictive coding perspective. Journal of Neural Transmission 123, 917–923
2016
-
[26]
Cherry, E. C. (1953). Some experiments on the recognition of speech, with one and with two ears. The Journal of the Acoustical Society of America 25, 975–979
1953
-
[27]
ventriloquist effect
Choe, C. S., Welch, R. B., Gilford, R. M., and Juola, J. F. (1975). The “ventriloquist effect”: Visual dominance or response bias? Perception & Psychophysics 18, 55–60
1975
-
[28]
K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y
Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y. (2015). Attention-based models for speech recognition. In Advances in Neural Information Processing Systems. 577–585
2015
-
[29]
S., Senior, A., Vinyals, O., and Zisserman, A
Chung, J. S., Senior, A., Vinyals, O., and Zisserman, A. (2017). Lip reading sentences in the wild. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE), 3444–3453
2017
-
[30]
S., Yeung, N., and Kadosh, R
Clayton, M. S., Yeung, N., and Kadosh, R. C. (2015). The roles of cortical oscillations in sustained attention. Trends in Cognitive Sciences 19, 188–195
2015
-
[31]
Colflesh, G. J. and Conway, A. R. (2007). Individual differences in working memory capacity and divided attention in dichotic listening. Psychonomic Bulletin & Review 14, 699–703
2007
-
[32]
S., and Yau, J
Convento, S., Rahman, M. S., and Yau, J. M. (2018). Selective attention gates the interactive crossmodal coupling between perceptual systems. Current Biology 28, 746–752
2018
-
[33]
R., Cowan, N., and Bunting, M
Conway, A. R., Cowan, N., and Bunting, M. F. (2001). The cocktail party phenomenon revisited: The importance of working memory capacity. Psychonomic Bulletin & Review 8, 331–335
2001
-
[34]
and Shulman, G
Corbetta, M. and Shulman, G. L. (2002). Control of goal -directed and stimulus -driven attention in the brain. Nature Reviews Neuroscience 3, 201
2002
-
[35]
Dai, B., Chen, C., Long, Y., Zheng, L., Zhao, H., Bai, X., et al. (2018). Neural mechanisms for selectively tuning in to the target speaker in a naturalistic noisy situation. Nature Communications 9, 2405 Fu et al. Selective attention mechanisms and modeling 22
2018
-
[36]
Dai, J., Li, Y., He, K., and Sun, J. (2016). R -fcn: Object detection via region -based fully convolutional networks. In Advances in Neural Information Processing Systems. 379–387
2016
-
[37]
Das, A., Agrawal, H., Zitnick, L., Parikh, D., and Batra, D. (2017). Human att ention in visual question answering: Do humans and deep networks look at the same regions? Computer Vision and Image Understanding 163, 90–100 Da´vila-Chaco´n, J., Liu, J., and Wermter, S. (2018). E...
2017
-
[38]
E., Neal, R
Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S. (1995). The Helmholtz machine. Neural Computation 7, 889–904
1995
-
[39]
and Duncan, J
Desimone, R. and Duncan, J. (1995). Neural mechanisms of selective visual attention. Annual Review of Neuroscience 18, 193–222
1995
-
[40]
Diehl, M. M. and Romanski, L. M. (2014). Responses of prefrontal multisensory neurons to mismatching faces and vocalizations. Journal of Neuroscience 34, 11233–11243
2014
-
[41]
and Simon, J
Ding, N. and Simon, J. Z. (2012). Emergence of neural encoding of auditory objects while listening to competing speakers. Proceedings of the National Academy of Sciences 109, 11854–11859
2012
-
[42]
Dipoppa, M., Szwed, M., and Gutkin, B. S. (2016). Controlling working memory operations by selective gating: the roles of oscillations and synchrony. Advances in Cognitive Psychology 12, 209–232
2016
-
[43]
J., Killinger, M
Dorkenwald, S., Schubert, P. J., Killinger, M. F., Urban, G., Mikula, S., Svara, F., et al. (2017). Automated synaptic connectivity inference for volume electron microscopy. Nature Methods 14, 435–442
2017
-
[44]
cocktail-party problem
Du, Y., Kong, L., Wang, Q., Wu, X., and Li, L. (2011). Auditory frequency -following response: a neurophysiological measure for studying the “cocktail-party problem”. Neuroscience & Biobehavioral Reviews 35, 2046–2057
2011
-
[45]
A., and Downar, J
Dunlop, K., Hanlon, C. A., and Downar, J. (2017). Noninvasive brain stimulation treatments for addiction and major depression. Annals of the New York Academy of Sciences 1394, 31–54
2017
-
[46]
and Roig, G
Dwivedi, K. and Roig, G. (2019). Representation similarity analysis for efficient task taxonomy & transfer learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12387–12396
2019
-
[47]
and Driver, J
Eimer, M. and Driver, J. (2001). Crossmodal links in endogenous and exogenous spatial attention: evidence from event-related brain potential studies. Neuroscience & Biobehavioral Reviews 25, 497–511
2001
-
[48]
Fan, J. (2014). An information theory account of cognitive control. Frontiers in Human Neuroscience 8, 680
2014
-
[49]
D., Fossella, J., Flombaum, J
Fan, J., McCandliss, B. D., Fossella, J., Flombaum, J. I., and Posner, M. I. (2005). The activation of attentional networks. Neuroimage 26, 471–479
2005
-
[50]
D., Sommer, T., Raz, A., and Posner, M
Fan, J., McCandliss, B. D., Sommer, T., Raz, A., and Posner, M. I. (2002). Testing the efficiency and independence of attentional networks. Journal of Cognitive Neuroscience 14, 340–347
2002
-
[51]
and Posner, M
Fan, J. and Posner, M. (2004). Human attentional networks. Psychiatrische Praxis 31, 210–214
2004
-
[52]
T., and Lee, B.-S
Fang, Y., Lin, W., Lau, C. T., and Lee, B.-S. (2011). A visual attention model combining top -down and bottom-up mechanisms for salient object detection. In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE), 1293–1296
2011
-
[53]
J., Wong, A
Farah, M. J., Wong, A. B., Monheit, M. A., and Morrow, L. A. (1989). Parietal lobe mechanisms of spatial attention: Modality-specific or supramodal? Neuropsychologia 27, 461–470
1989
-
[54]
and Friston, K
Feldman, H. and Friston, K. (2010). Attention, uncertainty, and free -energy. Frontiers in Human Neuroscience 4, 215 Fu et al. Selective attention mechanisms and modeling 23
2010
-
[55]
L., Remington, R
Folk, C. L., Remington, R. W., and Johnston, J. C. (1992). Involuntary covert orienting is contingent on attentional control settings. Journal of Experimental Psychology: Human Perception and Performance 18, 1030–1044
1992
-
[56]
Frintrop, S., Rome, E., and Christensen, H. I. (2010). Computational visual attention systems and their cognitive foundations: A survey. ACM Transactions on Applied Perception (TAP) 7, 6
2010
-
[57]
Friston, K. (2009). The free-energy principle: a rough guide to the brain? Trends in Cognitive Sciences 13, 293–301
2009
-
[58]
I., Wu, H., Magg, S., Liu, X., et al
Fu, D., Barros, P., Parisi, G. I., Wu, H., Magg, S., Liu, X., et al. (2018). Assessing the contribution of semantic congruency to multisensory integration and conflict resolution. In IROS 2018 Workshop on Crossmodal Learning for Intelligent Robotics
2018
-
[59]
Gao, G., Lauri, M., Zhang, J., and Frintrop, S. (2017a). Saliency -guided adaptive seeding for supervoxel segmentation. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE), 4938–4943
2017
-
[60]
Gao, L., Guo, Z., Zhang , H., Xu, X., and Shen, H. T. (2017b). Video captioning with attention -based LSTM and semantic consistency. IEEE Transactions on Multimedia 19, 2045–2055
2017
-
[61]
J., and Luck, S
Gaspelin, N., Leonard, C. J., and Luck, S. J. (2015). Direct evidence for active suppression of salient-but- irrelevant sensory inputs. Psychological Science 26, 1740–1750
2015
-
[62]
J., and Luck, S
Gaspelin, N., Leonard, C. J., and Luck, S. J. (2017). Suppression of overt attentional capture by salient-but- irrelevant color singletons. Attention, Perception, & Psychophysics 79, 45–62
2017
-
[63]
cocktail party
Golumbic, E. M. Z., Ding, N., Bickel, S., Lakatos, P., Schevon, C. A., McKhann, G. M., et al. (2013). Mechanisms underlying selective neuronal tracking of attended speech at a “cocktail party”. Neuron 77, 980–991
2013
-
[64]
M., Swets, J
Green, D. M., Swets, J. A., et al. (1966). Signal Detection Theory and Psychophysics, vol. 1 (Wiley New York)
1966
-
[65]
M., Goffart, L., and Krauzlis, R
Hafed, Z. M., Goffart, L., and Krauzlis, R. J. (2008). Superior colliculus inactivation causes stable offsets in eye position during tracking. Journal of Neuroscience 28, 8124–8137
2008
-
[66]
Hafed, Z. M. and Krauzlis, R. J. (2008). Goal representations dominate superior colliculus activity during extrafoveal tracking. Journal of Neuroscience 28, 9426–9439
2008
-
[67]
B., Weber, C., Kerzel, M., and Wermter, S
Hafez, M. B., Weber, C., Kerzel, M., and Wermter, S. (2019). Deep intrinsically motivated continuous actor-critic for efficient robotic visuomotor skill learning. Paladyn, Journal of Behavioral Robotics 10, 14–29 Ha¨kkinen, S. and Rinne, T. (2018). Intrinsic, stimulus-driven a...
2019
-
[68]
R., and Hanson, S
Hanson, C., Caglar, L. R., and Hanson, S. J. (2018). Attentional bias in human category learning: The case of deep learning. Frontiers in Psychology 9, 374–384
2018
-
[69]
-Y., Tuzel, O., and Farahmand, A
Hara, K., Liu, M. -Y., Tuzel, O., and Farahmand, A. -M. (2017). Attentional network for visual object detection. arXiv preprint arXiv:1702.01478
2017 arXiv
-
[70]
Henderson, J. M. and Hayes, T. R. (2017). Meaning-based guidance of attention in scenes as revealed by meaning maps. Nature Human Behaviour 1, 743–747
2017
-
[71]
and Amedi, A
Hertz, U. and Amedi, A. (2014). Flexibility and stability in sensory processing revealed using visual -to- auditory sensory substitution. Cerebral Cortex 25, 2049–2064
2014
-
[72]
C., McLaughlin, S
Higgins, N. C., McLaughlin, S. A., Rinne, T., and Stecker, G. C. (2017). Evidence for cue -independent spatial representation in the human auditory cortex during active listening. Proceedings of the National Academy of Sciences 114, E7602–E7611 Fu et al. Selective attention me...
2017
-
[73]
Hinz, T., Heinrich, S., and Wermter, S. (2019). Generating multiple objects at spatially distinct locations. In International Conference on Learning Representations (ICLR)
2019
-
[74]
M., Kahng, M., Pienta, R., and Chau, D
Hohman, F. M., Kahng, M., Pienta, R., and Chau, D. H. (2018). Visual analytics in deep learning: An interrogative survey for the next frontiers. IEEE Transactions on Visualization and Computer Graphics 25, 2674–2693
2018
-
[75]
and Baldi, P
Itti, L. and Baldi, P. (2009). Bayesian surprise attracts human attention. Vision Research 49, 1295–1306
2009
-
[76]
and Koch, C
Itti, L. and Koch, C. (2000). A saliency-based search mechanism for overt and covert shifts of visual attention. Vision Research 40, 1489–1506
2000
-
[77]
Itti, L., Koch, C., and Niebur, E. (1998). A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis & Machine Intelligence 11, 1254–1259
1998
-
[78]
and Mazaheri, A
Jensen, O. and Mazaheri, A. (2010). Shaping functional architecture by oscillator y alpha activity: gating by inhibition. Frontiers in Human Neuroscience 4, 186
2010
-
[79]
Jetley, S., Murray, N., and Vig , E. (2016). End -to-end saliency mapping via probability distribution prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5753–5761
2016
-
[80]
A., Robertson, I
Johnson, K. A., Robertson, I. H., Barry, E., Mulligan, A., Daibhis, A., Daly, M., et al. (2008). Impaired conflict resolution and alerting in children with ADHD: evidence from the attention network task (ANT). Journal of Child Psychology and Psychiatry 49, 1339–1347
2008
-
[81]
and Narayanan, S
Kalinli, O. and Narayanan, S. S. (2007). A saliency -based auditory attention model with applications to unsupervised prominent syllable detection in speech. In Eighth Annual Conference of the International Speech Communication Association
2007
-
[82]
Kaya, E. M. and Elhilali , M. (2017). Modelling auditory attention. Philosophical Transactions of the Royal Society B: Biological Sciences 372, 20160101
2017
-
[83]
I., Lippert, M., and Logothetis, N
Kayser, C., Petkov, C. I., Lippert, M., and Logothetis, N. K. (2005). Mechanisms for allocating auditory attention: an auditory saliency map. Current Biology 15, 1943–1947
2005
-
[84]
Khaligh-Razavi, S.-M., Henriksson, L., Kay, K., and Kriegeskorte, N. (2017). Fixed versus mixed RSA: Explaining visual representations by fixed and mixed feature sets from shallow and deep computational models. Journal of Mathematical Psychology 76, 184–197
2017
-
[85]
Klein, D. A. and Frintrop, S. (2011). Center -surround divergence of feature statistics for salient object detection. In 2011 International Conference on Computer Vision (IEEE), 2214–2219
2011
-
[86]
J., Lovejoy, L
Krauzlis, R. J., Lovejoy, L. P., and Ze´non, A. (2013). Superior colliculus and visual spatial attention. Annual Review of Neuroscience 36, 165–182
2013
-
[87]
Kriegeskorte, N., Mur, M., and Bandettini, P. A. (2008). Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience 2, 4
2008
-
[88]
S., Ayush, K., and Babu, R
Kruthiventi, S. S., Ayush, K., and Babu, R. V. (2017). Deepfix: A fully convolutional neural network for predicting human eye fixations. IEEE Transactions on Image Processing 26, 4446–4456
2017
-
[89]
E., Vaden Jr, K
Kuchinsky, S. E., Vaden Jr, K. I., Keren, N. I., Harris, K. C., Ahlstrom, J. B., Dubno, J. R., et al. (2012). Word intelligibility and age predict visual cortex activity during word listening. Cerebral Cortex 22, 1360–1371
2012
-
[90]
V., Atkinson, J., and Braddick, O
Kulke, L. V., Atkinson, J., and Braddick, O. (2016). Neural differences between covert and overt attention studied using EEG with simultaneous remote eye tracking. Frontiers in Human Neuroscience 10, 592
2016
-
[91]
S., Gatys, L
Kummerer, M., Wallis, T. S., Gatys, L. A., and Bethge, M. (2017). Understanding low -and high-level contributions to fixation prediction. In Proceedings of the IEEE International Conference on Computer Vision. 4789–4798 Fu et al. Selective attention mechanisms and modeling 25
2017
-
[92]
Lahat, D., Adali, T., and Jutten, C. (2015). Multimodal data fusion: an overview of methods, challenges, and prospects. Proceedings of the IEEE 103, 1449–1477
2015
-
[93]
K., Larson, E., Maddox, R
Lee, A. K., Larson, E., Maddox, R. K., and Shinn -Cunningham, B. G. (2014). Using neuroimaging to understand the cortical mechanisms of auditory selective attention. Hearing Research 307, 111–120
2014
-
[94]
and Choo, H
Lee, K. and Choo, H. (2013). A critical review of selective attention: an interdisciplinary perspective. Artificial Intelligence Review 40, 27–50
2013
-
[95]
and Getzmann, S
Lewald, J. and Getzmann, S. (2015). Electrophysiological correlates of cocktail-party listening. Behavioural Brain Research 292, 157–166
2015
-
[96]
Li, G., Gan, Y., Wu, H., Xiao, N., and Lin, L. (2019). Cross-modal attentional context learning for rgb-d object detection. IEEE Transactions on Image Processing 28, 1591–1601
2019
-
[97]
Li, W., Yuan, Z., Fang, X., and Wang, C. (2018). Knowing where to look? Analysis on attention of visual question answering system. In Proceedings of the European Conference on Computer Vision (ECCV). 1–8
2018
-
[98]
Li, Z. (1999). Contextual influences in V1 as a basis for pop out and asymmetry in visual search. Proceedings of the National Academy of Sciences 96, 10530–10535
1999
-
[99]
Li, Z. (2002). A saliency map in primary visual cortex. Trends in Cognitive Sciences 6, 9–16
2002
-
[100]
Lidestam, B., Holgersson, J., and Moradi, S. (2014). Comparison of informational vs. energetic masking effects on speechreading performance. Frontiers in Psychology 5, 639
2014
-
[101]
and Milanova, M
Liu, X. and Milanova, M. (2018). Visual attention in deep learning: a review. International Robotics and Automation Journal 4, 154–155
2018
-
[102]
F., and Sadjedi, H
Lotfi, Y., Mehrkian, S., Moossavi, A., Z adeh, S. F., and Sadjedi, H. (2016). Relation between working memory capacity and auditory stream segregation in children with auditory processing disorder. Iranian Journal of Medical Sciences 41, 110–117
2016
-
[103]
Lowe, D. G. et al. (1999). Object recognition from local scale -invariant features. In International Conference on Computer Vision. vol. 99, 1150–1157
1999
-
[104]
J., Fritz, J
Lu, K., Xu, Y., Yin, P., Oxenham, A. J., Fritz, J. B., and Shamma, S. A. (2017). Temporal coherence structure rapidly shapes neuronal interactions. Nature Communications 8, 13900
2017
-
[105]
-T., Pham, H., and Manning, C
Luong, M. -T., Pham, H., and Manning, C. D. (2015). Effective approaches to attention -based neural machine translation. In Conference on Empirical Methods in Natural Language Processing
2015
-
[106]
A., and Brown, G
Ma, N., Gonzalez, J. A., and Brown, G. J. (2018). Robust binaural localization of a target sound source by combining spectral source models and deep neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing 26, 2122–2131
2018
-
[107]
Ma, W. J. (2012). Organizing probabilistic models of perception. Trends in Cognitive Sciences 16, 511–518
2012
-
[108]
Mahdi, A., Qin, J., and Crosby, G. (2019). Deepfeat: A bottom-up and top-down saliency model based on deep features of convolutional neural nets. IEEE Transactions on Cognitive and Developmental Systems
2019
-
[109]
Mai, G., Schoof, T., and Howell, P. (2019). Modulation of phase-locked neural responses to speech during different arousal states is age-dependent. NeuroImage 189, 734–744
2019
-
[110]
J., Teder-Sa¨leja¨rvi, W
Mcdonald, J. J., Teder-Sa¨leja¨rvi, W. A., Russo, F. D., and Hillyard, S. A. (2003). Neural substrates of perceptual enhancement by cross-modal spatial attention. Journal of Cognitive Neuroscience 15, 10–19
2003
-
[111]
Melloni, L., van Leeuwen, S., Alink, A., and Mu¨ ller, N. G. (2012). Interaction between bottom-up saliency and top-down control: how saliency maps are created in the human brain. Cerebral Cortex 22, 2943–2952
2012
-
[112]
L., Fink, G
Mengotti, P., Boers, F., Dombert, P. L., Fink, G. R., and Vossel, S. (2018). Integrating modality-specific expectancies for the deployment of spatial attention. Scientific Reports 8, 1210 Fu et al. Selective attention mechanisms and modeling 26
2018
-
[113]
and Uddin, L
Menon, V. and Uddin, L. Q. (2010). Saliency, switching, attention and control: a network model of insula function. Brain Structure and Function 214, 655–667
2010
-
[114]
Meredith, M. A. (2002). On the neuronal basis for multisensory convergence: a brief overview. Cognitive Brain Research 14, 31–40
2002
-
[115]
T., Bearpark, H
Michie, P. T., Bearpark, H. M., Crawford, J. M., and Glue, L. C. (1990). The nature of selective attention effects on auditory event-related potentials. Biological Psychology 30, 219–250
1990
-
[116]
Misselhorn, J., Friese, U., and Engel, A. K. (2019). Frontal and parietal alpha oscillations reflect attentional modulation of cross-modal matching. Scientific Reports 9, 5030
2019
-
[117]
and Chartier, S
Morissette, L. and Chartier, S. (2015). Saliency model of auditory attention based on frequency, amplitude and spatial location. In Proceedings of International Joint Conference on Neural Networks (IJCNN) (IEEE), 1–5
2015
-
[118]
Mounts, J. R. (2000). Attentional capture by abrupt onsets and feature singletons produces inhibitory surrounds. Perception & Psychophysics 62, 1485–1493
2000
-
[119]
Mroueh, Y., Marcheret, E., and Goel, V. (2015). Deep multimodal learning for audio -visual speech recognition. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE), 2130–2134
2015
-
[120]
K., and Whittingstall, K
Musall, S., von Pfo¨ stl, V., Rauch, A., Logothetis, N. K., and Whittingstall, K. (2012). Effects of neural synchrony on surface EEG. Cerebral Cortex 24, 1045–1053
2012
-
[121]
Oldoni, D., De Coensel, B., Boes, M., Rademaker, M., De Baets, B., Van Renterghem, T., et al. (2013). A computational model of auditory attention for use in soundscape research. The Journal of the Acoustical Society of America 134, 852–861
2013
-
[122]
I., Barros, P., Fu, D., Magg, S., Wu, H., Liu, X., et al
Parisi, G. I., Barros, P., Fu, D., Magg, S., Wu, H., Liu, X., et al. (2018). A neurorobotic experiment for crossmodal conflict resolution in complex environments. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE), 2330–2335
2018
-
[123]
I., Barros, P., Kerzel, M., Wu, H., Yang, G., Li, Z., et al
Parisi, G. I., Barros, P., Kerzel, M., Wu, H., Yang, G., Li, Z., et al. (2017). A computational model of crossmodal processing for conflict resolution. In 2017 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob) (IEEE), 33–38
2017
-
[124]
M., and Lu, Z
Peng, Y., Yan, K., Sandfort, V., Summers, R. M., and Lu, Z. (2019). A self-attention based deep learning method for lesion attribute detection from CT reports. arXiv preprint arXiv:1904.13018
2019 arXiv
-
[125]
and Adolphs, R
Pessoa, L. and Adolphs, R. (2010). Emotion processing and the amygdala: from a’low road’to’many roads’ of evaluating biological significance. Nature Reviews Neuroscience 11, 773
2010
-
[126]
S., Maroy, R., and Bottlaender, M
Picard, F., Sadaghiani , S., Leroy, C., Courvoisier, D. S., Maroy, R., and Bottlaender, M. (2013). High density of nicotinic receptors in the cingulo-insular network. Neuroimage 79, 42–51
2013
-
[127]
A., and Poliakoff, E
Poole, D., Gowen, E., Warren, P. A., and Poliakoff, E. (2018). Visual-tactile selective attention in autism spectrum condition: An increased influence of visual distractors. Journal of Experimental Psychology: General 147, 1309
2018
-
[128]
and Rothbart, M
Posner, M. and Rothbart, M. (1998). Attention, self –regulation and consciousness. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 353, 1915–1927
1998
-
[129]
and Snyder, C
Posner, M. and Snyder, C. (1975). Attention and cognitive control In RL Solso,(Ed), Information processing and cognition (pp 55-85) Hillsdale (NJ Erlbaum)
1975
-
[130]
Posner, M. I. (1980). Orienting of attention. Quarterly Journal of Experimental Psychology 32, 3–25
1980
-
[131]
Posner, M. I. and Cohen, Y. (1984). Components of visual orienting. Attention and Performance X: Control of Language Processes 32, 531–556
1984
-
[132]
and Taylor, G
Ramachandram, D. and Taylor, G. W. (2017). Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Processing Magazine 34, 96–108 Fu et al. Selective attention mechanisms and modeling 27
2017
-
[133]
and Farhadi, A
Redmon, J. and Farhadi, A. (2017). Yolo9000: better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7263–7271
2017
-
[134]
Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r -cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems. 91–99
2015
-
[135]
and Palmer, S
Rock, I. and Palmer, S. (1990). The legacy of gestalt psychology. Scientific American 263, 84–91
1990
-
[136]
Roseboom, W., Kawabe, T., and Nishida, S. (2013). The cross-modal double flash illusion depends on featural similarity between cross-modal inducers. Scientific Reports 3, 3437
2013
-
[137]
and D’Esposito, M
Sadaghiani, S. and D’Esposito, M. (2014). Functional characterization of the cingulo-opercular network in the maintenance of tonic alertness. Cerebral Cortex 25, 2763–2773
2014
-
[138]
and Luck, S
Sawaki, R. and Luck, S. J. (2010). Capture versus suppression of attention by salient singletons: Electrophysiological evidence for an automatic attend -to-me signal. Attention, Perception, & Psychophysics 72, 1455–1470
2010
-
[139]
and Gutschalk, A
Schadwinkel, S. and Gutschalk, A. (2010). Activity associated with stream segregation in human auditory cortex is similar for spatial and pitch cues. Cerebral Cortex 20, 2863–2873
2010
-
[140]
K., Rosen, S., Wickham, L., and Wise, R
Scott, S. K., Rosen, S., Wickham, L., and Wise, R. J. (2004). A positron emission tomography study of the neural basis of informational and energetic masking effects in speech perception. The Journal of the Acoustical Society of America 115, 813–821
2004
-
[141]
R., Foxe, J
Senkowski, D., Schneider, T. R., Foxe, J. J., and Engel, A. K. (2008). Crossmodal binding through neural coherence: implications for multisensory processing. Trends in Neurosciences 31, 401–409
2008
-
[142]
and Kim, R
Shams, L. and Kim, R. (2010). Crossmodal influences on visual perception. Physics of Life Reviews 7, 269–284
2010
-
[143]
Shannon, C. E. (1948). A mathematical theory of com munication. Bell System Technical Journal 27, 379–423
1948
-
[144]
Shi, J., Xu, J., Liu, G., Xu, B., et al. (2018). Listen, think and listen again: capturing top -down auditory attention for speaker-independent speech separation. In Proceedings of the International Joint Conference on Artificial Intelligence. 4353–4360
2018
-
[145]
Shinn-Cunningham, B. G. (2008). Object-based auditory and visual attention. Trends in Cognitive Sciences 12, 182–186
2008
-
[146]
and Zisserman, A
Simonyan, K. and Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR)
2015
-
[147]
Skocaj, D., Leonardis, A., and Kruijff, G. -J. M. (2012). Cross-Modal Learning (Boston, MA: Springer US)
2012
-
[148]
Sloutsky, V. M. (2003). The role of similarity in the deve lopment of categorization. Trends in Cognitive Sciences 7, 246–251
2003
-
[149]
T., Rorden, C., and Jackson, S
Smith, D. T., Rorden, C., and Jackson, S. R. (2004). Exogenous orienting of attention depends upon the ability to execute eye movements. Current Biology 14, 792–795
2004
-
[150]
-H., Jeong, H.-W., Choi, I., Jeong, D., Kim, K., et al
Song, Y.-H., Kim, J. -H., Jeong, H.-W., Choi, I., Jeong, D., Kim, K., et al. (2017). A neural circuit for auditory dominance over visual perception. Neuron 93, 940–954
2017
-
[151]
Stein, B. E. and Stanford, T. R. (2008). Multisensory integration: current issues from the perspective of the single neuron. Nature Reviews Neuroscience 9, 255
2008
-
[152]
E., Wallace, M
Stein, B. E., Wallace, M. W., Stanford, T. R., and Jiang, W. (2002). Book review: Cortex governs multisensory integration in the midbrain. The Neuroscientist 8, 306–314 Strauß , A., Wo¨stmann, M., and Obleser, J. (2014). Cortical alpha oscillations as a tool for auditory selec...
2002
-
[153]
Styles, E. (2006). The psychology of attention (Psychology Press) Fu et al. Selective attention mechanisms and modeling 28
2006
-
[154]
Swets, J. A. (2014). Signal detection theory and ROC analysis in psychology and diagnostics: Collected papers (Psychology Press)
2014
-
[155]
Talsma, D., Senkowski, D., Soto -Faraco, S., and Woldorff, M. G. (2010). The multifaceted interplay between attention and multisensory integration. Trends in Cognitive Sciences 14, 400–410
2010
-
[156]
Theeuwes, J. (1991). Exogenous and endogenous control of attention: The effect of visual onsets and offsets. Perception & Psychophysics 49, 83–90
1991
-
[157]
ventriloquism effect
Thurlow, W. R. and Jack, C. E. (1973). Certain determinants of the “ventriloquism effect”. Perceptual and Motor Skills 36, 1171–1184
1973
-
[158]
Todd, J. T. and Van Gelder, P. (1979). Implications of a transient–sustained dichotomy for the measurement of human performance. Journal of Experimental Psychology: Human Perception and Performance 5, 625–638
1979
-
[159]
H., and Quigley, K
Togo, F., Lange, G., Natelson, B. H., and Quigley, K. S. (2015). Attention network test: assessment of cognitive function in chronic fatigue syndrome. Journal of Neuropsychology 9, 1–9
2015
-
[160]
and Gormican, S
Treisman, A. and Gormican, S. (1988). Feature analysis in early vision: evidence from search asymmetries. Psychological Review 95, 15–48
1988
-
[161]
Uddin, L. Q. (2015). Sa lience processing and insular cortical function and dysfunction. Nature Reviews Neuroscience 16, 55
2015
-
[162]
Uddin, L. Q. and Menon, V. (2009). The anterior insula in autism: under-connected and under-examined. Neuroscience & Biobehavioral Reviews 33, 1198–1203
2009
-
[163]
Urbanek, C., Weinges-Evers, N., Bellmann-Strobl, J., Bock, M., Do¨rr, J., Hahn, E., et al. (2010). Attention network test reveals alerting network dysfunction in multiple sclerosis. Multiple Sclerosis Journal 16, 93–99
2010
-
[164]
VanRullen, R. (2003). Visual saliency and spike timing in the ventral visual pathway. Journal of Physiology-Paris 97, 365–377
2003
-
[165]
Varela, F., Lachaux, J.-P., Rodriguez, E., and Martinerie, J. (2001). The brainweb: phase synchronization and large-scale integration. Nature Reviews Neuroscience 2, 229
2001
-
[166]
N., et al
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems. 5998–6008
2017
-
[167]
M., and Yoshida, M
Veale, R., Hafed, Z. M., and Yoshida, M. (2017). How is visual salience computed in the brain? insights from behaviour, neurobiology and modelling. Philosophical Transactions of the Royal Society B: Biological Sciences 372, 20160113
2017
-
[168]
Veen, V. v. and Carter, C. S. (2006). C onflict and cognitive control in the brain. Current Directions in Psychological Science 15, 237–240
2006
-
[169]
and Pelli, D
Verghese, P. and Pelli, D. G. (1992). The information capacity of visual attention. Vision Research 32, 983–995
1992
-
[170]
Vuilleumier, P. (2005). How brains beware: neural mechanisms of emotional attention. Trends in Cognitive Sciences 9, 585–594
2005
-
[171]
T., Meredith, M
Wallace, M. T., Meredith, M. A., and Stein, B. E. (1998). Multisensory integration in the superior colliculus of the alert cat. Journal of Neurophysiology 80, 1006–1010
1998
-
[172]
Wang, B., Yang, Y., Xu, X., Hanjalic, A., and Shen, H. T. (2017). Adversarial cross-modal retrieval. In Proceedings of the 25th ACM International Conference on Multimedia (ACM), 154–162
2017
-
[173]
and Chang, P
Wang, D. and Chang, P. (2008). An oscillatory correlation model of auditory streaming. Cognitive Neurodynamics 2, 7–19
2008
-
[174]
and Terman, D
Wang, D. and Terman, D. (1995). Locally excitatory glob ally inhibitory oscillator networks. IEEE Transactions on Neural Networks 6, 283–286 Fu et al. Selective attention mechanisms and modeling 29
1995
-
[175]
Wang, X.-J. (2010). Neurophysiological and computational principles of cortical rhythms in cognition. Physiological Reviews 90, 1195–1268
2010
-
[176]
Wang, Y., Huang, M., Zhao, L., et al. (2016). Attention -based LSTM for aspect -level sentiment classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 606–615
2016
-
[177]
compellingness
Warren, D. H., Welch, R. B., and McCarthy, T. J. (1981). The role of visual-auditory “compellingness” in the ventriloquism effect: Implications for transitivity among the spatial senses. Perception & Psychophysics 30, 557–564
1981
-
[178]
H., Gopalakrishnan, A., Hazlett, C., and Woldorff, M
Weissman, D. H., Gopalakrishnan, A., Hazlett, C., and Woldorff, M. (2004). Dorsal anterior cingulate cortex resolves conflict from distracting stimuli by boosting attention toward relevant events. Cerebral Cortex 15, 229–237
2004
-
[179]
Welch, R. B. and Warren, D. H. (1980). Immediate perceptual response to intersensory discrepancy. Psychological Bulletin 88, 638–667
1980
-
[180]
J., Berg, D
White, B. J., Berg, D. J., Kan, J. Y., Marino, R. A., Itti, L., and Munoz, D. P. (2017). Superior colliculus neurons encode a visual saliency map during free viewing of natural dynamic video. Nature Communications 8, 14263
2017
-
[181]
G., Gallen, C
Woldorff, M. G., Gallen, C. C., Hampson, S. A., Hillyard, S. A., Pantev, C., Sobel, D., et al. (1993). Modulation of early sensory processing in human auditory cortex during auditory selective attention. Proceedings of the National Academy of Sciences 90, 8722–8726 Wo¨ stmann,...
1993
-
[182]
Wrigley, S. N. and Brown, G. J. (2004). A computational model of auditory selective attention. IEEE Transactions on Neural Networks 15, 1151–1163
2004
-
[183]
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., et al. (2015). Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of International Conference on Machine Learning. 2048–2057
2015
-
[184]
Yang, G., Nan, W., Zheng, Y., Wu, H., Li, Q., and Liu, X. (2017). Distinct cognitive control mechanisms as revealed by modality-specific conflict adaptation effects. Journal of Experimental Psychology: Human Perception and Performance 43, 807
2017
-
[185]
and Jonides, J
Yantis, S. and Jonides, J. (1984). Abrupt visual onsets and selective attention: evidence from visual search. Journal of Experimental Psychology: Human Perception and Performance 10, 601
1984
-
[186]
Yao, X., Han, J., Zhang, D., and Nie, F. (2017). Revisiting co-saliency detection: A novel approach based on two-stage multi-view spectral rotation co -clustering. IEEE Transactions on Image Processing 26, 3196–3209
2017
-
[187]
B., and Shamma, S
Yin, P., Fritz, J. B., and Shamma, S. A. (2014). Rapid spectrotemporal plasticity in primary auditory cortex during behavior. Journal of Neuroscience 34, 4396–4408
2014
-
[188]
and Gong, Q
Zhang, X. and Gong, Q. (2019). Frequency-following responses to complex tones at different frequencies reflect different source configurations. Frontiers in Neuroscience 13, 130
2019
-
[189]
G., and Papani colaou, A
Zouridakis, G., Simos, P. G., and Papani colaou, A. C. (1998). Multiple bilaterally asymmetric cortical sources account for the auditory N1m component. Brain Topography 10, 183–189
1998
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.