Pith. sign in

REVIEW 5 major objections 6 minor 189 references

What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective

T0 review · 5 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper argues that an integrated framework combining visual, auditory, and audiovisual selective attention can bridge human behavioral and neural findings with computational simulation.

desk verdict A competent, useful review of selective attention with a smart audiovisual frame, but its Section 5.2 overstates what computational models have actually demonstrated. read the letter →

arxiv 1909.05654 v1 pith:VUC3GKVU submitted 2019-09-05 cs.CV cs.LG

classification cs.CVcs.LG
keywords selectiveattentioncrossmodallearningaudiovisualintegrationvisualauditorycocktailpartyeffectventriloquismdeep
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that selective attention, which psychology and neuroscience have mostly studied one sensory modality at a time, is better understood and better transferred to machines when vision, audition, and their interactions are treated as one integrated system. It argues that comparing the visual pop-out effect, the auditory cocktail party effect, and audiovisual conflict resolution reveals a common logic: attention allocates processing weight by priority, driven by bottom-up salience, top-down goals, and past selection history. If the paper is right, computational models can use these human findings as a roadmap for more capable audiovisual agents, and psychologists can use those models to test and refine attention theories.

What carries the argument

The organizing device is crossmodal integration and conflict resolution: the brain binds sights and sounds that plausibly come from one source, resolves mismatches by weighting the modality with higher reliability, and uses attention to gate the coupling between sensory areas. On the computational side, the load-bearing mechanisms are saliency maps with winner-take-all selection, locally excitatory and globally inhibitory oscillator networks for auditory stream segregation, and deep attention architectures with top-down prediction and conflict-monitoring modules, all of which the paper connects to neural mechanisms such as gamma-band enhancement and alpha-band inhibition.

What would settle it

Take an audiovisual model built on the paper's integrated mechanisms and compare it with a simple fusion baseline in a naturalistic incongruent scene, such as a speaker's voice arriving from a different location than the visible lip movements; if the human-derived crossmodal mechanisms never improve localization or conflict resolution beyond the baseline, the central bridge claim is not supported.

Watch

Extended reading notes

Core claim

The central claim is that the similarities and differences among selective attention mechanisms across modalities are exactly what an integrated framework needs to capture, and that such a framework can bridge human behavioral and neural patterns with intelligent system simulation. The paper assembles evidence that visual attention, auditory attention, and audiovisual integration all involve a common weight-allocation logic, but express it through different routes: saliency maps and winner-take-all selection in vision, stream segregation and top-down prediction in hearing, and modality-appropriate weighting and conflict monitoring when sights and sounds compete. It then maps these findings onto computational models, from classical saliency and oscillator networks to deep learning attention mechanisms, and argues that crossmodal modeling is not an optional extra but a necessary step for real-world robotics and human-robot interaction.

Load-bearing premise

The review's roadmap depends on the assumption that laboratory effects such as pop-out visual search, listening tasks with different sounds in each ear, and ventriloquism survive in rich real-world scenes and in autonomous agents; if those effects are task-specific, the proposed bridge between human attention and computational modeling weakens.

Editorial extensions

If this is right

  • If the integrated framework is correct, an audiovisual model should reproduce the ventriloquism effect: visual lip movements should bias sound localization more strongly when the auditory cue is unreliable.
  • Cocktail-party findings imply that top-down prediction of a target voice, not just bottom-up stream separation, should improve speech separation in multi-speaker noise.
  • Neural oscillation findings suggest that computational models could benefit from gating mechanisms analogous to alpha suppression of task-irrelevant inputs and gamma enhancement of task-relevant inputs.
  • A priority map that includes selection history and semantic meaning should outperform maps based only on physical salience when predicting where humans look in real scenes.
  • Co-saliency and meaning-map approaches could be combined to improve image and video interpretation by prioritizing the most informative content for humans.
  • Robots that resolve crossmodal conflicts by choosing the more reliable modality should localize sounds and recognize events more accurately in naturalistic environments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper, the reviewed findings suggest that modern self-attention architectures could be made more human-like by adding explicit modality-reliability weighting, so that auditory temporal precision or visual spatial precision wins depending on the situation.
  • A testable extension would be to train a deep network on audiovisual conflict data with and without a conflict-monitoring module, and to compare its error patterns directly with human performance on incongruent speaker-lip movement scenes.
  • The paper's emphasis on conflict-driven curiosity implies a developmental-robotics prediction: agents that treat crossmodal mismatches as learning signals, rather than errors to discard, should acquire new multimodal concepts faster.
  • An implicit consequence for experimental psychology is that computational models could serve as falsifiable implementations of attention theories, making theoretical disagreements such as stimulus-driven versus goal-driven capture testable in simulation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This review article surveys selective attention from psychology, neuroscience, and computational modeling, focusing on visual 'pop-out', auditory 'cocktail party', and audiovisual crossmodal integration/conflict resolution. It reviews theories, behavioral and neural mechanisms, and computational models for each, then discusses gaps and future directions. The paper's stated aims are to integrate unimodal and crossmodal selective attention findings and to bridge human behavioral/neural patterns with intelligent-system simulation, particularly in robotics.

Significance. If fully substantiated, the review would provide a useful interdisciplinary map and a roadmap for transferring psychological findings into computational models. As a compilation it is strong: it covers classic theories (top-down/bottom-up, priority maps, neural oscillations, free-energy), summarizes a broad literature, and organizes it around three well-chosen representative effects. It also explicitly acknowledges several limitations of the human literature (e.g., correlational neural evidence, Section 2.3). The central gap is the computational side of the claimed bridge: the key evidence in Section 5.2 is self-cited and not quantitatively validated against human data, so the roadmap is better described as a research program than a demonstrated bridge. The paper's explicit, falsifiable roadmap and clear admission of loose connections between psychology and computer science are assets; the missing cross-validation is the main obstacle to accepting the strongest claims.

major comments (5)
  1. [5.2] The paragraph beginning 'Many studies focus on multimodal fusion' reports human behavioral experiments (Parisi et al. 2017, 2018; Fu et al. 2018) and then states that 'human-like responses were modelled' and that the 'work above shows that DL can simulate humans' selective attention and conflict resolution.' However, no quantitative comparison is provided between model outputs and the human psychophysical results described in the same paragraph: there are no effect sizes for the congruence effect in the model, no statistical test against human performance, and no ablation isolating the contribution of the attention/crossmodal layer. Since the paper's central claim is that computational models can bridge human behavioral/neural patterns, this evidence gap is load-bearing. Please either add the quantitative comparison from the cited papers or soften the claim to 'a proof-of-concept simulation' and explicitly state the missing validation as a limitation.
  2. [5.2] The computational crossmodal review is based almost entirely on the authors' own prior work (Parisi et al. 2017, 2018; Fu et al. 2018; Barros et al. 2018). No independent work implementing audiovisual selective attention with quantitative validation is cited. This is not a fatal flaw for a review, but it means the general roadmap rests on a single line of evidence. Please either survey independent models (e.g., audiovisual saliency or multisensory integration models outside the authors' group) or explicitly frame the section as a case study from the authors' lab rather than a representative field-wide survey.
  3. [1, 5.1, 6] The Introduction and Section 5.1 set up a bridge between laboratory selective-attention effects (pop-out, cocktail party, ventriloquism) and autonomous agents, and Section 6 makes robotics recommendations based on this bridge. However, the review never addresses whether these simplified laboratory effects survive real-world complexity (e.g., moving cameras, reverberant audio, task-relevant semantic context). Section 2.5 itself concedes that 'the connection between computer science models and psychology is still loose and broad.' Please add a discussion of ecological validity for each of the three representative effects, or explicitly state that the transfer is an open research question.
  4. [2.3 vs 5.1] Section 2.3 correctly concedes that the neural-oscillation evidence is 'mainly correlations and descriptive results rather than causal relationships.' Yet Section 5.1 later states that the gamma-alpha oscillation pattern 'is proposed to be the information gating mechanism,' and the summary attributes this mechanism without repeating the correlational caveat. Please carry the same epistemic qualifier through to Section 5.1, or clearly distinguish correlational findings from mechanistic claims.
  5. [Abstract, Introduction, 3, 4, 5] The Abstract and Introduction promise an 'integrated framework' that 'combine[s] and compare[s] selective attention mechanisms from different modalities.' In practice, Sections 3 and 4 are parallel reviews with separate computational-model subsections, and Section 5 is largely independent; cross-cutting comparisons appear only in each section's closing paragraphs. A reader looking for an explicit side-by-side comparison of visual and auditory attention mechanisms (e.g., a table of shared and differing computational principles) will not find it. Please add such a comparison or temper the claim of integration.
minor comments (6)
  1. [2.5] The abbreviation 'LTSM' should be 'LSTM' (Long Short-Term Memory).
  2. [3.1] The phrase 'super colliculus' should be 'superior colliculus'.
  3. [6] The text reads 'One the one hand' and should read 'On the one hand'.
  4. [6] The phrase 'the -state-of-the-art approaches' contains a stray hyphen and should be 'the state-of-the-art approaches'.
  5. [Figures] Figures 1, 2, and 4 are explicitly labeled 'adapted from' previous publications; please verify that all required permissions for reuse have been obtained for the final version.
  6. [5.2] The iCub robot is mentioned without a citation; please add a reference to the iCub platform (e.g., Metta et al., 2008) or clarify the source.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the review is descriptive and rests on independent empirical literature; its concentrated self-citations in Section 5.2 raise validation rather than circularity concerns.

full rationale

This article is a narrative review rather than a derivation, so the circularity patterns do not apply. The psychological and neural sections are grounded in independent experimental literature (e.g., Cherry 1953; Desimone and Duncan 1995; Corbetta and Shulman 2002; Talsma et al. 2010), and the paper explicitly acknowledges open problems rather than claiming a closed formal result. Section 5.2 does rely heavily on the authors' own prior computational work (Parisi et al. 2017, 2018; Fu et al. 2018; Barros et al. 2018) to support the statement that deep learning can simulate crossmodal conflict resolution. That is a self-citation cluster and a validation/overclaim concern: the review does not provide quantitative model-human comparisons, and the concluding sentence in Section 5.2 exceeds what is demonstrated. However, it is not circular: the cited experiments are empirical studies with human participants and models, not definitions of the conclusion in terms of the premises. No equation is fitted and renamed a prediction, no uniqueness theorem or ansatz is imported from the authors' prior work, and no target quantity is defined in terms of an input quantity. The paper even concedes that 'research about selective attention and conflict resolution in computer science is limited' and that the connection between computer science models and psychology is 'still loose and broad.' Therefore no step in the paper reduces by construction to its own input; the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities appear because the paper is a literature review with no new model or fitting procedure. The two listed axioms are background domain assumptions that the review relies on for its cross-modal transfer arguments and its interpretation of neural measures.

assumptions (2)
  • domain assumption Laboratory findings such as pop-out, dichotic listening, and ventriloquism generalize to real-world scenarios and to computational agents.
    This transferability is the premise of the review's robotics recommendations and future directions, asserted in Sections 1 and 6 rather than demonstrated.
  • domain assumption Neural correlates of attention, such as ERP components, alpha oscillations, and fMRI activations, provide mechanistic evidence rather than mere correlation.
    Section 4.1 treats these signals as explanatory mechanisms, while Section 2.3 concedes that oscillation evidence is mainly correlational and descriptive.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective." pith.science (2026). https://pith.science/paper/VUC3GKVU

@misc{pith2026190905654,
  author       = {Pith},
  title        = {Pith review of: What can computational models learn from human selective attention? A review from an audiovisual crossmodal perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VUC3GKVU}},
  note         = {Machine review of arXiv:1909.05654}
}
read the original abstract

Selective attention plays an essential role in information acquisition and utilization from the environment. In the past 50 years, research on selective attention has been a central topic in cognitive science. Compared with unimodal studies, crossmodal studies are more complex but necessary to solve real-world challenges in both human experiments and computational modeling. Although an increasing number of findings on crossmodal selective attention have shed light on humans' behavioral patterns and neural underpinnings, a much better understanding is still necessary to yield the same benefit for computational intelligent agents. This article reviews studies of selective attention in unimodal visual and auditory and crossmodal audiovisual setups from the multidisciplinary perspectives of psychology and cognitive neuroscience, and evaluates different ways to simulate analogous mechanisms in computational models and robotics. We discuss the gaps between these fields in this interdisciplinary review and provide insights about how to use psychological findings and theories in artificial intelligence from different perspectives.

Figures

Figures reproduced from arXiv: 1909.05654 by the authors.

Figure 1
Figure 1. a) Neuroanatomical model of bottom-up and top-down attentional processing in the visual cortex. The dorsal system (green) executes the top-down attentional control. FEF: frontal eye field; IPS: intraparietal sulcus. The ventral system (red) executes the bottom-up processing. VFC: ventral frontal cortex; TPJ: temporoparietal junction (adapted from Corbetta and Shulman (2002)); b) Cortical oscillation model of attenti… view at source ↗
Figure 2
Figure 2. a) Auditory selective attention model with interaction between bottom-up processing and top￾down modulation. The compound sound enters the bottom-up processing in the form of segregated units and then the units are grouped into streams. After segregation and competition, foreground sound stands out from the background noise. The wider arrow represents the salient object with higher attentional weights. Top-down atte… view at source ↗
Figure 4
Figure 4. a) Visual saliency model. Features are extracted from the input image. The center-surround mechanism and normalization are used to generate the individual feature saliency maps. Finally, the saliency map is generated by a linear combination of different individual saliency maps (adapted from Itti et al. (1998)); b) Auditory saliency model. The structure of the model is similar to the visual saliency model by convert… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

189 extracted references · 80 canonical work pages

  1. [1]

    what” and “where

    Ahveninen, J., Ja¨a¨skela¨inen, I. P., Raij, T., Bonmassar, G., Devore, S., Ha¨ma¨la¨inen, M., et al. (2006). Task-modulated “what” and “where” pathways in human auditory cortex. Proceedings of the National Academy of Sciences 103, 14608–14613

  2. [2]

    What” and “where

    Alain, C., Arnott, S. R., Hevenor, S., Graham, S., and Grady, C. L. (2001). “What” and “where” in the human auditory system. Proceedings of the National Academy of Sciences 98, 12301–12306

  3. [3]

    and Burr, D

    Alais, D. and Burr, D. (2004). The ventriloquist effect results from near-optimal bimodal integration. Current Biology 14, 257–262

  4. [4]

    A., Laurent, P

    Anderson, B. A., Laurent, P. A., and Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences 108, 10367–10371

  5. [5]

    V., and Theeuwes, J

    Awh, E., Belopolsky, A. V., and Theeuwes, J. (2012). Top-down versus bottom-up attentional control: A failed theoretical dichotomy. Trends in Cognitive Sciences 16, 437–443

  6. [6]

    Aytar, Y., Castrejon, L., Vondrick, C., Pirsiavash, H., and Torralba, A. (2017). Cross-modal scene networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 40, 2303–2314

  7. [7]

    Ba, J., Mnih, V., and Kavukcuoglu, K. (2014). Multiple object recognition with visual attention. In International Conference on Learning Representations

  8. [8]

    Bacon, W. F. and Egeth, H. E. (1994). Overriding stimulus-driven attentional capture. Perception & Psychophysics 55, 485–496

Show all 189 references
  1. [9]

    Baddeley, A., Hitch, G., and Bower, G. (1974). Recent advances in learning and motivation. Working Memory 8, 647–667

  2. [10]

    Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations

  3. [11]

    and An, S

    Bai, S. and An, S. (2018). A survey on automatic image caption generation. Neurocomputing 311, 291–304

  4. [12]

    Barbey, A. K. (2018). Network neuroscience theory of human intelligence. Trends in Cognitive Sciences 22, 8–20 Fu et al. Selective attention mechanisms and modeling 21

  5. [13]

    I., Fu, D., Liu, X., and Wermter, S

    Barros, P., Parisi, G. I., Fu, D., Liu, X., and Wermter, S. (2018). Expectation learning and crossmodal modulation with a deep adversarial network. In 2018 International Joint Conference on Neural Networks (IJCNN) (IEEE), 1–8

  6. [14]

    Bee, M. A. and Micheyl, C. (2008). The cocktail party problem: what is it? how can it be solved? and why should animal behaviorists study it? Journal of Comparative Psychology 122, 235–251

  7. [15]

    Benes, F. M. (2000). Emerging principles of altered neural circuitry in schizophrenia. Brain Research Reviews 31, 251–269

  8. [16]

    Bizley, J. K. and Cohen, Y. E. (2013). The what, where and how of auditory -object perception. Nature Reviews Neuroscience 14, 693–707

  9. [17]

    and Jensen, O

    Bonnefond, M. and Jensen, O. (2015). Gamma activity coupled to alpha phase as a mechanism for top - down controlled gating. PloS One 10, e0128667

  10. [18]

    and Itti, L

    Borji, A. and Itti, L. (2012). State -of-the-art in visual attention modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence 35, 185–207

  11. [19]

    M., Braver, T

    Botvinick, M. M., Braver, T. S., Barch, D. M., Carter, C. S., and Cohen, J. D. (2001). Conflict monitoring and cognitive control. Psychological Review 108, 624–652

  12. [20]

    Bregman, A. S. (1994). Auditory scene analysis: The perceptual organization of sound (MIT press)

  13. [21]

    Broadbent, D. E. (2013). Perception and Communication (Elsevier)

  14. [22]

    Brungart, D. S. (2001). Informational and energetic masking effects in the perception of two simultaneous talkers. The Journal of the Acoustical Society of America 109, 1101–1109

  15. [23]

    and Sporns, O

    Bullmore, E. and Sporns, O. (2012). The economy of brain network organization. Nature Reviews Neuroscience 13, 336–349

  16. [24]

    Calvert, G. A. (2001). Crossmodal processing in the human brain: insights from functional neuroimaging studies. Cerebral Cortex 11, 1110–1123

  17. [25]

    S., Langer, A., and Kaiser, J

    Chan, J. S., Langer, A., and Kaiser, J. (2016). Temporal integration of multisensory stimuli in autism spectrum disorder: a predictive coding perspective. Journal of Neural Transmission 123, 917–923

  18. [26]

    Cherry, E. C. (1953). Some experiments on the recognition of speech, with one and with two ears. The Journal of the Acoustical Society of America 25, 975–979

  19. [27]

    ventriloquist effect

    Choe, C. S., Welch, R. B., Gilford, R. M., and Juola, J. F. (1975). The “ventriloquist effect”: Visual dominance or response bias? Perception & Psychophysics 18, 55–60

  20. [28]

    K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y

    Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y. (2015). Attention-based models for speech recognition. In Advances in Neural Information Processing Systems. 577–585

  21. [29]

    S., Senior, A., Vinyals, O., and Zisserman, A

    Chung, J. S., Senior, A., Vinyals, O., and Zisserman, A. (2017). Lip reading sentences in the wild. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE), 3444–3453

  22. [30]

    S., Yeung, N., and Kadosh, R

    Clayton, M. S., Yeung, N., and Kadosh, R. C. (2015). The roles of cortical oscillations in sustained attention. Trends in Cognitive Sciences 19, 188–195

  23. [31]

    Colflesh, G. J. and Conway, A. R. (2007). Individual differences in working memory capacity and divided attention in dichotic listening. Psychonomic Bulletin & Review 14, 699–703

  24. [32]

    S., and Yau, J

    Convento, S., Rahman, M. S., and Yau, J. M. (2018). Selective attention gates the interactive crossmodal coupling between perceptual systems. Current Biology 28, 746–752

  25. [33]

    R., Cowan, N., and Bunting, M

    Conway, A. R., Cowan, N., and Bunting, M. F. (2001). The cocktail party phenomenon revisited: The importance of working memory capacity. Psychonomic Bulletin & Review 8, 331–335

  26. [34]

    and Shulman, G

    Corbetta, M. and Shulman, G. L. (2002). Control of goal -directed and stimulus -driven attention in the brain. Nature Reviews Neuroscience 3, 201

  27. [35]

    Dai, B., Chen, C., Long, Y., Zheng, L., Zhao, H., Bai, X., et al. (2018). Neural mechanisms for selectively tuning in to the target speaker in a naturalistic noisy situation. Nature Communications 9, 2405 Fu et al. Selective attention mechanisms and modeling 22

  28. [36]

    Dai, J., Li, Y., He, K., and Sun, J. (2016). R -fcn: Object detection via region -based fully convolutional networks. In Advances in Neural Information Processing Systems. 379–387

  29. [37]

    Das, A., Agrawal, H., Zitnick, L., Parikh, D., and Batra, D. (2017). Human att ention in visual question answering: Do humans and deep networks look at the same regions? Computer Vision and Image Understanding 163, 90–100 Da´vila-Chaco´n, J., Liu, J., and Wermter, S. (2018). E...

  30. [38]

    E., Neal, R

    Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S. (1995). The Helmholtz machine. Neural Computation 7, 889–904

  31. [39]

    and Duncan, J

    Desimone, R. and Duncan, J. (1995). Neural mechanisms of selective visual attention. Annual Review of Neuroscience 18, 193–222

  32. [40]

    Diehl, M. M. and Romanski, L. M. (2014). Responses of prefrontal multisensory neurons to mismatching faces and vocalizations. Journal of Neuroscience 34, 11233–11243

  33. [41]

    and Simon, J

    Ding, N. and Simon, J. Z. (2012). Emergence of neural encoding of auditory objects while listening to competing speakers. Proceedings of the National Academy of Sciences 109, 11854–11859

  34. [42]

    Dipoppa, M., Szwed, M., and Gutkin, B. S. (2016). Controlling working memory operations by selective gating: the roles of oscillations and synchrony. Advances in Cognitive Psychology 12, 209–232

  35. [43]

    J., Killinger, M

    Dorkenwald, S., Schubert, P. J., Killinger, M. F., Urban, G., Mikula, S., Svara, F., et al. (2017). Automated synaptic connectivity inference for volume electron microscopy. Nature Methods 14, 435–442

  36. [44]

    cocktail-party problem

    Du, Y., Kong, L., Wang, Q., Wu, X., and Li, L. (2011). Auditory frequency -following response: a neurophysiological measure for studying the “cocktail-party problem”. Neuroscience & Biobehavioral Reviews 35, 2046–2057

  37. [45]

    A., and Downar, J

    Dunlop, K., Hanlon, C. A., and Downar, J. (2017). Noninvasive brain stimulation treatments for addiction and major depression. Annals of the New York Academy of Sciences 1394, 31–54

  38. [46]

    and Roig, G

    Dwivedi, K. and Roig, G. (2019). Representation similarity analysis for efficient task taxonomy & transfer learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12387–12396

  39. [47]

    and Driver, J

    Eimer, M. and Driver, J. (2001). Crossmodal links in endogenous and exogenous spatial attention: evidence from event-related brain potential studies. Neuroscience & Biobehavioral Reviews 25, 497–511

  40. [48]

    Fan, J. (2014). An information theory account of cognitive control. Frontiers in Human Neuroscience 8, 680

  41. [49]

    D., Fossella, J., Flombaum, J

    Fan, J., McCandliss, B. D., Fossella, J., Flombaum, J. I., and Posner, M. I. (2005). The activation of attentional networks. Neuroimage 26, 471–479

  42. [50]

    D., Sommer, T., Raz, A., and Posner, M

    Fan, J., McCandliss, B. D., Sommer, T., Raz, A., and Posner, M. I. (2002). Testing the efficiency and independence of attentional networks. Journal of Cognitive Neuroscience 14, 340–347

  43. [51]

    and Posner, M

    Fan, J. and Posner, M. (2004). Human attentional networks. Psychiatrische Praxis 31, 210–214

  44. [52]

    T., and Lee, B.-S

    Fang, Y., Lin, W., Lau, C. T., and Lee, B.-S. (2011). A visual attention model combining top -down and bottom-up mechanisms for salient object detection. In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE), 1293–1296

  45. [53]

    J., Wong, A

    Farah, M. J., Wong, A. B., Monheit, M. A., and Morrow, L. A. (1989). Parietal lobe mechanisms of spatial attention: Modality-specific or supramodal? Neuropsychologia 27, 461–470

  46. [54]

    and Friston, K

    Feldman, H. and Friston, K. (2010). Attention, uncertainty, and free -energy. Frontiers in Human Neuroscience 4, 215 Fu et al. Selective attention mechanisms and modeling 23

  47. [55]

    L., Remington, R

    Folk, C. L., Remington, R. W., and Johnston, J. C. (1992). Involuntary covert orienting is contingent on attentional control settings. Journal of Experimental Psychology: Human Perception and Performance 18, 1030–1044

  48. [56]

    Frintrop, S., Rome, E., and Christensen, H. I. (2010). Computational visual attention systems and their cognitive foundations: A survey. ACM Transactions on Applied Perception (TAP) 7, 6

  49. [57]

    Friston, K. (2009). The free-energy principle: a rough guide to the brain? Trends in Cognitive Sciences 13, 293–301

  50. [58]

    I., Wu, H., Magg, S., Liu, X., et al

    Fu, D., Barros, P., Parisi, G. I., Wu, H., Magg, S., Liu, X., et al. (2018). Assessing the contribution of semantic congruency to multisensory integration and conflict resolution. In IROS 2018 Workshop on Crossmodal Learning for Intelligent Robotics

  51. [59]

    Gao, G., Lauri, M., Zhang, J., and Frintrop, S. (2017a). Saliency -guided adaptive seeding for supervoxel segmentation. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE), 4938–4943

  52. [60]

    Gao, L., Guo, Z., Zhang , H., Xu, X., and Shen, H. T. (2017b). Video captioning with attention -based LSTM and semantic consistency. IEEE Transactions on Multimedia 19, 2045–2055

  53. [61]

    J., and Luck, S

    Gaspelin, N., Leonard, C. J., and Luck, S. J. (2015). Direct evidence for active suppression of salient-but- irrelevant sensory inputs. Psychological Science 26, 1740–1750

  54. [62]

    J., and Luck, S

    Gaspelin, N., Leonard, C. J., and Luck, S. J. (2017). Suppression of overt attentional capture by salient-but- irrelevant color singletons. Attention, Perception, & Psychophysics 79, 45–62

  55. [63]

    cocktail party

    Golumbic, E. M. Z., Ding, N., Bickel, S., Lakatos, P., Schevon, C. A., McKhann, G. M., et al. (2013). Mechanisms underlying selective neuronal tracking of attended speech at a “cocktail party”. Neuron 77, 980–991

  56. [64]

    M., Swets, J

    Green, D. M., Swets, J. A., et al. (1966). Signal Detection Theory and Psychophysics, vol. 1 (Wiley New York)

  57. [65]

    M., Goffart, L., and Krauzlis, R

    Hafed, Z. M., Goffart, L., and Krauzlis, R. J. (2008). Superior colliculus inactivation causes stable offsets in eye position during tracking. Journal of Neuroscience 28, 8124–8137

  58. [66]

    Hafed, Z. M. and Krauzlis, R. J. (2008). Goal representations dominate superior colliculus activity during extrafoveal tracking. Journal of Neuroscience 28, 9426–9439

  59. [67]

    B., Weber, C., Kerzel, M., and Wermter, S

    Hafez, M. B., Weber, C., Kerzel, M., and Wermter, S. (2019). Deep intrinsically motivated continuous actor-critic for efficient robotic visuomotor skill learning. Paladyn, Journal of Behavioral Robotics 10, 14–29 Ha¨kkinen, S. and Rinne, T. (2018). Intrinsic, stimulus-driven a...

  60. [68]

    R., and Hanson, S

    Hanson, C., Caglar, L. R., and Hanson, S. J. (2018). Attentional bias in human category learning: The case of deep learning. Frontiers in Psychology 9, 374–384

  61. [69]

    -Y., Tuzel, O., and Farahmand, A

    Hara, K., Liu, M. -Y., Tuzel, O., and Farahmand, A. -M. (2017). Attentional network for visual object detection. arXiv preprint arXiv:1702.01478

  62. [70]

    Henderson, J. M. and Hayes, T. R. (2017). Meaning-based guidance of attention in scenes as revealed by meaning maps. Nature Human Behaviour 1, 743–747

  63. [71]

    and Amedi, A

    Hertz, U. and Amedi, A. (2014). Flexibility and stability in sensory processing revealed using visual -to- auditory sensory substitution. Cerebral Cortex 25, 2049–2064

  64. [72]

    C., McLaughlin, S

    Higgins, N. C., McLaughlin, S. A., Rinne, T., and Stecker, G. C. (2017). Evidence for cue -independent spatial representation in the human auditory cortex during active listening. Proceedings of the National Academy of Sciences 114, E7602–E7611 Fu et al. Selective attention me...

  65. [73]

    Hinz, T., Heinrich, S., and Wermter, S. (2019). Generating multiple objects at spatially distinct locations. In International Conference on Learning Representations (ICLR)

  66. [74]

    M., Kahng, M., Pienta, R., and Chau, D

    Hohman, F. M., Kahng, M., Pienta, R., and Chau, D. H. (2018). Visual analytics in deep learning: An interrogative survey for the next frontiers. IEEE Transactions on Visualization and Computer Graphics 25, 2674–2693

  67. [75]

    and Baldi, P

    Itti, L. and Baldi, P. (2009). Bayesian surprise attracts human attention. Vision Research 49, 1295–1306

  68. [76]

    and Koch, C

    Itti, L. and Koch, C. (2000). A saliency-based search mechanism for overt and covert shifts of visual attention. Vision Research 40, 1489–1506

  69. [77]

    Itti, L., Koch, C., and Niebur, E. (1998). A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis & Machine Intelligence 11, 1254–1259

  70. [78]

    and Mazaheri, A

    Jensen, O. and Mazaheri, A. (2010). Shaping functional architecture by oscillator y alpha activity: gating by inhibition. Frontiers in Human Neuroscience 4, 186

  71. [79]

    Jetley, S., Murray, N., and Vig , E. (2016). End -to-end saliency mapping via probability distribution prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5753–5761

  72. [80]

    A., Robertson, I

    Johnson, K. A., Robertson, I. H., Barry, E., Mulligan, A., Daibhis, A., Daly, M., et al. (2008). Impaired conflict resolution and alerting in children with ADHD: evidence from the attention network task (ANT). Journal of Child Psychology and Psychiatry 49, 1339–1347

  73. [81]

    and Narayanan, S

    Kalinli, O. and Narayanan, S. S. (2007). A saliency -based auditory attention model with applications to unsupervised prominent syllable detection in speech. In Eighth Annual Conference of the International Speech Communication Association

  74. [82]

    Kaya, E. M. and Elhilali , M. (2017). Modelling auditory attention. Philosophical Transactions of the Royal Society B: Biological Sciences 372, 20160101

  75. [83]

    I., Lippert, M., and Logothetis, N

    Kayser, C., Petkov, C. I., Lippert, M., and Logothetis, N. K. (2005). Mechanisms for allocating auditory attention: an auditory saliency map. Current Biology 15, 1943–1947

  76. [84]

    Khaligh-Razavi, S.-M., Henriksson, L., Kay, K., and Kriegeskorte, N. (2017). Fixed versus mixed RSA: Explaining visual representations by fixed and mixed feature sets from shallow and deep computational models. Journal of Mathematical Psychology 76, 184–197

  77. [85]

    Klein, D. A. and Frintrop, S. (2011). Center -surround divergence of feature statistics for salient object detection. In 2011 International Conference on Computer Vision (IEEE), 2214–2219

  78. [86]

    J., Lovejoy, L

    Krauzlis, R. J., Lovejoy, L. P., and Ze´non, A. (2013). Superior colliculus and visual spatial attention. Annual Review of Neuroscience 36, 165–182

  79. [87]

    Kriegeskorte, N., Mur, M., and Bandettini, P. A. (2008). Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience 2, 4

  80. [88]

    S., Ayush, K., and Babu, R

    Kruthiventi, S. S., Ayush, K., and Babu, R. V. (2017). Deepfix: A fully convolutional neural network for predicting human eye fixations. IEEE Transactions on Image Processing 26, 4446–4456

  81. [89]

    E., Vaden Jr, K

    Kuchinsky, S. E., Vaden Jr, K. I., Keren, N. I., Harris, K. C., Ahlstrom, J. B., Dubno, J. R., et al. (2012). Word intelligibility and age predict visual cortex activity during word listening. Cerebral Cortex 22, 1360–1371

  82. [90]

    V., Atkinson, J., and Braddick, O

    Kulke, L. V., Atkinson, J., and Braddick, O. (2016). Neural differences between covert and overt attention studied using EEG with simultaneous remote eye tracking. Frontiers in Human Neuroscience 10, 592

  83. [91]

    S., Gatys, L

    Kummerer, M., Wallis, T. S., Gatys, L. A., and Bethge, M. (2017). Understanding low -and high-level contributions to fixation prediction. In Proceedings of the IEEE International Conference on Computer Vision. 4789–4798 Fu et al. Selective attention mechanisms and modeling 25

  84. [92]

    Lahat, D., Adali, T., and Jutten, C. (2015). Multimodal data fusion: an overview of methods, challenges, and prospects. Proceedings of the IEEE 103, 1449–1477

  85. [93]

    K., Larson, E., Maddox, R

    Lee, A. K., Larson, E., Maddox, R. K., and Shinn -Cunningham, B. G. (2014). Using neuroimaging to understand the cortical mechanisms of auditory selective attention. Hearing Research 307, 111–120

  86. [94]

    and Choo, H

    Lee, K. and Choo, H. (2013). A critical review of selective attention: an interdisciplinary perspective. Artificial Intelligence Review 40, 27–50

  87. [95]

    and Getzmann, S

    Lewald, J. and Getzmann, S. (2015). Electrophysiological correlates of cocktail-party listening. Behavioural Brain Research 292, 157–166

  88. [96]

    Li, G., Gan, Y., Wu, H., Xiao, N., and Lin, L. (2019). Cross-modal attentional context learning for rgb-d object detection. IEEE Transactions on Image Processing 28, 1591–1601

  89. [97]

    Li, W., Yuan, Z., Fang, X., and Wang, C. (2018). Knowing where to look? Analysis on attention of visual question answering system. In Proceedings of the European Conference on Computer Vision (ECCV). 1–8

  90. [98]

    Li, Z. (1999). Contextual influences in V1 as a basis for pop out and asymmetry in visual search. Proceedings of the National Academy of Sciences 96, 10530–10535

  91. [99]

    Li, Z. (2002). A saliency map in primary visual cortex. Trends in Cognitive Sciences 6, 9–16

  92. [100]

    Lidestam, B., Holgersson, J., and Moradi, S. (2014). Comparison of informational vs. energetic masking effects on speechreading performance. Frontiers in Psychology 5, 639

  93. [101]

    and Milanova, M

    Liu, X. and Milanova, M. (2018). Visual attention in deep learning: a review. International Robotics and Automation Journal 4, 154–155

  94. [102]

    F., and Sadjedi, H

    Lotfi, Y., Mehrkian, S., Moossavi, A., Z adeh, S. F., and Sadjedi, H. (2016). Relation between working memory capacity and auditory stream segregation in children with auditory processing disorder. Iranian Journal of Medical Sciences 41, 110–117

  95. [103]

    Lowe, D. G. et al. (1999). Object recognition from local scale -invariant features. In International Conference on Computer Vision. vol. 99, 1150–1157

  96. [104]

    J., Fritz, J

    Lu, K., Xu, Y., Yin, P., Oxenham, A. J., Fritz, J. B., and Shamma, S. A. (2017). Temporal coherence structure rapidly shapes neuronal interactions. Nature Communications 8, 13900

  97. [105]

    -T., Pham, H., and Manning, C

    Luong, M. -T., Pham, H., and Manning, C. D. (2015). Effective approaches to attention -based neural machine translation. In Conference on Empirical Methods in Natural Language Processing

  98. [106]

    A., and Brown, G

    Ma, N., Gonzalez, J. A., and Brown, G. J. (2018). Robust binaural localization of a target sound source by combining spectral source models and deep neural networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing 26, 2122–2131

  99. [107]

    Ma, W. J. (2012). Organizing probabilistic models of perception. Trends in Cognitive Sciences 16, 511–518

  100. [108]

    Mahdi, A., Qin, J., and Crosby, G. (2019). Deepfeat: A bottom-up and top-down saliency model based on deep features of convolutional neural nets. IEEE Transactions on Cognitive and Developmental Systems

  101. [109]

    Mai, G., Schoof, T., and Howell, P. (2019). Modulation of phase-locked neural responses to speech during different arousal states is age-dependent. NeuroImage 189, 734–744

  102. [110]

    J., Teder-Sa¨leja¨rvi, W

    Mcdonald, J. J., Teder-Sa¨leja¨rvi, W. A., Russo, F. D., and Hillyard, S. A. (2003). Neural substrates of perceptual enhancement by cross-modal spatial attention. Journal of Cognitive Neuroscience 15, 10–19

  103. [111]

    Melloni, L., van Leeuwen, S., Alink, A., and Mu¨ ller, N. G. (2012). Interaction between bottom-up saliency and top-down control: how saliency maps are created in the human brain. Cerebral Cortex 22, 2943–2952

  104. [112]

    L., Fink, G

    Mengotti, P., Boers, F., Dombert, P. L., Fink, G. R., and Vossel, S. (2018). Integrating modality-specific expectancies for the deployment of spatial attention. Scientific Reports 8, 1210 Fu et al. Selective attention mechanisms and modeling 26

  105. [113]

    and Uddin, L

    Menon, V. and Uddin, L. Q. (2010). Saliency, switching, attention and control: a network model of insula function. Brain Structure and Function 214, 655–667

  106. [114]

    Meredith, M. A. (2002). On the neuronal basis for multisensory convergence: a brief overview. Cognitive Brain Research 14, 31–40

  107. [115]

    T., Bearpark, H

    Michie, P. T., Bearpark, H. M., Crawford, J. M., and Glue, L. C. (1990). The nature of selective attention effects on auditory event-related potentials. Biological Psychology 30, 219–250

  108. [116]

    Misselhorn, J., Friese, U., and Engel, A. K. (2019). Frontal and parietal alpha oscillations reflect attentional modulation of cross-modal matching. Scientific Reports 9, 5030

  109. [117]

    and Chartier, S

    Morissette, L. and Chartier, S. (2015). Saliency model of auditory attention based on frequency, amplitude and spatial location. In Proceedings of International Joint Conference on Neural Networks (IJCNN) (IEEE), 1–5

  110. [118]

    Mounts, J. R. (2000). Attentional capture by abrupt onsets and feature singletons produces inhibitory surrounds. Perception & Psychophysics 62, 1485–1493

  111. [119]

    Mroueh, Y., Marcheret, E., and Goel, V. (2015). Deep multimodal learning for audio -visual speech recognition. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (IEEE), 2130–2134

  112. [120]

    K., and Whittingstall, K

    Musall, S., von Pfo¨ stl, V., Rauch, A., Logothetis, N. K., and Whittingstall, K. (2012). Effects of neural synchrony on surface EEG. Cerebral Cortex 24, 1045–1053

  113. [121]

    Oldoni, D., De Coensel, B., Boes, M., Rademaker, M., De Baets, B., Van Renterghem, T., et al. (2013). A computational model of auditory attention for use in soundscape research. The Journal of the Acoustical Society of America 134, 852–861

  114. [122]

    I., Barros, P., Fu, D., Magg, S., Wu, H., Liu, X., et al

    Parisi, G. I., Barros, P., Fu, D., Magg, S., Wu, H., Liu, X., et al. (2018). A neurorobotic experiment for crossmodal conflict resolution in complex environments. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE), 2330–2335

  115. [123]

    I., Barros, P., Kerzel, M., Wu, H., Yang, G., Li, Z., et al

    Parisi, G. I., Barros, P., Kerzel, M., Wu, H., Yang, G., Li, Z., et al. (2017). A computational model of crossmodal processing for conflict resolution. In 2017 Joint IEEE International Conference on Development and Learning and Epigenetic Robotics (ICDL-EpiRob) (IEEE), 33–38

  116. [124]

    M., and Lu, Z

    Peng, Y., Yan, K., Sandfort, V., Summers, R. M., and Lu, Z. (2019). A self-attention based deep learning method for lesion attribute detection from CT reports. arXiv preprint arXiv:1904.13018

  117. [125]

    and Adolphs, R

    Pessoa, L. and Adolphs, R. (2010). Emotion processing and the amygdala: from a’low road’to’many roads’ of evaluating biological significance. Nature Reviews Neuroscience 11, 773

  118. [126]

    S., Maroy, R., and Bottlaender, M

    Picard, F., Sadaghiani , S., Leroy, C., Courvoisier, D. S., Maroy, R., and Bottlaender, M. (2013). High density of nicotinic receptors in the cingulo-insular network. Neuroimage 79, 42–51

  119. [127]

    A., and Poliakoff, E

    Poole, D., Gowen, E., Warren, P. A., and Poliakoff, E. (2018). Visual-tactile selective attention in autism spectrum condition: An increased influence of visual distractors. Journal of Experimental Psychology: General 147, 1309

  120. [128]

    and Rothbart, M

    Posner, M. and Rothbart, M. (1998). Attention, self –regulation and consciousness. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 353, 1915–1927

  121. [129]

    and Snyder, C

    Posner, M. and Snyder, C. (1975). Attention and cognitive control In RL Solso,(Ed), Information processing and cognition (pp 55-85) Hillsdale (NJ Erlbaum)

  122. [130]

    Posner, M. I. (1980). Orienting of attention. Quarterly Journal of Experimental Psychology 32, 3–25

  123. [131]

    Posner, M. I. and Cohen, Y. (1984). Components of visual orienting. Attention and Performance X: Control of Language Processes 32, 531–556

  124. [132]

    and Taylor, G

    Ramachandram, D. and Taylor, G. W. (2017). Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Processing Magazine 34, 96–108 Fu et al. Selective attention mechanisms and modeling 27

  125. [133]

    and Farhadi, A

    Redmon, J. and Farhadi, A. (2017). Yolo9000: better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7263–7271

  126. [134]

    Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r -cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems. 91–99

  127. [135]

    and Palmer, S

    Rock, I. and Palmer, S. (1990). The legacy of gestalt psychology. Scientific American 263, 84–91

  128. [136]

    Roseboom, W., Kawabe, T., and Nishida, S. (2013). The cross-modal double flash illusion depends on featural similarity between cross-modal inducers. Scientific Reports 3, 3437

  129. [137]

    and D’Esposito, M

    Sadaghiani, S. and D’Esposito, M. (2014). Functional characterization of the cingulo-opercular network in the maintenance of tonic alertness. Cerebral Cortex 25, 2763–2773

  130. [138]

    and Luck, S

    Sawaki, R. and Luck, S. J. (2010). Capture versus suppression of attention by salient singletons: Electrophysiological evidence for an automatic attend -to-me signal. Attention, Perception, & Psychophysics 72, 1455–1470

  131. [139]

    and Gutschalk, A

    Schadwinkel, S. and Gutschalk, A. (2010). Activity associated with stream segregation in human auditory cortex is similar for spatial and pitch cues. Cerebral Cortex 20, 2863–2873

  132. [140]

    K., Rosen, S., Wickham, L., and Wise, R

    Scott, S. K., Rosen, S., Wickham, L., and Wise, R. J. (2004). A positron emission tomography study of the neural basis of informational and energetic masking effects in speech perception. The Journal of the Acoustical Society of America 115, 813–821

  133. [141]

    R., Foxe, J

    Senkowski, D., Schneider, T. R., Foxe, J. J., and Engel, A. K. (2008). Crossmodal binding through neural coherence: implications for multisensory processing. Trends in Neurosciences 31, 401–409

  134. [142]

    and Kim, R

    Shams, L. and Kim, R. (2010). Crossmodal influences on visual perception. Physics of Life Reviews 7, 269–284

  135. [143]

    Shannon, C. E. (1948). A mathematical theory of com munication. Bell System Technical Journal 27, 379–423

  136. [144]

    Shi, J., Xu, J., Liu, G., Xu, B., et al. (2018). Listen, think and listen again: capturing top -down auditory attention for speaker-independent speech separation. In Proceedings of the International Joint Conference on Artificial Intelligence. 4353–4360

  137. [145]

    Shinn-Cunningham, B. G. (2008). Object-based auditory and visual attention. Trends in Cognitive Sciences 12, 182–186

  138. [146]

    and Zisserman, A

    Simonyan, K. and Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR)

  139. [147]

    Skocaj, D., Leonardis, A., and Kruijff, G. -J. M. (2012). Cross-Modal Learning (Boston, MA: Springer US)

  140. [148]

    Sloutsky, V. M. (2003). The role of similarity in the deve lopment of categorization. Trends in Cognitive Sciences 7, 246–251

  141. [149]

    T., Rorden, C., and Jackson, S

    Smith, D. T., Rorden, C., and Jackson, S. R. (2004). Exogenous orienting of attention depends upon the ability to execute eye movements. Current Biology 14, 792–795

  142. [150]

    -H., Jeong, H.-W., Choi, I., Jeong, D., Kim, K., et al

    Song, Y.-H., Kim, J. -H., Jeong, H.-W., Choi, I., Jeong, D., Kim, K., et al. (2017). A neural circuit for auditory dominance over visual perception. Neuron 93, 940–954

  143. [151]

    Stein, B. E. and Stanford, T. R. (2008). Multisensory integration: current issues from the perspective of the single neuron. Nature Reviews Neuroscience 9, 255

  144. [152]

    E., Wallace, M

    Stein, B. E., Wallace, M. W., Stanford, T. R., and Jiang, W. (2002). Book review: Cortex governs multisensory integration in the midbrain. The Neuroscientist 8, 306–314 Strauß , A., Wo¨stmann, M., and Obleser, J. (2014). Cortical alpha oscillations as a tool for auditory selec...

  145. [153]

    Styles, E. (2006). The psychology of attention (Psychology Press) Fu et al. Selective attention mechanisms and modeling 28

  146. [154]

    Swets, J. A. (2014). Signal detection theory and ROC analysis in psychology and diagnostics: Collected papers (Psychology Press)

  147. [155]

    Talsma, D., Senkowski, D., Soto -Faraco, S., and Woldorff, M. G. (2010). The multifaceted interplay between attention and multisensory integration. Trends in Cognitive Sciences 14, 400–410

  148. [156]

    Theeuwes, J. (1991). Exogenous and endogenous control of attention: The effect of visual onsets and offsets. Perception & Psychophysics 49, 83–90

  149. [157]

    ventriloquism effect

    Thurlow, W. R. and Jack, C. E. (1973). Certain determinants of the “ventriloquism effect”. Perceptual and Motor Skills 36, 1171–1184

  150. [158]

    Todd, J. T. and Van Gelder, P. (1979). Implications of a transient–sustained dichotomy for the measurement of human performance. Journal of Experimental Psychology: Human Perception and Performance 5, 625–638

  151. [159]

    H., and Quigley, K

    Togo, F., Lange, G., Natelson, B. H., and Quigley, K. S. (2015). Attention network test: assessment of cognitive function in chronic fatigue syndrome. Journal of Neuropsychology 9, 1–9

  152. [160]

    and Gormican, S

    Treisman, A. and Gormican, S. (1988). Feature analysis in early vision: evidence from search asymmetries. Psychological Review 95, 15–48

  153. [161]

    Uddin, L. Q. (2015). Sa lience processing and insular cortical function and dysfunction. Nature Reviews Neuroscience 16, 55

  154. [162]

    Uddin, L. Q. and Menon, V. (2009). The anterior insula in autism: under-connected and under-examined. Neuroscience & Biobehavioral Reviews 33, 1198–1203

  155. [163]

    Urbanek, C., Weinges-Evers, N., Bellmann-Strobl, J., Bock, M., Do¨rr, J., Hahn, E., et al. (2010). Attention network test reveals alerting network dysfunction in multiple sclerosis. Multiple Sclerosis Journal 16, 93–99

  156. [164]

    VanRullen, R. (2003). Visual saliency and spike timing in the ventral visual pathway. Journal of Physiology-Paris 97, 365–377

  157. [165]

    Varela, F., Lachaux, J.-P., Rodriguez, E., and Martinerie, J. (2001). The brainweb: phase synchronization and large-scale integration. Nature Reviews Neuroscience 2, 229

  158. [166]

    N., et al

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems. 5998–6008

  159. [167]

    M., and Yoshida, M

    Veale, R., Hafed, Z. M., and Yoshida, M. (2017). How is visual salience computed in the brain? insights from behaviour, neurobiology and modelling. Philosophical Transactions of the Royal Society B: Biological Sciences 372, 20160113

  160. [168]

    Veen, V. v. and Carter, C. S. (2006). C onflict and cognitive control in the brain. Current Directions in Psychological Science 15, 237–240

  161. [169]

    and Pelli, D

    Verghese, P. and Pelli, D. G. (1992). The information capacity of visual attention. Vision Research 32, 983–995

  162. [170]

    Vuilleumier, P. (2005). How brains beware: neural mechanisms of emotional attention. Trends in Cognitive Sciences 9, 585–594

  163. [171]

    T., Meredith, M

    Wallace, M. T., Meredith, M. A., and Stein, B. E. (1998). Multisensory integration in the superior colliculus of the alert cat. Journal of Neurophysiology 80, 1006–1010

  164. [172]

    Wang, B., Yang, Y., Xu, X., Hanjalic, A., and Shen, H. T. (2017). Adversarial cross-modal retrieval. In Proceedings of the 25th ACM International Conference on Multimedia (ACM), 154–162

  165. [173]

    and Chang, P

    Wang, D. and Chang, P. (2008). An oscillatory correlation model of auditory streaming. Cognitive Neurodynamics 2, 7–19

  166. [174]

    and Terman, D

    Wang, D. and Terman, D. (1995). Locally excitatory glob ally inhibitory oscillator networks. IEEE Transactions on Neural Networks 6, 283–286 Fu et al. Selective attention mechanisms and modeling 29

  167. [175]

    Wang, X.-J. (2010). Neurophysiological and computational principles of cortical rhythms in cognition. Physiological Reviews 90, 1195–1268

  168. [176]

    Wang, Y., Huang, M., Zhao, L., et al. (2016). Attention -based LSTM for aspect -level sentiment classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 606–615

  169. [177]

    compellingness

    Warren, D. H., Welch, R. B., and McCarthy, T. J. (1981). The role of visual-auditory “compellingness” in the ventriloquism effect: Implications for transitivity among the spatial senses. Perception & Psychophysics 30, 557–564

  170. [178]

    H., Gopalakrishnan, A., Hazlett, C., and Woldorff, M

    Weissman, D. H., Gopalakrishnan, A., Hazlett, C., and Woldorff, M. (2004). Dorsal anterior cingulate cortex resolves conflict from distracting stimuli by boosting attention toward relevant events. Cerebral Cortex 15, 229–237

  171. [179]

    Welch, R. B. and Warren, D. H. (1980). Immediate perceptual response to intersensory discrepancy. Psychological Bulletin 88, 638–667

  172. [180]

    J., Berg, D

    White, B. J., Berg, D. J., Kan, J. Y., Marino, R. A., Itti, L., and Munoz, D. P. (2017). Superior colliculus neurons encode a visual saliency map during free viewing of natural dynamic video. Nature Communications 8, 14263

  173. [181]

    G., Gallen, C

    Woldorff, M. G., Gallen, C. C., Hampson, S. A., Hillyard, S. A., Pantev, C., Sobel, D., et al. (1993). Modulation of early sensory processing in human auditory cortex during auditory selective attention. Proceedings of the National Academy of Sciences 90, 8722–8726 Wo¨ stmann,...

  174. [182]

    Wrigley, S. N. and Brown, G. J. (2004). A computational model of auditory selective attention. IEEE Transactions on Neural Networks 15, 1151–1163

  175. [183]

    Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., et al. (2015). Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of International Conference on Machine Learning. 2048–2057

  176. [184]

    Yang, G., Nan, W., Zheng, Y., Wu, H., Li, Q., and Liu, X. (2017). Distinct cognitive control mechanisms as revealed by modality-specific conflict adaptation effects. Journal of Experimental Psychology: Human Perception and Performance 43, 807

  177. [185]

    and Jonides, J

    Yantis, S. and Jonides, J. (1984). Abrupt visual onsets and selective attention: evidence from visual search. Journal of Experimental Psychology: Human Perception and Performance 10, 601

  178. [186]

    Yao, X., Han, J., Zhang, D., and Nie, F. (2017). Revisiting co-saliency detection: A novel approach based on two-stage multi-view spectral rotation co -clustering. IEEE Transactions on Image Processing 26, 3196–3209

  179. [187]

    B., and Shamma, S

    Yin, P., Fritz, J. B., and Shamma, S. A. (2014). Rapid spectrotemporal plasticity in primary auditory cortex during behavior. Journal of Neuroscience 34, 4396–4408

  180. [188]

    and Gong, Q

    Zhang, X. and Gong, Q. (2019). Frequency-following responses to complex tones at different frequencies reflect different source configurations. Frontiers in Neuroscience 13, 130

  181. [189]

    G., and Papani colaou, A

    Zouridakis, G., Simos, P. G., and Papani colaou, A. C. (1998). Multiple bilaterally asymmetric cortical sources account for the auditory N1m component. Brain Topography 10, 183–189

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.