Pith. sign in

REVIEW 2 major objections 5 minor 47 references

Free-Lunch Saliency via Attention in Atari Agents

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A small attention module lets Atari agents emit saliency maps with no score penalty.

desk verdict Worth reviewing: an honest, reproducible evaluation of a simple attention module for Atari agents, whose central 'no performance cost' claim needs statistical grounding. read the letter →

arxiv 1908.02511 v2 pith:SYZFVQBO submitted 2019-08-07 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords saliencymapsattentionmechanismdeepreinforcementlearningAtariinterpretabilityeyetrackingPPOfeaturevisualization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that interpretability and game performance need not trade off in deep reinforcement learning. It inserts a trainable attention module, called FLS (Free Lunch Saliency), into the standard Nature CNN agent and claims the modified agent scores about the same as the baseline while producing saliency maps as a side effect. A second Dense variant yields crisper maps at the cost of lower scores. The authors evaluate the maps against human eye fixations from Atari-HEAD and report that the attention-based models generally beat random at matching where humans look. If the claim is right, explaining an RL agent's decisions can be built into training rather than added on afterward.

What carries the argument

The FLS module is a soft self-attention block inserted between the convolutional body and the fully-connected layers: two 1x1 or 3x3 convolutions ending in a SoftPlus activation with no normalization, whose output mask multiplies the feature map. The same mask, rendered through transposed convolution with a unit kernel, becomes the saliency map. This machinery makes interpretability a byproduct of the features the policy already uses, rather than a separate post-hoc explanation step, and it is what the paper credits for keeping performance on par with the baseline.

What would settle it

Re-run the same comparison with many more seeds and report per-seed score distributions; if a paired statistical test shows Sparse FLS is meaningfully worse than the Nature CNN baseline on any of the six games, the free-lunch claim is refuted. Alternatively, check whether FLS saliency maps fall to shuffled AUC at or below 0.5 on Atari-HEAD, which would refute that they encode human-like attention.

Watch

Extended reading notes

Core claim

The central claim is that adding the FLS module to the established Nature CNN feature extractor yields an agent whose performance is similar to the baseline across six Atari games, making the saliency maps effectively free. The FLS module sits between the convolutional and fully-connected layers, outputs a soft self-attention mask, and multiplies the convolutional features by that mask; upscaling the mask with a transposed convolution produces the saliency map. In the paper's experiments, Sparse FLS matches or slightly exceeds the baseline on most games (for example, Breakout 624 vs 618, BeamRider 6634 vs 6949), and a Dense FLS variant produces sharper visualizations but lower scores. On the Atari-HEAD human-gaze benchmark, the paper finds that no single attention model is a clear winner across NSS, KL divergence, and shuffled AUC, but the maps are broadly better than chance. The paper concludes that FLS can serve as a drop-in replacement for the baseline agent without sacrificing performance.

Load-bearing premise

The claim of 'no performance cost' rests on five training runs per setting and 8192 evaluation episodes, with no significance tests or confidence intervals, so a real drop in score could pass unnoticed amid the large variance.

Editorial extensions

If this is right

  • An FLS-equipped agent can produce saliency maps at inference time with no separate explanation pipeline, making interpretation a free byproduct of the policy.
  • The Sparse FLS agent can replace the Nature CNN baseline under the same PPO hyperparameters while preserving score, so existing RL training setups need little modification.
  • Training-time saliency maps can reveal strategy formation, such as Breakout tunneling behavior and Seaquest agents attending to the oxygen bar.
  • The Dense FLS variant offers a practical trade-off: users who value high-fidelity visualizations can accept lower scores, while those who need score preservation use Sparse FLS.
  • Human eye-tracking data can serve as a benchmark for agent explanations, since the paper shows attention maps are closer to human fixations than chance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the FLS module is architecture-level rather than RL-specific, a natural extension is to supervised tasks like image classification, where a similar attention mask could yield built-in explanations; the paper itself closes by suggesting this direction.
  • The 'free' claim is best read as 'no large performance cost' rather than exact equality: with only five seeds and no significance tests, a modest real drop could hide inside the reported variance, so a larger-seed replication would likely sharpen the conclusion.
  • Attention maps that expose systematic blind spots, such as the Seaquest agent ignoring targets at the top of the screen, could be used as a debugging tool to find policy weaknesses, not just as a human-friendly visualization.
  • A promising follow-up would be to use the attention masks as a training signal, for example by adding a KL or entropy loss to push agent attention toward human gaze or toward temporally consistent regions; the paper identifies similar loss-based ideas as future work.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces two attention-based modifications of the Nature CNN feature extractor for PPO agents on six Atari games: Sparse FLS (a SoftPlus attention module inserted after the convolutional stack) and Dense FLS (a higher-resolution variant requiring sum-pooling). It claims that Sparse FLS matches the baseline's game scores while providing built-in saliency maps, and that Dense FLS trades some performance for sharper visualizations. Saliency maps are evaluated against human eye fixations from Atari-HEAD using NSS, KL divergence, and shuffled AUC. The authors also present ablations (sum-pooling, normalization, SoftPlus variants, module placement) and report training/evaluation curves with code released.

Significance. If established, the Sparse FLS result is genuinely useful: a drop-in interpretability module with no measured performance cost on the tested games, validated against human gaze data. The paper's strengths include a reasonably extensive experimental protocol (5 seeds, 6 games, 8192 evaluation episodes per model), quantitative saliency evaluation with three metrics, an honest discussion of the Breakout score cap with an additional BreakoutInfinite evaluation, and publicly released code. The central performance-equivalence claim is not yet established statistically, and one saliency claim is contradicted by the paper's own table.

major comments (2)
  1. [4.2, Table 3] The central claim that Sparse FLS has 'no performance cost' and 'can be used as a drop-in replacement' is an equivalence claim, but the paper only reports means and standard deviations for 5 seeds and gives no confidence intervals, paired tests, or equivalence bounds. With the reported variance the data cannot distinguish a real drop from noise: on SpaceInvaders the Sparse FLS mean is 9359±13230 versus 3867±3627 for Nature CNN, and on BeamRider Sparse FLS is 315 points below the baseline (6634±2361 versus 6949±2569). The 8192 evaluation episodes per model reduce within-model error, but the between-seed standard deviations dominate. To support the 'free lunch' claim, please add per-game confidence intervals for the differences, use a paired test if the seeds are matched, or run a two-one-sided equivalence test with a predefined margin; without one of these, a 20-30% performance degradation cannot be ruled out on several games.
  2. [4.3, Table 4] The text states that 'all models perform better than random' in the saliency metrics, but Table 4 contradicts this: Dense FLS+SP has NSS = -0.136±0.188 on MsPacman and -0.230±0.557 on SpaceInvaders, and shuffled AUC = 0.419±0.110 on SpaceInvaders, which are at or below chance. Either restrict the claim to the models that actually exceed chance, or add a statistical test against chance for each model and correct for multiple comparisons. The sentence 'no model can be singled out as a clear winner' is similarly unsupported without pairwise significance tests or confidence intervals on the metric differences.
minor comments (5)
  1. [4.3, Eq. (2)] The KL divergence is asymmetric and the text does not state that lower values are better; please add a sentence clarifying the direction of the metric.
  2. [References, Section 4.3] Atari-HEAD is cited as [46] in Section 4.3 but as [45] in Section 3, while [46] is the AGIL paper; please correct the citation to [45] for Atari-HEAD.
  3. [6.2, Table 5] The BreakoutInfinite description says that replacing the score with 432 'triggers the built-in game logic for respawning blocks'; please clarify whether the score reset also affects the reward used for evaluation, since Table 5 reports scores on the modified environment.
  4. [Figures 3, 6, 7] With many overlapping lines and shading, the reward curves are hard to read; consider distinct line styles or separate subplots per architecture.
  5. [Abstract and Section 4.2] The description of Dense FLS as scoring 'slightly lower' is an understatement on BeamRider (866±415 versus 6949±2569); please rephrase to 'lower, sometimes substantially lower'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the FLS module is trained with the RL objective only, and saliency quality is evaluated against the external Atari-HEAD human gaze dataset, so the saliency results are measurements, not constructions from the input.

full rationale

The paper's claimed contributions are (1) an attention module that produces saliency maps as a side effect of RL training, and (2) an empirical claim that Sparse FLS performs similarly to the Nature CNN baseline. Neither claim reduces to its inputs by construction. The FLS module is inserted into the baseline feature extractor and trained with PPO using the standard RL reward objective; no term in the training loss involves the Atari-HEAD gaze data, the reported saliency metrics, or the baseline performance numbers. The saliency maps are generated by transposed convolution of the attention activations and then compared with human eye fixations from Atari-HEAD, which is an external benchmark not used during training. Thus, the NSS/KL/sAUC results are measured against an outside dataset rather than derived from the architecture's own design choices. The 'free lunch' claim is supported only by a 5-seed, 8192-episode comparison with reported means and standard deviations and no confidence intervals or equivalence tests; with rows such as SpaceInvaders Sparse FLS 9359±13230 versus Nature CNN 3867±3627, the evidence is statistically weak. However, underpowered statistics are a correctness risk, not circularity: no parameter is fitted to the baseline scores and then called a prediction, and no equation in the paper makes the performance claim true by definition. The paper also contains no load-bearing self-citations; its architectural precedents are external works (Mnih et al., Sorokin et al., Yang et al.) and the Atari-HEAD dataset is cited for evaluation only. Consequently, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or mathematical entities. The FLS module is a trainable architectural component, not an invented entity with independent falsifiable predictions. All evaluation is empirical, relying on standard RL training and an external human gaze dataset.

assumptions (4)
  • domain assumption PPO with default OpenAI Baselines hyperparameters is a fair and standard training regime for comparing RL architectures.
    The paper relies on this in Section 4.1 to make cross-model comparisons meaningful.
  • domain assumption Atari-HEAD eye fixations are a valid ground truth for evaluating saliency maps.
    Section 4.3 uses NSS, KL, and shuffled AUC against human gaze to judge saliency quality.
  • domain assumption The transposed-convolution upscaling of FLS activations produces a faithful saliency map.
    Section 3 states this visualization method 'strikes the balance' between bilinear upscaling and Jacobian noise, but no evidence links it to actual agent decision making.
  • domain assumption The six selected Atari games are representative for evaluating interpretability methods in RL.
    Section 4.1 chooses games based on the set from [42] plus Breakout, a common benchmark.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Free-Lunch Saliency via Attention in Atari Agents." pith.science (2026). https://pith.science/paper/SYZFVQBO

@misc{pith2026190802511,
  author       = {Pith},
  title        = {Pith review of: Free-Lunch Saliency via Attention in Atari Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SYZFVQBO}},
  note         = {Machine review of arXiv:1908.02511}
}
read the original abstract

We propose a new approach to visualize saliency maps for deep neural network models and apply it to deep reinforcement learning agents trained on Atari environments. Our method adds an attention module that we call FLS (Free Lunch Saliency) to the feature extractor from an established baseline (Mnih et al., 2015). This addition results in a trainable model that can produce saliency maps, i.e., visualizations of the importance of different parts of the input for the agent's current decision making. We show experimentally that a network with an FLS module exhibits performance similar to the baseline (i.e., it is "free", with no performance cost) and can be used as a drop-in replacement for reinforcement learning agents. We also design another feature extractor that scores slightly lower but provides higher-fidelity visualizations. In addition to attained scores, we report saliency metrics evaluated on the Atari-HEAD dataset of human gameplay.

Figures

Figures reproduced from arXiv: 1908.02511 by the authors.

Figure 1
Figure 1. Convolutional and attention blocks. At time [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Model architectures; st – input frames at time t: (a) Nature CNN [24]; (b) DAQN, inspired by [36]; (c) LTIAA [42]; (d) our architecture [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Reward curves during training. The horizontal axis [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Game visualizations: (a-d) Breakout; (e-h) Seaquest; (a,b,e,f) [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Atari-HEAD visualizations: (a,d) Breakout; (b,e) Enduro; (c,f) Seaquest. Each image shows the same frame with [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Reward curves for some models omitted in Fig. [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 8
Figure 8. Figure 8: Breakout scatterplot. 0 500 1000 1500 2000 2500 3000 Nature CNN Seed 1 Seed 9 Seed 17 Seed 25 Seed 33 0 500 1000 1500 2000 2500 3000 DAQN 0 500 1000 1500 2000 2500 3000 RS-PPO 0 500 1000 1500 2000 2500 3000 RS-PPO w/o padding 0 500 1000 1500 2000 2500 3000 Sparse FLS 0…
Figure 10
Figure 10. Figure 10: BeamRider scatterplot. 0 2000 4000 6000 8000 10000 12000 14000 Nature CNN Seed 1 Seed 9 Seed 17 Seed 25 Seed 33 0 2000 4000 6000 8000 10000 12000 14000 DAQN 0 2000 4000 6000 8000 10000 12000 14000 RS-PPO 0 2000 4000 6000 8000 10000 12000 14000 RS-PPO w/o padding 0 200…
Figure 12
Figure 12. Figure 12: SpaceInvaders scatterplot. 0 2500 5000 7500 10000 12500 15000 Nature CNN Seed 1 Seed 9 Seed 17 Seed 25 Seed 33 0 2500 5000 7500 10000 12500 15000 DAQN 0 2500 5000 7500 10000 12500 15000 RS-PPO 0 2500 5000 7500 10000 12500 15000 RS-PPO w/o padding 0 2500 5000 7500 1000…
Figure 14
Figure 14. Figure 14: Seaquest scatterplot. Note the difference in scale [PITH_FULL_IMAGE:figures/full_fig_p015_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 37 canonical work pages

  1. [42]

    Z. Yang, S. Bai, L. Zhang, and P. H. S. Torr. Learn to interpret Atari agents. 2018. 3, 4, 5, 6, 8, 10

  2. [1]

    Adebayo, J

    J. Adebayo, J. Gilmer, I. Goodfellow, and B. Kim. Local explanation methods for deep neural networks lack sensitivity to parameter values. arXiv preprint arXiv:1810.03307, 2018. 2

  3. [2]

    Arulkumaran, M

    K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath. Deep reinforcement learning: A brief survey. IEEE Signal Processing Magazine, 34(6):26–38, Nov 2017. 1

  4. [3]

    S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PloS one, 10(7):e0130140, 2015. 2

  5. [4]

    Bahdanau, K

    D. Bahdanau, K. Cho, and Y . Bengio. Neural machine trans- lation by jointly learning to align and translate. 2014. 1, 3

  6. [5]

    M. G. Bellemare, Y . Naddaf, J. Veness, and M. Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47:253–279, jun 2013. 4

  7. [6]

    Borji, D

    A. Borji, D. N. Sihite, and L. Itti. Quantitative analysis of human-model agreement in visual saliency modeling: A comparative study. IEEE Transactions on Image Processing, 22(1):55–69, 2012. 7

  8. [7]

    Brockman, V

    G. Brockman, V . Cheung, L. Pettersson, J. Schneider, J. Schul- man, J. Tang, and W. Zaremba. OpenAI Gym. arXiv preprint arXiv:1606.01540, 2016. 4

Show all 47 references
  1. [8]

    Chakraborty, R

    S. Chakraborty, R. Tomsett, R. Raghavendra, D. Harborne, M. Alzantot, F. Cerutti, M. Srivastava, A. Preece, S. Julier, R. M. Rao, T. D. Kelley, D. Braines, M. Sensoy, C. J. Willis, and P. Gurram. Interpretability of deep learning models: A survey of results. In 2017 IEEE Smart...

  2. [9]

    Y . Chen, R. Zhu, and H. Liang. 10703 course project final report: Observe, attend and act: Attention mechanisms in DQN. page 8, 2017. 2, 3, 4

  3. [10]

    Choi, B.-J

    J. Choi, B.-J. Lee, and B.-T. Zhang. Multi-focus attention network for efficient deep reinforcement learning. 2017. 3

  4. [11]

    Dhariwal, C

    P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y . Wu, and P. Zhokhov. OpenAI Baselines. https://github.com/openai/ baselines, 2017. 4

  5. [12]

    Dugas, Y

    C. Dugas, Y . Bengio, F. Bélisle, C. Nadeau, and R. Garcia. Incorporating second-order functional knowledge for better option pricing. In Advances in neural information processing systems, pages 472–478, 2001. 3

  6. [13]

    Espeholt, H

    L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V . Mnih, T. Ward, Y . Doron, V . Firoiu, T. Harley, I. Dunning, et al. Im- pala: Scalable distributed deep-rl with importance weighted actor-learner architectures. arXiv preprint arXiv:1802.01561,

  7. [14]

    François-Lavet, P

    V . François-Lavet, P. Henderson, R. Islam, M. G. Bellemare, and J. Pineau. An introduction to deep reinforcement learning. CoRR, abs/1811.12560, 2018. 1

  8. [15]

    Greydanus, A

    S. Greydanus, A. Koul, J. Dodge, and A. Fern. Visualizing and understanding atari agents. 2017. 2, 7

  9. [16]

    Harel, C

    J. Harel, C. Koch, and P. Perona. Graph-based visual saliency. In Proceedings of the 19th International Conference on Neu- ral Information Processing Systems, NIPS’06, pages 545–552, Cambridge, MA, USA, 2006. MIT Press. 2

  10. [17]

    Hausknecht and P

    M. Hausknecht and P. Stone. Deep recurrent q-learning for partially observable MDPs. 2015. 2, 3

  11. [18]

    Henderson, R

    P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger. Deep reinforcement learning that matters. In Thirty-Second AAAI Conference on Artificial Intelligence,

  12. [19]

    Hessel, J

    M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Os- trovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver. Rainbow: Combining improvements in deep re- inforcement learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18...

  13. [20]

    Hooker, D

    S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim. Evaluating feature importance estimates. arXiv preprint arXiv:1806.10758, 2018. 2

  14. [21]

    L. Itti, C. Koch, and E. Niebur. A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis & Machine Intelligence, (11):1254–1259,

  15. [22]

    Manchin, E

    A. Manchin, E. Abbasnejad, and A. v. d. Hengel. Reinforce- ment learning with attention that works: A self-supervised approach. 2019. 3

  16. [23]

    V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller. Playing Atari with deep reinforcement learning. page 9, 2013. 1, 2, 3, 6

  17. [24]

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through deep ...

  18. [25]

    A. Mott, D. Zoran, M. Chrzanowski, D. Wierstra, and D. J. Rezende. Towards interpretable reinforcement learning using attention augmented agents. CoRR, abs/1906.02500, 2019. 3, 5

  19. [26]

    Mousavi, M

    S. Mousavi, M. Schukat, E. Howley, A. Borji, and N. Moza- yani. Learning to predict where to look in interactive en- vironments using deep recurrent q-learning. 2016. 2, 3, 4, 7

  20. [27]

    H. Noh, A. Araujo, J. Sim, T. Weyand, and B. Han. Large- scale image retrieval with attentive deep local features. 2016. 3

  21. [28]

    R. J. Peters, A. Iyer, L. Itti, and C. Koch. Components of bottom-up gaze allocation in natural images. Vision research, 45(18):2397–2416, 2005. 7

  22. [29]

    Riche, M

    N. Riche, M. Duvinage, M. Mancas, B. Gosselin, and T. Du- toit. Saliency and human fixations: State-of-the-art and study of comparison metrics. In Proceedings of the IEEE inter- national conference on computer vision, pages 1153–1160,

  23. [30]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. 2017. 4

  24. [31]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Pro- ceedings of the IEEE International Conference on Computer Vision, pages 618–626, 2017. 2

  25. [32]

    J. Seo, J. Choe, J. Koo, S. Jeon, B. Kim, and T. Jeon. Noise- adding methods of saliency map as series of higher order partial derivative. arXiv preprint arXiv:1806.03000, 2018. 2

  26. [33]

    Shrikumar, P

    A. Shrikumar, P. Greenside, and A. Kundaje. Learning impor- tant features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 3145–3153. JMLR. org, 2017. 2

  27. [34]

    Simonyan, A

    K. Simonyan, A. Vedaldi, and A. Zisserman. Deep Inside Con- volutional Networks: Visualising Image Classification Mod- els and Saliency Maps. arXiv e-prints, page arXiv:1312.6034, Dec 2013. 1, 2

  28. [35]

    Smilkov, N

    D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg. Smoothgrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825, 2017. 2

  29. [36]

    Sorokin, A

    I. Sorokin, A. Seleznev, M. Pavlov, A. Fedorov, and A. Ignat- eva. Deep attention recurrent q-network. 2015. 2, 3, 4, 5, 6, 8, 10

  30. [37]

    J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Ried- miller. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806, 2014. 2

  31. [38]

    Sundararajan, A

    M. Sundararajan, A. Taly, and Q. Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 3319–

  32. [39]

    R. S. Sutton and A. G. Barto. Reinforcement Learning. MIT Press, Cambridge, MA, 2nd edition, 2018. 1

  33. [40]

    Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas. Dueling network architectures for deep reinforcement learning. 2015. 2

  34. [41]

    Weitkamp, E

    L. Weitkamp, E. van der Pol, and Z. Akata. Visual rationaliza- tions in deep reinforcement learning for Atari games. 2019. 2

  35. [43]

    Yuezhang, R

    L. Yuezhang, R. Zhang, and D. H. Ballard. An initial attempt of combining visual selective attention with deep reinforce- ment learning. 2018. 3

  36. [44]

    Zahavy, N

    T. Zahavy, N. Ben-Zrihem, and S. Mannor. Graying the black box: Understanding DQNs. In International Conference on Machine Learning, pages 1899–1908, 2016. 2

  37. [45]

    Zhang, Z

    R. Zhang, Z. Liu, L. Guan, L. Zhang, M. M. Hayhoe, and D. H. Ballard. Atari-HEAD: Atari human eye-tracking and demonstration dataset. CoRR, abs/1903.06754, 2019. 2, 3

  38. [46]

    Zhang, Z

    R. Zhang, Z. Liu, L. Zhang, J. A. Whritner, K. S. Muller, M. M. Hayhoe, and D. H. Ballard. AGIL: Learning attention from human for visuomotor tasks. In Proceedings of the European Conference on Computer Vision (ECCV) , pages 663–679, 2018. 2, 3, 6, 7

  39. [47]

    BreakoutInfinite

    Appendix 6.1. Performance details Figs. 6–7 show curves for some models omitted in Fig. 3. Figs. 8–14 show in detail performance evaluations sum- marized in Table 3. Each scatterplot corresponds to one model. Each circle in each scatterplot corresponds to one completed episode...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.