REVIEW 3 major objections 5 minor 53 references
An entity-centric attention RL controller for traffic lights cuts delay while producing real-time lane-to-phase maps that match traffic engineering practice.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 00:31 UTC pith:CHTUFBSL
load-bearing objection Solid entity-attention PPO with hard NEMA masks that works in SUMO; the trust claim rests on correlational heatmaps, not causal tests. the 3 major comments →
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Treating intersection lanes and signal phases as separate high-dimensional entities, then relating them with sequential multi-head cross-attention and self-attention inside a masked PPO policy, yields both lower (or equal) delay than strong baselines and a real-time affinity matrix whose weights align with established traffic-engineering principles of queue service and conflict resolution.
What carries the argument
The hierarchical relational attention block: multi-head cross-attention that produces an explicit Nl-by-Np lane-to-phase affinity matrix, followed by multi-head self-attention among lanes, with a deterministic action mask that zeros the probability of every illegal NEMA phase transition before sampling.
Load-bearing premise
That the numerical attention weights extracted from the cross-attention layer are a faithful causal account of why the policy chose a phase, rather than a convenient visual correlate that merely tracks queues.
What would settle it
Retrain or re-evaluate the identical agent on an intersection geometry or demand regime where the attention maps systematically highlight lanes that are not the ones receiving green, while delay still improves; if that occurs, the claimed explanatory fidelity fails.
If this is right
- Traffic agencies can inspect a live affinity matrix to decide whether an RL controller’s phase choice is justified before accepting the action.
- Engineers can inject domain knowledge by inspecting or lightly editing the attention patterns without having to retrain from scratch.
- The same entity-plus-attention pattern can be reused for multi-intersection networks once lane and phase entities are defined across a corridor.
- Deterministic action masking guarantees that any deployed policy remains inside NEMA ring-barrier and clearance constraints even during exploration or distribution shift.
Where Pith is reading between the lines
- If attention fidelity holds under real detector noise and imperfect demand counts, the architecture may also serve as a diagnostic for sensor failure: sudden attention on an empty lane would flag a bad detector.
- The same dual-stage attention could be attached to other constrained sequential decision problems (e.g., ramp metering or transit priority) where both performance and operator trust are required.
- A natural next test is whether freezing the attention heads after training still preserves most of the delay reduction; if so, the explainability layer can be audited offline without continuous re-learning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an entity-centric PPO agent for adaptive traffic signal control that decomposes intersection state into lane entities (Queries) and phase entities (Keys/Values), embeds them into a shared latent space, and processes them with sequential multi-head cross-attention followed by self-attention. The resulting real-time lane–phase affinity matrix (Eq. 8) is presented as an intrinsic explanation of decisions, while a deterministic action mask enforces NEMA ring-barrier, minimum-green and clearance constraints. Evaluated in SUMO on a single four-leg intersection under three training OD patterns and two held-out demand scenarios (N=10 runs), the method matches or reduces delay relative to Synchro-optimized actuated control and a flat-vector RL ablation; attention heatmaps are shown to track queue formation, phase switches and spillback mitigation.
Significance. If the performance and interpretability claims hold, the work supplies a practical bridge between high-performing model-free RL and the auditability requirements of transportation agencies. Strengths include the topology-preserving entity embedding, the dual-stage attention that yields an explicit affinity matrix, the hard action-masking interface integrated into PPO, multi-environment training, and evaluation on structurally dissimilar held-out volumes. The visual alignment of attention with established traffic-engineering principles (queue clearance, coordinated phases) is a concrete step beyond purely post-hoc XAI methods. These elements make the architecture potentially deployable and fine-tunable by domain experts, addressing a recognized barrier to RL adoption in safety-critical infrastructure.
major comments (3)
- [§V-C, Eq. (8), Figs. 7–9] §V-C, Eq. (8) and Figs. 7–9: The central claim that the cross-attention matrix A_cross “quantifies the direct influence of signal phases on specific approach volumes” and thereby yields an “auditable, trust-enabling” architecture rests solely on observational co-occurrence of high vehicle counts, elevated attention weights, and subsequent phase activation. No intervention, attention ablation, gradient-based attribution, or counterfactual experiment is reported to establish that the policy head actually uses these weights rather than residual or flattened pathways (H_lane + Z_1 → Z_t). Because the actor receives only the post-attention vectorized representation, visual alignment under SUMO does not demonstrate that an operator can audit or fine-tune decisions by inspecting A_cross. This leap from correlation to causal explainability is load-bearing for the abstract and §VI trust claims an
- [§V-B, Figs. 5–6] §V-B, Figs. 5–6: Performance comparisons are based on N=10 independent runs, yet neither error bars, confidence intervals, nor formal statistical tests (paired t-test, Wilcoxon, etc.) are reported for movement-level or network delay differences. Statements such as “comparable or superior” and “consistently outperforms” therefore cannot be assessed for significance, especially under the highly asymmetric Scenario 2 where only selected movements improve. Adding variance measures and hypothesis tests is necessary to substantiate the delay-minimization claim relative to ASC and RL-NoAtt.
- [§IV-A, abstract, §VI] §IV-A / Fig. 3 and abstract: All training and evaluation occur on a single isolated four-leg intersection. While the entity-attention architecture is in principle layout-agnostic, the abstract and conclusions assert a “deployable architecture for next-generation adaptive traffic control systems” and position the work against a multi-agent / network-level literature (§II). Without at least one multi-intersection or corridor experiment, the generalization and coordination claims remain untested; either the scope language should be narrowed or additional network-level results supplied.
minor comments (5)
- [Abstract, §V-B] Abstract and §V-B: “outperforms state-of-the-art baselines” is overstated; the only RL baseline is the authors’ own flat-vector ablation (RL-NoAtt). Comparison to at least one recent published attention- or graph-based TSC method would strengthen the claim.
- [Figs. 7–9] Figs. 7–9 span multiple pages with split (a)/(b) captions; the temporal intervals and active-phase annotations are hard to follow. Consider consolidating into single multi-panel figures with consistent color scales and a shared legend.
- [Eq. (18), Table II] Eq. (18): the normalization constant D is listed among free parameters but never given a concrete value or selection procedure; state it explicitly for reproducibility.
- [§VI and throughout] Minor typographical issues: “that have that have” (§VI), missing spaces around some equation references, and inconsistent use of “phase” vs. “phases” in action descriptions.
- [§II-B] Related-work discussion of Hu et al. (2025) and Luo et al. (2024) correctly notes their limitations, yet the paper does not quantify how its lane-level (vs. vehicle-level) representation improves scalability or sensor compatibility; a short quantitative remark would help.
Circularity Check
No circular derivation; only a minor non-load-bearing self-citation of the authors' prior flat-vector RL model used strictly as an ablation baseline.
specific steps
-
self citation load bearing
[Section V-B Comparative Performance Analysis]
"RL-NoAtt was formulated and trained on the same volume patterns as RL-Att in an earlier study [53]."
The flat-vector baseline against which the attention model is claimed to match or outperform is the authors' own prior work. This is ordinary ablation practice and does not force any numerical result or attention weight by construction (ASC remains an independent external baseline and all metrics are simulation-measured), so the circularity is minor and non-load-bearing for the central claims.
full rationale
The paper's architecture (entity embeddings of lanes as Queries and phases as Keys/Values, dual-stage multi-head cross- then self-attention producing A_cross, deterministic NEMA action masking inside PPO, negative-delay reward) is fully specified from first principles and standard components without any quantity being defined in terms of the claimed output. Performance claims are measured against an external optimized ASC baseline (Synchro) plus free-flow delay in SUMO under held-out demand patterns; the attention-alignment claim is a post-hoc observational inspection of logged weights against independent queue and phase traces (Figs. 7-9), not a quantity forced by construction or by a fitted parameter. The sole self-citation ([53]) supplies only the RL-NoAtt ablation comparator trained on the same volumes; it is not invoked to justify uniqueness, force the attention matrices, or underwrite the safety mask or reward. No equation equates a prediction to its own input, no uniqueness theorem is imported from the authors, and no ansatz is smuggled. The derivation chain is therefore self-contained and non-circular; the low score reflects only the ordinary presence of one non-essential self-citation.
Axiom & Free-Parameter Ledger
free parameters (4)
- embedding dimension de =
32
- number of attention heads (Hc, Hs) =
4
- PPO clipping epsilon, entropy coefficient, learning rates, GAE lambda =
0.2 / 0.01 / 1e-4 / 0.95
- reward normalization constant D
axioms (4)
- domain assumption NEMA dual-ring barrier diagram and associated minimum-green, yellow, and all-red constraints define the only legal phase transitions.
- domain assumption Vehicle counts per lane plus current green durations constitute a sufficient Markov state for optimal signal control.
- standard math Scaled dot-product multi-head attention followed by residual LayerNorm yields stable, interpretable relational weights.
- domain assumption SUMO microscopic kinematics and the chosen free-flow delay definition accurately proxy real-world intersection delay and queue dynamics.
invented entities (2)
-
Lane entities as Queries and phase entities as Keys/Values in a shared de-dimensional space
no independent evidence
-
Dual-stage hierarchical attention (cross-attention then self-attention) producing real-time affinity matrices for explainability
no independent evidence
read the original abstract
Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastructure like traffic control, the opaque, black-box nature of deep RL models poses challenges for transportation agency acceptance, regulatory compliance, operational trust, troubleshooting, and fine-tuning. To bridge this gap between high-performance optimization and human-comprehensible interpretability, this effort introduces a novel, explainable entity centric RL framework for safe and transparent traffic signal control. Rather than processing traffic states through monolithic, flat vectors, the proposed architecture disaggregates real-time intersection observations into distinct, high-dimensional lane entities and phase temporal configurations to inherently preserve the structural topology and geometric configurations of the intersection. Relational dependencies and inter-lane conflicts are dynamically extracted via a dual-stage attention network featuring sequential multi-head cross-attention and self-attention blocks. This design yields a real time affinity matrix that quantifies the direct influence of signal phases on specific approach volumes and queues, providing full visual and analytical interpretability. To ensure strict operational reliability, a deterministic action-masking interface is integrated directly into the Proximal Policy Optimization pipeline, explicitly blocking invalid phase transitions to guarantee absolute compliance with established signal timing and safety constraints. Evaluated in a microscopic simulation environment, outperforms state-of-the-art baselines in delay minimization. More importantly, the emergent attention weights align precisely with established traffic engineering principles, offering an auditable, trust-enabling, and deployable architecture for next-generation adaptive traffic control systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Shen, ”Hierarchical reinforcement learning-based traffic signal con- trol,”Scientific Reports, vol
J. Shen, ”Hierarchical reinforcement learning-based traffic signal con- trol,”Scientific Reports, vol. 15, no. 1, 2025, doi: 10.1038/s41598-025- 18449-1
-
[2]
K. Othman, X. Wang, A. Shalaby, and B. Abdulhai, ”Integrating Multimodality and Partial Observability Solutions Into Decentralized Multiagent Reinforcement Learning Adaptive Traffic Signal Control,” IEEE Open Journal of Intelligent Transportation Systems, vol. 6, pp. 322–334, 2025, doi: 10.1109/OJITS.2025.3550312
-
[3]
X. Li, X. Wang, I. Smirnov, S. Sanner, and B. Abdulhai, ”Multi-Hop Upstream Anticipatory Traffic Signal Control With Deep Reinforcement Learning,”IEEE Open Journal of Intelligent Transportation Systems, vol. 6, pp. 554–567, 2025, doi: 10.1109/OJITS.2025.3562757
-
[4]
D. Xu, Z. Yu, X. Liao, and H. Guo, ”A Graph Deep Reinforcement Learning Traffic Signal Control for Multiple Intersections Considering Missing Data,”IEEE Transactions on Vehicular Technology, vol. 73, no. 12, pp. 18307–18319, 2024, doi: 10.1109/TVT.2024.3444475
-
[5]
X. Wang, A. Taitler, I. Smirnov, S. Sanner, and B. Abdulhai, ”eMAR- LIN: Distributed Coordinated Adaptive Traffic Signal Control with Topology-Embedding Propagation,”Transportation Research Record, vol. 2678, no. 4, pp. 189–202, 2024, doi: 10.1177/03611981231184250
-
[6]
S. Shuvoet al., ”Deep Multi-Agent Reinforcement Learning for Real- World Signalized Traffic Corridor Control,” inProceedings of the Inter- national Conference on Machine Learning and Applications (ICMLA), 2024, pp. 644–651, doi: 10.1109/ICMLA61862.2024.00093
-
[7]
T. Shi, F. X. Devailly, D. Larocque, and L. Charlin, ”Improving the Generalizability and Robustness of Large-Scale Traffic Signal Control,” IEEE Open Journal of Intelligent Transportation Systems, vol. 5, pp. 2–15, 2024, doi: 10.1109/OJITS.2023.3331689
-
[8]
D. Kwesiga, S. Vishnoi, A. Guin, and M. Hunter, ”Integrating Transit Signal Priority into Multi-Agent Reinforcement Learning based Traffic Signal Control,” arXiv preprint arXiv:2411.19359, 2024. IEEE TRANSACTIONS ON RELIABILITY 15
Pith/arXiv arXiv 2024
-
[9]
D. Kwesiga, A. Guin, and M. Hunter, ”Adaptive Transit Signal Pri- ority Based on Deep Reinforcement Learning and Connected Vehi- cles in a Traffic Micro Simulation Environment,” inProceedings of the Winter Simulation Conference (WSC), 2024, pp. 1422–1433, doi: 10.1109/WSC63780.2024.10838807
-
[10]
X. Jia, M. Guo, Y . Lyu, J. Qu, D. Li, and F. Guo, ”Adaptive Traffic Signal Control Based on Graph Neural Networks and Dynamic Entropy- Constrained Soft Actor–Critic,”Electronics, vol. 13, no. 23, 2024, doi: 10.3390/electronics13234794
-
[11]
A. Chang, Y . Ji, C. Wang, and Y . Bie, ”CVDMARL: A Communication- Enhanced Value Decomposition Multi-Agent Reinforcement Learning Traffic Signal Control Method,”Sustainability, vol. 16, no. 5, 2024, doi: 10.3390/su16052160
-
[12]
M. Long and E. Chung, ”Transit Signal Priority for Arterial Road with Deep Reinforcement Learning,” inProceedings of the International Conference on Models and Technologies for Intelligent Transportation Systems (MT-ITS), 2023, doi: 10.1109/MT-ITS56129.2023.10241759
-
[13]
X.-Y . Liu, M. Zhu, S. Borst, and A. Walid, ”Deep Reinforcement Learning for Traffic Light Control in Intelligent Transportation Systems,” arXiv preprint arXiv:2302.03669, 2023
Pith/arXiv arXiv 2023
-
[14]
D. Liu and L. Li, ”A traffic light control method based on multi-agent deep reinforcement learning algorithm,”Scientific Reports, vol. 13, no. 1, 2023, doi: 10.1038/s41598-023-36606-2
-
[15]
M. Kolat, B. K ˝ov´ari, T. B ´ecsi, and S. Aradi, ”Multi-Agent Reinforce- ment Learning for Traffic Signal Control: A Cooperative Approach,” Sustainability, vol. 15, no. 4, 2023, doi: 10.3390/su15043479
-
[16]
X. Fu, S. Chen, Q. Liang, and Y . Li, ”Research on Multi-Agent Reinforcement Learning Traffic Control,” inProceedings of the IEEE International Conference on Control, Electronics and Com- puter Technology (ICCECT), 2023, pp. 231–239, doi: 10.1109/IC- CECT57938.2023.10140678
doi:10.1109/ic- 2023
-
[17]
S. Bouktif, A. Cheniki, A. Ouni, and H. El-Sayed, ”Deep reinforcement learning for traffic signal control with consistent state and reward design approach,”Knowledge-Based Systems, vol. 267, 2023, doi: 10.1016/j.knosys.2023.110440
-
[18]
R. Bokade, X. Jin, and C. Amato, ”Multi-Agent Reinforcement Learn- ing Based on Representational Communication for Large-Scale Traffic Signal Control,”IEEE Access, vol. 11, pp. 47646–47658, 2023, doi: 10.1109/ACCESS.2023.3275883
-
[19]
M. Aslani, M. S. Mesgari, S. Seipel, and M. Wiering, ”Developing adaptive traffic signal control by actor-critic and direct exploration methods,”Proceedings of the Institution of Civil Engineers: Transport, vol. 172, no. 5, pp. 289–298, 2019, doi: 10.1680/jtran.17.00085
-
[20]
K. B ´alint, T. Tam´as, and B. Tam´as, ”Deep Reinforcement Learning based approach for Traffic Signal Control,”Transportation Research Procedia, vol. 62, pp. 278–285, 2022, doi: 10.1016/j.trpro.2022.02.035
-
[21]
S. Bouktif, A. Cheniki, and A. Ouni, ”Traffic signal control using hybrid action space deep reinforcement learning,”Sensors, vol. 21, no. 7, 2021, doi: 10.3390/s21072302
-
[22]
N. Casas, ”Deep Deterministic Policy Gradient for Urban Traffic Light Control,” arXiv preprint arXiv:1703.09035, 2017
Pith/arXiv arXiv 2017
-
[23]
H. Lee, Y . Han, Y . Kim, and Y . H. Kim, ”Effects analysis of reward functions on reinforcement learning for traffic signal control,”PLoS ONE, vol. 17, no. 11, 2022, doi: 10.1371/journal.pone.0277813
-
[24]
D. Li, J. Wu, M. Xu, Z. Wang, and K. Hu, ”Adaptive Traffic Sig- nal Control Model on Intersections Based on Deep Reinforcement Learning,”Journal of Advanced Transportation, vol. 2020, 2020, doi: 10.1155/2020/6505893
-
[25]
L. Li, Y . Lv, and F.-Y . Wang, ”Traffic Signal Timing via Deep Rein- forcement Learning,”IEEE Transactions on Intelligent Transportation Systems, vol. 3, 2016
2016
-
[26]
Y . Li, J. He, and Y . Gao, ”Intelligent Traffic Signal Control with Deep Reinforcement Learning at Single Intersection,” inProceedings of the ACM International Conference Proceeding Series, 2021, pp. 399–406, doi: 10.1145/3467707.3467767
-
[27]
S. Liu, G. Wu, and M. Barth, ”A Complete State Transition-Based Traf- fic Signal Control Using Deep Reinforcement Learning,” inProceedings of the IEEE Conference on Technologies for Sustainability (SusTech), 2022, pp. 100–107, doi: 10.1109/SusTech53338.2022.9794168
-
[28]
P. Chanlohaet al., ”Traffic Signal Control with Cell Transmission Model Using Reinforcement Learning for Total Delay Minimisation,” International Journal of Computers, Communications & Control, vol. 10, no. 5, pp. 627–642, 2015, doi: 10.15837/ijccc.2015.5.2025
-
[29]
S. El-Tantawy, B. Abdulhai, and H. Abdelgawad, ”Design of reinforce- ment learning parameters for seamless application of adaptive traffic signal control,”Journal of Intelligent Transportation Systems, vol. 18, no. 3, pp. 227–245, 2014, doi: 10.1080/15472450.2013.810991
-
[30]
X. Liang, X. Du, G. Wang, and Z. Han, ”A Deep Reinforcement Learning Network for Traffic Light Cycle Control,”IEEE Transactions on Vehicular Technology, vol. 68, no. 2, pp. 1243–1253, 2019, doi: 10.1109/TVT.2018.2890726
-
[31]
H. Pang and W. Gao, ”Deep Deterministic Policy Gradient for Traffic Signal Control of Single Intersection,” inProceedings of the Chinese Control and Decision Conference (CCDC), 2019, pp. 5861–5866, doi: 10.1109/CCDC.2019.8832406
-
[32]
S. M. A. Shabestary, B. Abdulhai, H. Ma, and Y . Huo, ”Cycle-level vs. Second-by-Second Adaptive Traffic Signal Control using Deep Reinforcement Learning,” inProceedings of the IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), 2020, pp. 1–8, doi: 10.1109/ITSC45102.2020.9294171
-
[33]
X. Chen, G. Xiong, Y . Lv, Y . Chen, B. Song, and F. Y . Wang, ”A Collaborative Communication-Qmix Approach for Large-scale Net- worked Traffic Signal Control,” inProceedings of the IEEE International Conference on Intelligent Transportation Systems (ITSC), 2021, pp. 3450–3455, doi: 10.1109/ITSC48978.2021.9564683
-
[34]
T. Chu, J. Wang, L. Codec `a, and Z. Li, ”Multi-Agent Deep Reinforce- ment Learning for Large-scale Traffic Signal Control,” arXiv preprint arXiv:1903.04527, 2019
Pith/arXiv arXiv 1903
-
[35]
F. X. Devailly, D. Larocque, and L. Charlin, ”Model-Based Graph Reinforcement Learning for Inductive Traffic Signal Control,”IEEE Open Journal of Intelligent Transportation Systems, vol. 5, pp. 238– 250, 2024, doi: 10.1109/OJITS.2024.3376583
-
[36]
H. Weiet al., ”Colight: Learning network-level cooperation for traffic signal control,” inProceedings of the ACM International Conference on Information and Knowledge Management (CIKM), 2019, pp. 1913– 1922, doi: 10.1145/3357384.3357902
-
[37]
Z. Luo, J. Xu, and F. Chen, ”Multi-agent Reinforcement Traffic Signal Control based on Interpretable Influence Mechanism and Biased ReLU Approximation,” arXiv preprint arXiv:2403.13639, 2024
Pith/arXiv arXiv 2024
-
[38]
Y . Hu, L. Du, and S. M. Easa, ”Explainable reinforcement learning for improved traffic signal control,”Computer-Aided Civil and In- frastructure Engineering, vol. 40, no. 24, pp. 3911–3933, 2025, doi: 10.1111/mice.70037
-
[39]
D. Bahdanau, K. Cho, and Y . Bengio, ”Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014
Pith/arXiv arXiv 2014
-
[40]
K. He, X. Zhang, S. Ren, and J. Sun, ”Deep Residual Learning for Image Recognition,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90
-
[41]
J. L. Ba, J. R. Kiros, and G. E. Hinton, ”Layer normalization,” arXiv preprint arXiv:1607.06450, 2016
Pith/arXiv arXiv 2016
-
[42]
Vaswaniet al., ”Attention is all you need,” inAdvances in Neural Information Processing Systems (NeurIPS), 2017, pp
A. Vaswaniet al., ”Attention is all you need,” inAdvances in Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008
2017
-
[43]
A. Dosovitskiyet al., ”An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020
Pith/arXiv arXiv 2010
-
[44]
Jumperet al., ”Highly accurate protein structure prediction with AlphaFold,”Nature, vol
J. Jumperet al., ”Highly accurate protein structure prediction with AlphaFold,”Nature, vol. 596, no. 7873, pp. 583–589, 2021, doi: 10.1038/s41586-021-03819-2
-
[45]
M. McGibbonet al., ”From intuition to AI: evolution of small molecule representations in drug discovery,”Briefings in Bioinformatics, vol. 25, no. 1, p. bbad422, 2024, doi: 10.1093/bib/bbad422
-
[46]
E. Choi, M. T. Bahadori, J. Sun, J. Kulas, A. Schuetz, and W. Stewart, ”Retain: An interpretable predictive model for healthcare using reverse time attention mechanism,” inAdvances in Neural Information Process- ing Systems (NeurIPS), 2016, pp. 3504–3512
2016
-
[47]
J. Mullenbach, S. Wiegreffe, J. Duke, J. Sun, and J. Eisenstein, ”Ex- plainable prediction of medical codes from clinical text,” inProceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2018, pp. 1101–1111, doi: 10.18653/v1/N18-1100
-
[48]
Xuet al., ”Show, attend and tell: Neural image caption generation with visual attention,” inProceedings of the International Conference on Machine Learning (ICML), 2015, pp
K. Xuet al., ”Show, attend and tell: Neural image caption generation with visual attention,” inProceedings of the International Conference on Machine Learning (ICML), 2015, pp. 2048–2057
2015
-
[49]
A. Mott, D. Zoran, M. Chrzanowski, D. Wierstra, and D. J. Rezende, ”Towards interpretable reinforcement learning using attention aug- mented agents,” inAdvances in Neural Information Processing Systems (NeurIPS), 2019, pp. 12360–12369
2019
-
[50]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, ”Prox- imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017
Pith/arXiv arXiv 2017
-
[51]
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli, ”Image Quality Assessment: From Error Visibility to Structural Similarity,” inIEEE IEEE TRANSACTIONS ON RELIABILITY 16 Transactions on Image Processing, vol. 13, no. 4, 2004, pp. 600–612, doi: 10.1109/TIP.2003.819861
-
[52]
P. A. Lopezet al., ”Microscopic Traffic Simulation using SUMO,” inProceedings of the IEEE International Conference on Intel- ligent Transportation Systems (ITSC), 2018, pp. 2575–2582, doi: 10.1109/ITSC.2018.8569938
-
[53]
D. Kwesiga, A. Guin, K. Abdelghany, and M. Hunter, ”Evaluating the Robustness of Reinforcement Learning based Adaptive Traffic Signal Control,”arXiv preprint arXiv:2603.15283, 2026. BIOGRAPHYSECTION Dickens Kwesigais a Research Engineer with the School of Civil and Environmental Engineering, Georgia Institute of Technology, Atlanta, GA, USA. His current r...
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.