{"work":{"id":"c7f2f5a9-ae4b-48db-aff0-24b9d0528995","openalex_id":"https://openalex.org/W2509830164","doi":"10.1145/2939672.2939875","arxiv_id":"1412.3555","raw_key":null,"title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","authors":null,"authors_text":"Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, Yoshua Bengio","year":2014,"venue":"cs.NE","abstract":"In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long short-term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). We evaluate these recurrent units on the tasks of polyphonic music modeling and speech signal modeling. Our experiments revealed that these advanced recurrent units are indeed better than more traditional recurrent units such as tanh units. Also, we found GRU to be comparable to LSTM.","external_url":"https://arxiv.org/abs/1412.3555","cited_by_count":554,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1412.3555","created_at":"2026-05-09T03:35:57.469145+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","render_title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling"},"hub":{"state":{"work_id":"c7f2f5a9-ae4b-48db-aff0-24b9d0528995","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":134,"external_cited_by_count":554,"distinct_field_count":23,"first_pith_cited_at":"2017-06-12T17:57:34+00:00","last_pith_cited_at":"2026-07-08T22:22:37+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T17:59:24.274802+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":8},{"context_role":"method","n":3},{"context_role":"baseline","n":1},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":8},{"context_polarity":"use_method","n":3},{"context_polarity":"baseline","n":1},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","claims":[{"claim_text":"In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long short-term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). We evaluate these recurrent units on the tasks of polyphonic music modeling and speech signal modeling. Our experiments revealed that these advanced recurrent units are indeed better than more traditional recurrent units such as tanh units. Also, we found GRU to be comparable to LSTM.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Most actions exhibit dependencies on preceding actions. These preceding trigger actions may not necessarily be con- fined to the immediate past but can extend back to the long- term history. Consequently, many methods aim to exploit the information contained in the long-term history. Recurrent networks. Having an internal memory, recur- rent networks, including LSTMs [138] and GRUs [146], are often adopted for exploiting the sequential action his- tory [45], [75], [77], [98]. Qi et al. [147] emp","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The GRU extends this capability to temporal sequences, enabling the model to learn dynamic behaviors and dependencies over time. Unlike traditional recurrent neural networks, GRUs introduce gating mechanisms that regulate information flow between successive time steps, allowing the network to retain long-term dependencies while also mitigating vanishing or exploding gradient issues [25]. The GRU's hidden state acts as a compact memory of past observations, updated through reset and update gates ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"CL] 2 May 2026 Sample GenerationChallenging Sample Re-GenerationQuery Hint SFT/DPOQuery RationalePrediction 𝐷!\"# iReMedi RationalePrediction QueryRationalePrediction ×𝑘 𝐷$%& Correct Answer IncorrectAnswer Figure 1: Overview ofReMedi, which operates iteratively across three stages: (1) Sample Generation, (2) Challenging Sample Re-Generation, and (3) Model Training. The dotted orange line represents the data processing pipeline for DPO, while the solid blue line denotes the pipeline for SFT. prove","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"International Conference on Learning Representations (ICLR) . 2021. [20] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. \"PaLM: Scaling Language Modeling with Path- ways\". In:Journal of Machine Learning Research24.240 (2023), pp. 1-113.url: http://jmlr.org/papers/v24/22- 1144.html. [21] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. \"Empirical Evaluation of G","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"International Conference on Learning Representations (ICLR) . 2021. [16] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. \"PaLM: Scaling Language Modeling with Pathways\". In: Journal of Machine Learning Research 24.240 (2023), pp. 1-113. url: http://jmlr.org/papers/v24/22- 1144.html. [17] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. \"Empirical Evaluation of ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 Large Language Models: Foundations and Capabili- ties The study of language models has a long and rich history [74], beginning with early statistical language models [75] and smaller neural network architectures [76]. Building on these foundational concepts, recent advancements have focused on transformer-based LLMs, such as the Generative Pre-trained Transformers (GPTs) [77]. These models are pretrained on extensive text corpora and feature significantly larger model sizes, validating scaling","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":3,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-30T21:40:32.480594+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"ea9fd9b8-f844-40bd-b210-1f4e12881c9c","orcid":null,"display_name":"Junyoung Chung"},{"id":"73429235-8714-4af9-939e-2fcda3390d48","orcid":null,"display_name":"Caglar Gulcehre"},{"id":"47976fb5-b1ea-42da-8948-b8959ffe2992","orcid":null,"display_name":"Kyunghyun Cho"},{"id":"5a7a1034-5ab5-46b2-9f24-66204fbd9091","orcid":null,"display_name":"Yoshua Bengio"}]},"error":null,"updated_at":"2026-06-30T21:40:32.768656+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:20:07.377428+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":5},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":4},{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","work_id":"4ee75248-1199-492c-a52f-6661e0f4adff","shared_citers":4},{"title":"Retentive Network: A Successor to Transformer for Large Language Models","work_id":"5b0449ac-92b0-41f2-8b4f-586c2b5a08b6","shared_citers":4},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3},{"title":"Eagle and finch: Rwkv with matrix- valued states and dynamic recurrence.arXiv preprint arXiv:2404.05892","work_id":"b993c2c1-c4d3-4ade-a380-dc5464bcb940","shared_citers":3},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":3},{"title":"Hochreiter and J","work_id":"c3b0bfa7-6764-45f1-a40d-45baaee9d22c","shared_citers":3},{"title":"HyperNetworks","work_id":"45baa084-f34a-44f4-bb45-6d2b9131cb67","shared_citers":3},{"title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer","work_id":"2c6b3f6d-54e4-4df7-baa7-475a490799af","shared_citers":3},{"title":"RWKV: Reinventing RNNs for the Transformer Era","work_id":"524dc80d-f4ef-4f89-bf1a-9a8c1e4b6a81","shared_citers":3},{"title":"Searching for Activation Functions","work_id":"3a43a02d-e005-47ad-8373-c166e20c9ee9","shared_citers":3},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":3},{"title":"Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex","work_id":"27d00efc-1c8b-4af4-bd30-5dfa575d0989","shared_citers":2},{"title":"An attention free transformer","work_id":"49358264-ed48-4ca7-9326-d1365796f89a","shared_citers":2},{"title":"Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering","work_id":"b9d68fc0-5b23-4def-bc5e-6ad71d64eec6","shared_citers":2},{"title":"Combining Recurrent, Convolutional, and Continuous-time Models with the Linear State Space Layer","work_id":"6a784a12-0336-43ba-97ba-761d1d4ce254","shared_citers":2},{"title":"CosFormer: Rethinking Softmax in Attention","work_id":"0f8f7dc1-c7a6-4757-87d7-508b56aba24a","shared_citers":2},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":2},{"title":"Deep variational bayes filters: Unsupervised learning of state space models from raw data","work_id":"f695e0b3-79e3-4fa9-8d65-6d2851078829","shared_citers":2},{"title":"Diagonal State Spaces are as Effective as Structured State Spaces","work_id":"b8804c9b-93c0-4d6c-947a-af4e3131cb71","shared_citers":2},{"title":"Efficiently Modeling Long Sequences with Structured State Spaces","work_id":"3ca3e8df-d89a-4505-a4c8-ceea67cc5e5e","shared_citers":2},{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","shared_citers":2}],"time_series":[{"n":1,"year":2017},{"n":1,"year":2023},{"n":1,"year":2024},{"n":1,"year":2025},{"n":29,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:19:51.309825+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:19:46.526832+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","claims":[{"claim_text":"In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long short-term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). We evaluate these recurrent units on the tasks of polyphonic music modeling and speech signal modeling. Our experiments revealed that these advanced recurrent units are indeed better than more traditional recurrent units such as tanh units. Also, we found GRU to be comparable to LSTM.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Most actions exhibit dependencies on preceding actions. These preceding trigger actions may not necessarily be con- fined to the immediate past but can extend back to the long- term history. Consequently, many methods aim to exploit the information contained in the long-term history. Recurrent networks. Having an internal memory, recur- rent networks, including LSTMs [138] and GRUs [146], are often adopted for exploiting the sequential action his- tory [45], [75], [77], [98]. Qi et al. [147] emp","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The GRU extends this capability to temporal sequences, enabling the model to learn dynamic behaviors and dependencies over time. Unlike traditional recurrent neural networks, GRUs introduce gating mechanisms that regulate information flow between successive time steps, allowing the network to retain long-term dependencies while also mitigating vanishing or exploding gradient issues [25]. The GRU's hidden state acts as a compact memory of past observations, updated through reset and update gates ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"CL] 2 May 2026 Sample GenerationChallenging Sample Re-GenerationQuery Hint SFT/DPOQuery RationalePrediction 𝐷!\"# iReMedi RationalePrediction QueryRationalePrediction ×𝑘 𝐷$%& Correct Answer IncorrectAnswer Figure 1: Overview ofReMedi, which operates iteratively across three stages: (1) Sample Generation, (2) Challenging Sample Re-Generation, and (3) Model Training. The dotted orange line represents the data processing pipeline for DPO, while the solid blue line denotes the pipeline for SFT. prove","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"International Conference on Learning Representations (ICLR) . 2021. [20] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. \"PaLM: Scaling Language Modeling with Path- ways\". In:Journal of Machine Learning Research24.240 (2023), pp. 1-113.url: http://jmlr.org/papers/v24/22- 1144.html. [21] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. \"Empirical Evaluation of G","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"International Conference on Learning Representations (ICLR) . 2021. [16] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. \"PaLM: Scaling Language Modeling with Pathways\". In: Journal of Machine Learning Research 24.240 (2023), pp. 1-113. url: http://jmlr.org/papers/v24/22- 1144.html. [17] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. \"Empirical Evaluation of ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 Large Language Models: Foundations and Capabili- ties The study of language models has a long and rich history [74], beginning with early statistical language models [75] and smaller neural network architectures [76]. Building on these foundational concepts, recent advancements have focused on transformer-based LLMs, such as the Generative Pre-trained Transformers (GPTs) [77]. These models are pretrained on extensive text corpora and feature significantly larger model sizes, validating scaling","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":3,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-30T21:40:32.772242+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","claims":[{"claim_text":"In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long short-term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). We evaluate these recurrent units on the tasks of polyphonic music modeling and speech signal modeling. Our experiments revealed that these advanced recurrent units are indeed better than more traditional recurrent units such as tanh units. Also, we found GRU to be comparable to LSTM.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:19:36.179529+00:00"}},"summary":{"title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","claims":[{"claim_text":"In this paper we compare different types of recurrent units in recurrent neural networks (RNNs). Especially, we focus on more sophisticated units that implement a gating mechanism, such as a long short-term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). We evaluate these recurrent units on the tasks of polyphonic music modeling and speech signal modeling. Our experiments revealed that these advanced recurrent units are indeed better than more traditional recurrent units such as tanh units. Also, we found GRU to be comparable to LSTM.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":5},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":4},{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","work_id":"4ee75248-1199-492c-a52f-6661e0f4adff","shared_citers":4},{"title":"Retentive Network: A Successor to Transformer for Large Language Models","work_id":"5b0449ac-92b0-41f2-8b4f-586c2b5a08b6","shared_citers":4},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3},{"title":"Eagle and finch: Rwkv with matrix- valued states and dynamic recurrence.arXiv preprint arXiv:2404.05892","work_id":"b993c2c1-c4d3-4ade-a380-dc5464bcb940","shared_citers":3},{"title":"Gaussian Error Linear Units (GELUs)","work_id":"0466fd22-03a1-4a61-af0a-a900e77bb023","shared_citers":3},{"title":"Hochreiter and J","work_id":"c3b0bfa7-6764-45f1-a40d-45baaee9d22c","shared_citers":3},{"title":"HyperNetworks","work_id":"45baa084-f34a-44f4-bb45-6d2b9131cb67","shared_citers":3},{"title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer","work_id":"2c6b3f6d-54e4-4df7-baa7-475a490799af","shared_citers":3},{"title":"RWKV: Reinventing RNNs for the Transformer Era","work_id":"524dc80d-f4ef-4f89-bf1a-9a8c1e4b6a81","shared_citers":3},{"title":"Searching for Activation Functions","work_id":"3a43a02d-e005-47ad-8373-c166e20c9ee9","shared_citers":3},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":3},{"title":"Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex","work_id":"27d00efc-1c8b-4af4-bd30-5dfa575d0989","shared_citers":2},{"title":"An attention free transformer","work_id":"49358264-ed48-4ca7-9326-d1365796f89a","shared_citers":2},{"title":"Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering","work_id":"b9d68fc0-5b23-4def-bc5e-6ad71d64eec6","shared_citers":2},{"title":"Combining Recurrent, Convolutional, and Continuous-time Models with the Linear State Space Layer","work_id":"6a784a12-0336-43ba-97ba-761d1d4ce254","shared_citers":2},{"title":"CosFormer: Rethinking Softmax in Attention","work_id":"0f8f7dc1-c7a6-4757-87d7-508b56aba24a","shared_citers":2},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":2},{"title":"Deep variational bayes filters: Unsupervised learning of state space models from raw data","work_id":"f695e0b3-79e3-4fa9-8d65-6d2851078829","shared_citers":2},{"title":"Diagonal State Spaces are as Effective as Structured State Spaces","work_id":"b8804c9b-93c0-4d6c-947a-af4e3131cb71","shared_citers":2},{"title":"Efficiently Modeling Long Sequences with Structured State Spaces","work_id":"3ca3e8df-d89a-4505-a4c8-ceea67cc5e5e","shared_citers":2},{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","shared_citers":2}],"time_series":[{"n":1,"year":2017},{"n":1,"year":2023},{"n":1,"year":2024},{"n":1,"year":2025},{"n":29,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"73429235-8714-4af9-939e-2fcda3390d48","orcid":null,"display_name":"Caglar Gulcehre","source":"manual","import_confidence":0.72},{"id":"ea9fd9b8-f844-40bd-b210-1f4e12881c9c","orcid":null,"display_name":"Junyoung Chung","source":"manual","import_confidence":0.72},{"id":"47976fb5-b1ea-42da-8948-b8959ffe2992","orcid":null,"display_name":"Kyunghyun Cho","source":"manual","import_confidence":0.72},{"id":"5a7a1034-5ab5-46b2-9f24-66204fbd9091","orcid":null,"display_name":"Yoshua Bengio","source":"manual","import_confidence":0.72}]}}