{"work":{"id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","openalex_id":"https://openalex.org/W2242818861","doi":"10.48550/arxiv.1308.3432","arxiv_id":"1308.3432","raw_key":null,"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","authors":null,"authors_text":"Yoshua Bengio, Nicholas L\\'eonard, Aaron Courville","year":2013,"venue":"cs.LG","abstract":"Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we \"back-propagate\" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A second approach, introduced here, decomposes the operation of a binary stochastic neuron into a stochastic binary part and a smooth differentiable part, which approximates the expected effect of the pure stochatic binary neuron to first order. A third approach involves the injection of additive or multiplicative noise in a computational graph that is otherwise differentiable. A fourth approach heuristically copies the gradient with respect to the stochastic output directly as an estimator of the gradient with respect to the sigmoid argument (we call this the straight-through estimator). To explore a context where these estimators are useful, we consider a small-scale version of {\\em conditional computation}, where sparse stochastic units form a distributed representation of gaters that can turn off in combinatorially many ways large chunks of the computation performed in the rest of the neural network. In this case, it is important that the gating units produce an actual 0 most of the time. The resulting sparsity can be potentially be exploited to greatly reduce the computational cost of large deep networks for which conditional computation would be useful.","external_url":"https://arxiv.org/abs/1308.3432","cited_by_count":2009,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1308.3432","created_at":"2026-05-09T03:35:49.635087+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","render_title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation"},"hub":{"state":{"work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":210,"external_cited_by_count":2009,"distinct_field_count":27,"first_pith_cited_at":"2016-11-03T19:48:08+00:00","last_pith_cited_at":"2026-07-08T18:18:41+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-20T11:39:29.834249+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"method","n":16},{"context_role":"background","n":10}],"polarity_counts":[{"context_polarity":"use_method","n":16},{"context_polarity":"background","n":9},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","claims":[{"claim_text":"Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we \"back-propagate\" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A s","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"𝑖 =𝑓 𝜃 (𝑥 tr 𝑖 ),( ˆ𝑧tr 𝑖 , 𝑠tr 𝑖 )=𝑄(𝑧 tr 𝑖 ), ˜𝑥 tr 𝑖 =ℎ 𝜙 ( ˆ𝑧tr 𝑖 ), 𝑧ta 𝑖 =𝑓 𝜃 (𝑥 ta 𝑖 ),( ˆ𝑧ta 𝑖 , 𝑠ta 𝑖 )=𝑄(𝑧 ta 𝑖 ), ˜𝑥 ta 𝑖 =ℎ 𝜙 ( ˆ𝑧ta 𝑖 ). (1) Here, ˆ𝑧tr 𝑖 and ˆ𝑧ta 𝑖 are the quantized embeddings, while 𝑠tr 𝑖 and 𝑠ta 𝑖 are the corresponding SIDs. Since nearest-codeword lookup is non- differentiable, we adopt the straight-through estimator (STE) [2] to pass gradients from the quantized embeddings to the encoder-side representations during back-propagation. The forward pipeline above de","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"yielding superior rendering quality and improved efficiency for generalizable novel view synthesis. C. Dynamic Neural Networks. Dynamic neural networks [54]-[61] are intended to adap- tively adjust their weights or structure to handle given input with appropriate states, offering a more flexible alternative to static architectures. Recently, this paradigm has evolved from basic conditional computation [62] toward sophisticated Mixture-of- Experts (MoE) architectures, which effectively scale mode","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"machine learning, pages 3734-3743. pmlr, 2019. [51] Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax.arXiv preprint arXiv:1611.01144, 2016. [52] Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation.arXiv preprint arXiv:1308.3432, 2013. [53] Miles Macklin. Warp: A high-performance python framework for gpu simulation and graphics, March 2022. NVIDIA GPU Technology Co","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"We do not anneal this threshold and do not use additional entropy or sparsity regularization on the gate. The final linear head is zero-initialized so that the gate starts from a neutral routing distribution. The forward routing decision is binary, Gl i = round(pl i),(2) and gradients are passed through this binary threshold with a straight-through estimator [ 4]. The effective PE for subject tokens is then P El eff,i =G l i P Ebase,i + (1−G l i)P Eswap,i, i∈ S l,(3) where P Ebase denotes the or","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"to model p by composition of (P q)⊤ (q→u ) and P p (u→p ). This way, the objective of eq. (3.2) becomes nX p̸=q LX i=1 ⟨W p i , P p i (P q i )⊤W q i (P p i−1(P q i−1)⊤)⊤⟩= nX p̸=q LX i=1 ⟨(P p i )⊤W p i P p i−1,(P q i )⊤W q i P q i−1⟩. (3.3) As stated by Theorem 3.2.1, the permutations we obtain using eq. (3.3) are cycle consistent. We refer the reader to Bernard et al.[15] for the proof and a complete discussion of the subject. Theorem 3.2.1(Restated from Bernard et al. 15).Given a set of n mod","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"be interpreted as a soft dot product between the sorted Doppler values and the Gaussian weight distribution. Given the thresholdτt, each Doppler valuevi t is compared against it to obtain a soft motion indicator: si t =σ \u0012 vi t −τ t γ \u0013 , (2) whereγcontrols the sharpness of the transition. The resulting scores are then binarized using a Straight-Through Estimator (STE) [4,10], maintaining ex- plicit partitioning in the forward pass while propagating gradients through the soft probabilitiessi t i","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (14 contexts).","role_counts":[{"n":14,"context_role":"method"},{"n":8,"context_role":"background"}]},"error":null,"updated_at":"2026-05-19T13:21:29.435305+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"5a7a1034-5ab5-46b2-9f24-66204fbd9091","orcid":null,"display_name":"Yoshua Bengio"},{"id":"3379e339-092f-45c5-bea1-cc99f846a75a","orcid":null,"display_name":"Nicholas L\\'eonard"},{"id":"038e7d59-3d4b-4bfc-9006-867c776aec8f","orcid":null,"display_name":"Aaron Courville"}]},"error":null,"updated_at":"2026-05-19T13:21:29.687263+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T08:28:09.304600+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":7},{"title":"Categorical Reparameterization with Gumbel-Softmax","work_id":"050fb4ab-62d4-4150-866a-77bc502eca22","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":4},{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":4},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":3},{"title":"and Bengio, Y","work_id":"73fcd90c-53bb-4d2e-87ef-284402d02867","shared_citers":3},{"title":"//arxiv.org/abs/1811.04551","work_id":"146fc4a4-6db2-43a9-a57f-4dd133c0d315","shared_citers":3},{"title":"arXiv preprint arXiv:2408.15664 , year=","work_id":"267500ca-1512-478f-8a1b-6ecbdb09771d","shared_citers":3},{"title":"Bitnet: Scaling 1-bit transformers for large language models","work_id":"28ad8f61-4291-4894-b120-1d42fc9937a3","shared_citers":3},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":3},{"title":"Continuous control with deep reinforcement learning","work_id":"41a65444-c819-4303-a1f1-b075aa86d40c","shared_citers":3},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":3},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":3},{"title":"Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients","work_id":"ff82bd1e-0b64-4426-81a0-91594dba0a20","shared_citers":3},{"title":"Fast and accurate deep network learning by exponential linear units (elus)","work_id":"619b409a-ddb8-4401-b55e-d8ca367322ce","shared_citers":3},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":3},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":3},{"title":"Jaderberg, V","work_id":"8ac80d7b-b24c-4128-9798-b3a9b02ad56d","shared_citers":3},{"title":"Kaiser, M","work_id":"edc1a23e-c421-4569-ab9e-83b204eeb0fa","shared_citers":3},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":3},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":3},{"title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer","work_id":"2c6b3f6d-54e4-4df7-baa7-475a490799af","shared_citers":3}],"time_series":[{"n":1,"year":2016},{"n":1,"year":2017},{"n":2,"year":2019},{"n":2,"year":2022},{"n":1,"year":2023},{"n":68,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T08:27:56.049360+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T08:28:21.095576+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","claims":[{"claim_text":"Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we \"back-propagate\" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A s","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"𝑖 =𝑓 𝜃 (𝑥 tr 𝑖 ),( ˆ𝑧tr 𝑖 , 𝑠tr 𝑖 )=𝑄(𝑧 tr 𝑖 ), ˜𝑥 tr 𝑖 =ℎ 𝜙 ( ˆ𝑧tr 𝑖 ), 𝑧ta 𝑖 =𝑓 𝜃 (𝑥 ta 𝑖 ),( ˆ𝑧ta 𝑖 , 𝑠ta 𝑖 )=𝑄(𝑧 ta 𝑖 ), ˜𝑥 ta 𝑖 =ℎ 𝜙 ( ˆ𝑧ta 𝑖 ). (1) Here, ˆ𝑧tr 𝑖 and ˆ𝑧ta 𝑖 are the quantized embeddings, while 𝑠tr 𝑖 and 𝑠ta 𝑖 are the corresponding SIDs. Since nearest-codeword lookup is non- differentiable, we adopt the straight-through estimator (STE) [2] to pass gradients from the quantized embeddings to the encoder-side representations during back-propagation. The forward pipeline above de","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"yielding superior rendering quality and improved efficiency for generalizable novel view synthesis. C. Dynamic Neural Networks. Dynamic neural networks [54]-[61] are intended to adap- tively adjust their weights or structure to handle given input with appropriate states, offering a more flexible alternative to static architectures. Recently, this paradigm has evolved from basic conditional computation [62] toward sophisticated Mixture-of- Experts (MoE) architectures, which effectively scale mode","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"machine learning, pages 3734-3743. pmlr, 2019. [51] Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax.arXiv preprint arXiv:1611.01144, 2016. [52] Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation.arXiv preprint arXiv:1308.3432, 2013. [53] Miles Macklin. Warp: A high-performance python framework for gpu simulation and graphics, March 2022. NVIDIA GPU Technology Co","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"We do not anneal this threshold and do not use additional entropy or sparsity regularization on the gate. The final linear head is zero-initialized so that the gate starts from a neutral routing distribution. The forward routing decision is binary, Gl i = round(pl i),(2) and gradients are passed through this binary threshold with a straight-through estimator [ 4]. The effective PE for subject tokens is then P El eff,i =G l i P Ebase,i + (1−G l i)P Eswap,i, i∈ S l,(3) where P Ebase denotes the or","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"to model p by composition of (P q)⊤ (q→u ) and P p (u→p ). This way, the objective of eq. (3.2) becomes nX p̸=q LX i=1 ⟨W p i , P p i (P q i )⊤W q i (P p i−1(P q i−1)⊤)⊤⟩= nX p̸=q LX i=1 ⟨(P p i )⊤W p i P p i−1,(P q i )⊤W q i P q i−1⟩. (3.3) As stated by Theorem 3.2.1, the permutations we obtain using eq. (3.3) are cycle consistent. We refer the reader to Bernard et al.[15] for the proof and a complete discussion of the subject. Theorem 3.2.1(Restated from Bernard et al. 15).Given a set of n mod","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"be interpreted as a soft dot product between the sorted Doppler values and the Gaussian weight distribution. Given the thresholdτt, each Doppler valuevi t is compared against it to obtain a soft motion indicator: si t =σ \u0012 vi t −τ t γ \u0013 , (2) whereγcontrols the sharpness of the transition. The resulting scores are then binarized using a Straight-Through Estimator (STE) [4,10], maintaining ex- plicit partitioning in the forward pass while propagating gradients through the soft probabilitiessi t i","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (14 contexts).","role_counts":[{"n":14,"context_role":"method"},{"n":8,"context_role":"background"}]},"error":null,"updated_at":"2026-05-19T13:21:29.692661+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","claims":[{"claim_text":"Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we \"back-propagate\" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A s","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T08:28:04.630551+00:00"}},"summary":{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","claims":[{"claim_text":"Stochastic neurons and hard non-linearities can be useful for a number of reasons in deep learning models, but in many cases they pose a challenging problem: how to estimate the gradient of a loss function with respect to the input of such stochastic or non-smooth neurons? I.e., can we \"back-propagate\" through these stochastic neurons? We examine this question, existing approaches, and compare four families of solutions, applicable in different settings. One of them is the minimum variance unbiased gradient estimator for stochatic binary neurons (a special case of the REINFORCE algorithm). A s","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":7},{"title":"Categorical Reparameterization with Gumbel-Softmax","work_id":"050fb4ab-62d4-4150-866a-77bc502eca22","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":4},{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":4},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":3},{"title":"and Bengio, Y","work_id":"73fcd90c-53bb-4d2e-87ef-284402d02867","shared_citers":3},{"title":"//arxiv.org/abs/1811.04551","work_id":"146fc4a4-6db2-43a9-a57f-4dd133c0d315","shared_citers":3},{"title":"arXiv preprint arXiv:2408.15664 , year=","work_id":"267500ca-1512-478f-8a1b-6ecbdb09771d","shared_citers":3},{"title":"Bitnet: Scaling 1-bit transformers for large language models","work_id":"28ad8f61-4291-4894-b120-1d42fc9937a3","shared_citers":3},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":3},{"title":"Continuous control with deep reinforcement learning","work_id":"41a65444-c819-4303-a1f1-b075aa86d40c","shared_citers":3},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":3},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":3},{"title":"Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients","work_id":"ff82bd1e-0b64-4426-81a0-91594dba0a20","shared_citers":3},{"title":"Fast and accurate deep network learning by exponential linear units (elus)","work_id":"619b409a-ddb8-4401-b55e-d8ca367322ce","shared_citers":3},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":3},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":3},{"title":"Jaderberg, V","work_id":"8ac80d7b-b24c-4128-9798-b3a9b02ad56d","shared_citers":3},{"title":"Kaiser, M","work_id":"edc1a23e-c421-4569-ab9e-83b204eeb0fa","shared_citers":3},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":3},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":3},{"title":"Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer","work_id":"2c6b3f6d-54e4-4df7-baa7-475a490799af","shared_citers":3}],"time_series":[{"n":1,"year":2016},{"n":1,"year":2017},{"n":2,"year":2019},{"n":2,"year":2022},{"n":1,"year":2023},{"n":68,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"038e7d59-3d4b-4bfc-9006-867c776aec8f","orcid":null,"display_name":"Aaron Courville","source":"manual","import_confidence":0.72},{"id":"3379e339-092f-45c5-bea1-cc99f846a75a","orcid":null,"display_name":"Nicholas L\\'eonard","source":"manual","import_confidence":0.72},{"id":"5a7a1034-5ab5-46b2-9f24-66204fbd9091","orcid":null,"display_name":"Yoshua Bengio","source":"manual","import_confidence":0.72}]}}