Pith. sign in

Paper Citation Record · LEDGER

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.22578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22578 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:00.259466Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf65120c-5e39-42e5-b796-7692035685d2 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, and Ruosong Wang.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Wei Hu, Zhiyuan Li, and Ruosong Wang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:09.203822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.332413Z digest=sha256:86f8c61adfa45f894ca321e94707ed19482ee9f43266d451407c2d25dc71b7c8

Observation 83dd3c12-b08d-4f3a-81a2-57cb4c270077 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.939482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.377545Z digest=sha256:2e03dd644fe0e027d335f49a221a2c75f45e3b2553af0b610d47aa4a5f1cc88b

Observation 74c7d6ca-ef8d-41ef-a13f-d5ba1e3ad86e · outbound

This paper cites Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.668772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.451054Z digest=sha256:b2b91410340ac4d29e0ddbfcf1a1b5335209f38c288a0b71da77c2d48e5181b7

Observation 9c1ef4dc-9322-4aff-a239-1c5c4fb6f1ca · outbound

This paper cites Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T13:13:01.118361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.531975Z digest=sha256:2a887f13722ea7bd17a81c4d824b6a944dc242393f76b79a76d185edb1e2836c

Observation f6a9376d-c56c-4135-ae90-47898d9cfd9c · outbound

This paper cites Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.496137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.590775Z digest=sha256:acb1f300e633fe5cfc1b74a46bf1b3e4ee55d296757d47e924a773e68f8039d3

Observation 792d70e4-35af-48d8-be78-090179834cc0 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.357491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.664933Z digest=sha256:06a9299e1bf51e436612141c1e679edfaf4cdf15b2c4a9e3cd16a400bc7cd1bf

Observation 8f0db553-e388-4387-8853-a948e9210d09 · outbound

This paper cites How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.217059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.756758Z digest=sha256:d59b849bfcc97702b90be4620c9267ed272159c3805712de5e33fa12ed5c3226

Observation 9b1854d3-34d3-4136-9f96-688a9f2403c1 · outbound

This paper cites Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.843877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.914089Z digest=sha256:829a9fb58fac10b018a931377015fd37b96c91d9aacd528cae690d34d1619728

Observation 03f525e9-9160-44a3-a045-08d516f2efe6 · outbound

This paper cites Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.001695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.032252Z digest=sha256:25084a7cd83482fc8c1b8f3c9ace704ab8997f8688844fbe6011cd20850fab6b

Observation 071cdf86-ea16-4c38-a119-33a3a0a780ff · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.949186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.152527Z digest=sha256:9d112be33754a048b4b52cb737ecf2531cf46f6d68a7c0d78c594f310f80c5cc

Observation 38f3a440-21cb-4a86-8869-33e040ebb0ae · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.918542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.306196Z digest=sha256:0ab6452530499bd812c3605a170955d9b76f46b4263d463ccc5cc01ec78d5f84

Observation 64cdbab3-38e7-43c0-a369-5bd2ede0cb6c · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.701863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.474552Z digest=sha256:02c573a7719a706b16c97bed50dde5a6b146d4d2583a5d73f582d6ac06100c1e

Observation e30d8cd1-5d7c-4bdb-91ac-65168652e6dd · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.590174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.590174Z digest=sha256:a50912aa5e1c871758832adc109365425c2991732c29062ac1cbbe4a5ce62484

Observation b4bd2716-0494-4d02-bfb8-c1da6de70724 · outbound

This paper cites Bach, and Loucas Pillaud - Vivien.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bach, and Loucas Pillaud - Vivien

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.630343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.704138Z digest=sha256:6db7dc226f4bded3b7296803a62dc35bdb6a97d16a5d6fb92402a63dfcd96270

Observation e8e52731-abf1-486f-b05d-36faeb95cf96 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.503025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.846688Z digest=sha256:f9936820913979275d1c0a907c2121ed9f75e744befcedbf27e27d1cbcce29bd

Observation 79819969-13bd-4fc5-b1f9-a5e6b3778476 · outbound

This paper cites Kakade, and Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Kakade, and Jason D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.969941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.969941Z digest=sha256:a89297c5a659a66bcd62dde3e543372e519ab670d60cce0440ccbf56b650ce44

Observation 174a3a10-6587-46e7-8fa9-61538a6fc583 · outbound

This paper cites https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.256033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.058268Z digest=sha256:05b3ea10a9e9fa1fbce9c49962605a9b7fd3603af2fe9915b92539d926439dff

Observation 28dc0c1e-59b1-41da-9bd7-e02db3e4f05f · outbound

This paper cites Hamprecht.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Hamprecht

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.951243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.230864Z digest=sha256:ef82646a4265e3734b21364ac0dffc0bc7ab6eac496a99904b4774e214f2b4cb

Observation 2e16810a-a1fd-4a45-a54e-25e522cbd06a · outbound

This paper cites Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.708498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.343084Z digest=sha256:f1b20e761d177a6a8e35912f75f7fe3eb3e7312075eb1b07f82242f1c58c6d5e

Observation 0b01b0f9-19ea-4923-afb6-ec89ca45688d · outbound

This paper cites Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.540935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.443251Z digest=sha256:f2bed80523bbdfdf976e4c67926bf120762a0bea9edbf58ceb70eba53cfbbba0

Observation b7873b5f-3ba9-4ee0-8a70-c6cd7daa7e7c · outbound

This paper cites An introduction to probability theory and its applications, volume 2.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization An introduction to probability theory and its applications, volume 2

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.364081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.535095Z digest=sha256:f871db510927953574b404c31f17436166599515716f5ddda31337a7d9a9dbba

Observation 51ada76a-9c7e-49b9-b168-d8626d6c5b38 · outbound

This paper cites Vetrov, and Andrew Gordon Wilson.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Vetrov, and Andrew Gordon Wilson

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.086766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.639192Z digest=sha256:4b7546712062dccdd912fcf1cc72a64a69711a00894c638e0d38ef3ddc378192

Observation f8d588e8-3ef4-433d-ae93-2cedfb5b5b9a · outbound

This paper cites https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.911555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.757397Z digest=sha256:9e12961d3f80e371247398c50741a3dafd24c0f0226e3856c9b0c6a061a6c0f4

Observation 28a6c832-38c5-4355-9f20-5bb655161c75 · outbound

This paper cites Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.787591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.826153Z digest=sha256:4af46dffef246957b1e43c1d23be51eaa852bd7e0884d0b0d10f79d095f9fcd6

Observation febe0f18-62a5-4f05-a621-8c1522f71328 · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.938493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:56.938493Z digest=sha256:40f1d8949b8170fc9bfd54bb18ca83664fbb272ae8e5104b7cd29edd0c7d8ada

Observation a96253ce-1c46-4c1f-b9cc-8169e7197c6b · outbound

This paper cites Haeffele and Ren \' e Vidal.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Haeffele and Ren \' e Vidal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.023287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.023287Z digest=sha256:84acd73420f59b9988fad58067b54163e25566851ecabca754ab2321a562baf4

Observation 12dd8075-579f-4650-b68b-bc03a8f58350 · outbound

This paper cites Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.676439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.197644Z digest=sha256:57567835b37338117270de4b5a3f12104ea7013d381e499c9067acb54bed8856

Observation c6928d55-1f75-46db-bb7c-805e90b5bf41 · outbound

This paper cites Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.324053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.324053Z digest=sha256:4e3c51d8f88eaff5306e088986b7f22e6c778f890cca714d5417181281814551

Observation 752329bc-34da-4c9c-97ac-02aaa5a0259e · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.456856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.442933Z digest=sha256:e1dc54e16038d1717f0e753675cd91f97b120527c359c5296df743e7ae877325

Observation 16283734-4f5a-42f1-bc36-49195400c4fa · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.560687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.560687Z digest=sha256:c351b8262c8a88e036b1c766e01e919330fd9edec7c1f3b0beab6b1229d83d44

Observation f4c597be-e459-448e-9c63-6d5b0870e528 · outbound

This paper cites Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.283568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.708312Z digest=sha256:40ff5806b122f4f9eb15663a23c1be8d66d10db350e4f06257df7b4ac28ab403

Observation 46685940-4451-4b0e-8239-74ac682d9155 · outbound

This paper cites Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.082243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.835214Z digest=sha256:93ffd119ae0f86d5b48265121088722aa9bc9942d12898344d993faafd8569f0

Observation eee97ce2-6b4e-4dad-983b-60b3b5db89fb · outbound

This paper cites Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.832441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.920365Z digest=sha256:d21fdf612e90bbcbeab228c7ae5a19b40ab477332861b4e8420f2b843a5ecd3d

Observation 2dc68690-56d3-48db-8bf9-d4c0fb558842 · outbound

This paper cites Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.610229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.982098Z digest=sha256:4711bcd48fd12649829c1fca507ab03bac1566ea49c61dced3ba4548cb6d8bf7

Observation e05247a5-00b8-498d-a4fc-57a01db631b3 · outbound

This paper cites Michaud, and Max Tegmark.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Michaud, and Max Tegmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.456563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.075285Z digest=sha256:cf6d7c8b3b5d4bc9fbe1860def5d637d6588fd2a04b479ce4d68947c80b1fc81

Observation 79238af5-cf16-4216-898e-969faaaa50b1 · outbound

This paper cites Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.326855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.144580Z digest=sha256:c69e2f627f1fc19aaa509f1697e2c5217f3979bac7d2d67f61c25f7dc12a4073

Observation 7dcba60d-5130-4eb8-a06f-75bed81ef595 · outbound

This paper cites Lee, and Wei Hu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Lee, and Wei Hu

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.151036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.265403Z digest=sha256:61882a865f2cae04f931103f9322acc58c7adfea2722be215acc515159400d6a

Observation 7f4f4e97-1a83-4c3e-95aa-f8f1cb223d96 · outbound

This paper cites Gradient Descent Quantizes ReLU Network Features.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent Quantizes ReLU Network Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.348208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.348208Z digest=sha256:31174eb3a6b4499c8e597fef67c6ff236a3d961e30d6a31ff33506339aafdb01

Observation bb55700f-f0c9-4f44-a56d-d5e518060c56 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.426905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.426905Z digest=sha256:8aa606a3dfff76590c60a3c0799b101034453c1d6ab42873398340f7db5667e1

Observation 349c5a1b-fa92-42ed-b34e-603a66389385 · outbound

This paper cites Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.976474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.544465Z digest=sha256:cb8202b0671c1946fa5b3b56a4c7b83eaa4dff0e8cf4b629071842bc85f88e95

Observation 279e937b-4b17-45e4-bad8-8d78436d72b9 · outbound

This paper cites Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.780828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.664378Z digest=sha256:7e6b3c1be250f6eaa59f54285b2101f98af921dfebc4bdb9728b75372eb80c8c

Observation 14ad19bc-d854-492b-9efd-5753fc7dc9e5 · outbound

This paper cites In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.738470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.738470Z digest=sha256:a12f57a6a28fba04d0eb286032beefd064270d368a9a5d64603241d3591c7d94

Observation 7b1bb096-e0cc-491c-b36e-c123b365e76b · outbound

This paper cites On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.632268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.852854Z digest=sha256:05beaf949403c72906b2a87bafd242ecc52c64d8096f9b1e2c52df94b59c6686

Observation 9331022f-4f6d-48ce-b281-a344eee577ab · outbound

This paper cites A Note on Connectivity of Sublevel Sets in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A Note on Connectivity of Sublevel Sets in Deep Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.964229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.964229Z digest=sha256:3e813ba01605e49610fb7f66b8d55262ee2867a4a43e9b1a02213ab22e2721da

Observation 0d225fce-7953-4a55-ae08-b58b9f4cdc38 · outbound

This paper cites When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.463658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.027410Z digest=sha256:fc6de8dd244e51bcddb5c631b6e9d32927ed8bd3be3776aff7038e26c83cb864

Observation ded24de2-dbb3-4229-b5a3-351500f6c86a · outbound

This paper cites Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:13:01.452769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.146372Z digest=sha256:446434fded84c7f59407e19a2fbbb443acaa7676441483dbbbb447ede49c3d25

Observation 027e40ea-cedb-4cb4-bbe5-9e2f036d8106 · outbound

This paper cites Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.250869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.231078Z digest=sha256:a6fe8df90a1bcd33f975ef46b21d356b948612bff45738b5e391f1cf643a2d22

Observation af995fc2-0e66-43c6-827f-2fdf0ae520e5 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.316611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.316611Z digest=sha256:8d163025df6e290651149b68e66fd6c8fd5db46ba3dc202470b481ada6d6513c

Observation 147fe6d0-4954-400d-b601-98d8cdd4522c · outbound

This paper cites Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.419770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.419770Z digest=sha256:052adccef84aff25f1c4a8e487d4d665994748c316e5c1fd4b5df2d6b6617d3d

Observation a7b54f09-d21b-4f87-9ddd-80e61a9a7a81 · outbound

This paper cites Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.111915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.498281Z digest=sha256:33ba78f51c8507cd433c7e57112add3351ca6824a970fe979170b9744d198de9

Observation f60e7662-4104-42ba-b431-718fd99987c1 · outbound

This paper cites How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.939431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.586553Z digest=sha256:487554391387a2139fdae10032dc7bd2ce5b7671f37208dfeec5dfe31b1cd4e8

Observation 52ee293f-9430-491e-915b-45413cc09286 · outbound

This paper cites Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.653892Z digest=sha256:d88fcd0de4c8ec9cb3e610853cbdc52eb0be1d23a39f1ac3a3db0f62a7bc0f68

Observation 3a176d66-6b8b-45b8-83d3-617a6339a7b5 · outbound

This paper cites Jamaloddin Golestani.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Jamaloddin Golestani

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.705425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.707457Z digest=sha256:3573ff105478eb1bbf39574768f8d294cc04fe8621a4f42de0285ed8c45ef022

Observation 0855d4a2-ac27-4f9b-bfdc-55c07735c983 · outbound

This paper cites Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.467783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.787288Z digest=sha256:69ffb4405ce4ab0f2fe971496b366728e98f9ad9eece04cc83a8b973e9127ba3

Observation 8fc4d611-2594-4d34-ab6c-1d366d687115 · outbound

This paper cites The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.866232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.866232Z digest=sha256:1076851967057565b2b247c687eff903ce9028b23a9e27d74552d9fa15a60c9a

Observation 42386525-5e0c-41b5-b985-efdf1ddaf1e1 · outbound

This paper cites Bandeira, and Joan Bruna.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bandeira, and Joan Bruna

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.276608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.972636Z digest=sha256:4a47752fbcbb400b82552bee37b83cfba80eca96070c98d31c4822663645e9d2

Observation c461e396-90eb-48de-954c-3a24e8ad053d · outbound

This paper cites The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.091759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.068517Z digest=sha256:95afd38739742be6897ff2355b81001307a84d62136852d649f2bfed0db4a0c9

Observation 064d0a4b-9a0b-4ce9-a5ac-6bb0e04fee23 · outbound

This paper cites On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:00.129590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:13:00.129590Z digest=sha256:daf982ce5698ed84bec4b6fca0db263ac8cfdd1383ef8e3638bc748a4a3c15e7

Observation 047842c8-cae0-4b6f-88e8-63aebf6436ef · outbound

This paper cites Woodworth, Suriya Gunasekar, Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Woodworth, Suriya Gunasekar, Jason D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.882611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.191389Z digest=sha256:2e6ae34f105d632983fd9810980f4cbf5e21c6c456633f4f828b8b26846cfd7f

Observation a1e10686-1355-4d63-8293-7c13b534aa61 · outbound

This paper cites Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.699375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.259466Z digest=sha256:4f40713bb037c9ab64c47d48fd47dbac1cb3368905e460b1042e6c03684e7944

Pith citing papers

No inbound Pith citation observations are available.