Pith. sign in

Paper Citation Record · LEDGER

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.22578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22578 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:00.259466Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf65120c-5e39-42e5-b796-7692035685d2 · outbound

This paper cites Du, Wei Hu, Zhiyuan Li, and Ruosong Wang.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Wei Hu, Zhiyuan Li, and Ruosong Wang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:09.203822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.332413Z digest=sha256:fec04e97053db4dbad4148852d9cde3a09d77cc50356e7965e6903da454b7ce5

Observation 83dd3c12-b08d-4f3a-81a2-57cb4c270077 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.939482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.377545Z digest=sha256:bb78fcb12e63213744243b285bec0456f9aa1854114e158ff3bb95eb690233ef

Observation 74c7d6ca-ef8d-41ef-a13f-d5ba1e3ad86e · outbound

This paper cites Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Penalising the biases in norm regularisation enforces sparsity https://papers.neurips.cc/paper_files/paper/2023/hash/b444ad72520a5f5c467343be88e352ed-Abstract-Conference.html

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.668772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.451054Z digest=sha256:e0e4dcf961fc8f9fb772e8115d1e20781ded0cb591fed5db9d9869766eafcd98

Observation 9c1ef4dc-9322-4aff-a239-1c5c4fb6f1ca · outbound

This paper cites Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early alignment in two-layer networks training is a two-edged sword 10.48550/arxiv.2401.10791

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T13:13:01.118361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.531975Z digest=sha256:a240e3a1f4a9cb6f78e9245adaae1ad523768e4ca32e5d36538538ef42ddafc9

Observation f6a9376d-c56c-4135-ae90-47898d9cfd9c · outbound

This paper cites Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Simplicity bias and optimization threshold in two-layer ReLU networks https://openreview.net/forum?id=qAarsvflTa

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.496137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.590775Z digest=sha256:c1065c189005029d394861efd28782f2647c05d8a28e63ece47c8452fc282d1e

Observation 792d70e4-35af-48d8-be78-090179834cc0 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:08.357491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.664933Z digest=sha256:d21066c37b0ae7f8b5d86b6985df87300854b1260fa5ff087611e60a018f5ead

Observation 8f0db553-e388-4387-8853-a948e9210d09 · outbound

This paper cites How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers https://openreview.net/forum?id=3eHNvPHL9Z

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.217059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.756758Z digest=sha256:801c50b4a001abcc9787e9de5dcf5a745229586c073547bf91bdb40b303dec2b

Observation 9b1854d3-34d3-4136-9f96-688a9f2403c1 · outbound

This paper cites Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convergence of gradient descent for deep neural networks 10.48550/arxiv.2203.16462

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.843877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:54.914089Z digest=sha256:adb21ccb2d7462e67403bf049af97fe20baeb37ba85ba36b2579b9a3f09b9771

Observation 03f525e9-9160-44a3-a045-08d516f2efe6 · outbound

This paper cites Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Loss Landscapes are All You Need: Neural Network Generalization Can Be Explained Without the Implicit Bias of Gradient Descent https://openreview.net/forum?id=QC10RmRbZy9

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:08.001695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.032252Z digest=sha256:7fa724368e184b0d6bed94db0a7fbf009df5da860640f07275f28cf7ec62ddd5

Observation 071cdf86-ea16-4c38-a119-33a3a0a780ff · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.949186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.152527Z digest=sha256:443674bbfbcd16af9d325d6f1f3cb03772875c26e1038483caeb3feb3e97b746

Observation 38f3a440-21cb-4a86-8869-33e040ebb0ae · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.918542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.306196Z digest=sha256:f41999d3d8958e7b3efaaaf086fc1fc64a6c19f03db4d49662f3f7be45125bd9

Observation 64cdbab3-38e7-43c0-a369-5bd2ede0cb6c · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.701863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.474552Z digest=sha256:c97dfcd771d3d716aa61b4f46055c3e093f08d38e07eafc6ffab856bd2109c05

Observation e30d8cd1-5d7c-4bdb-91ac-65168652e6dd · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.590174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.590174Z digest=sha256:1d967eeb745fcbde4705433422340fa493cde78dfe87de5f629ddee4837a0bd1

Observation b4bd2716-0494-4d02-bfb8-c1da6de70724 · outbound

This paper cites Bach, and Loucas Pillaud - Vivien.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bach, and Loucas Pillaud - Vivien

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T13:13:00.630343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.704138Z digest=sha256:b4791c35493eb0c9f7d0baf7250f204c50955074cafa4420ed552751a105b285

Observation e8e52731-abf1-486f-b05d-36faeb95cf96 · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:13:07.503025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:55.846688Z digest=sha256:4d67cd82cdf3f867321a4058d420a23c7121b406cafc0873b03484f993b63ff3

Observation 79819969-13bd-4fc5-b1f9-a5e6b3778476 · outbound

This paper cites Kakade, and Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Kakade, and Jason D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.969941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.969941Z digest=sha256:9c731eb6f5740edd320734e1a6b67713c8123f48a42b37d1beacaecfd4007f52

Observation 174a3a10-6587-46e7-8fa9-61538a6fc583 · outbound

This paper cites https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://www.jmlr.org/papers/v17/15-408.html CVXPY : A P ython-embedded modeling language for convex optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:07.256033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.058268Z digest=sha256:094d0f59d885ce437508bd40eb9e6990f2e40de6fc1cc44c5d8aecc1bf754fc5

Observation 28dc0c1e-59b1-41da-9bd7-e02db3e4f05f · outbound

This paper cites Hamprecht.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Hamprecht

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.951243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.230864Z digest=sha256:1518e618cb51230e13839cf7082ecaf75c454b1062197244624a011d750fcc6e

Observation 2e16810a-a1fd-4a45-a54e-25e522cbd06a · outbound

This paper cites Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Du, Xiyu Zhai, Barnab \' a s P \' o czos, and Aarti Singh

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.708498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.343084Z digest=sha256:83ccc1913b34a1ee190873903714221165c318ac9a274c259f2dee9305520898

Observation 0b01b0f9-19ea-4923-afb6-ec89ca45688d · outbound

This paper cites Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Convex Geometry and Duality of Over-parameterized Neural Networks http://jmlr.org/papers/v22/20-1447.html

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.540935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.443251Z digest=sha256:abf35abe2739feb2305af06988a9ca151f93a531bb1463809a752bdd4b3c79eb

Observation b7873b5f-3ba9-4ee0-8a70-c6cd7daa7e7c · outbound

This paper cites An introduction to probability theory and its applications, volume 2.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization An introduction to probability theory and its applications, volume 2

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.364081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.535095Z digest=sha256:1ed8fa769ccbeee07cf2ef1325897c26967e886c6289feb138b8ba797c823c96

Observation 51ada76a-9c7e-49b9-b168-d8626d6c5b38 · outbound

This paper cites Vetrov, and Andrew Gordon Wilson.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Vetrov, and Andrew Gordon Wilson

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:06.086766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.639192Z digest=sha256:89dfad6a906ad689eb854e436f85baf19ed339f8fb949e1ebf5020e9b429baf2

Observation f8d588e8-3ef4-433d-ae93-2cedfb5b5b9a · outbound

This paper cites https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization https://openreview.net/forum?id=HgOJlxzB16 SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.911555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.757397Z digest=sha256:2d7e11aa8bb2ac34b2e81d1d96558ddae576e296c48be4b375fc243518508e81

Observation 28a6c832-38c5-4355-9f20-5bb655161c75 · outbound

This paper cites Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Truth or backpropaganda? An empirical investigation of deep learning theory https://openreview.net/forum?id=HyxyIgHFvr

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.787591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:56.826153Z digest=sha256:322b094276b3f200684fa6a79778854cad4f0ebb096e9a2207927e36c2ee92a6

Observation febe0f18-62a5-4f05-a621-8c1522f71328 · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:56.938493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:56.938493Z digest=sha256:f9bd51e3d4ab8b3cbc1a44c8ab677a3e54643aaecaa52420bfc9b37e592b696d

Observation a96253ce-1c46-4c1f-b9cc-8169e7197c6b · outbound

This paper cites Haeffele and Ren \' e Vidal.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Haeffele and Ren \' e Vidal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.023287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.023287Z digest=sha256:b751cdece35fea14660ea20a99bb1f1074660bf5a73a5228a1b878456b0a18bc

Observation 12dd8075-579f-4650-b68b-bc03a8f58350 · outbound

This paper cites Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Piecewise linear activations substantially shape the loss surfaces of neural networks https://openreview.net/forum?id=B1x6BTEKwr

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.676439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.197644Z digest=sha256:b5735fdf055238900b7f3dc79ad0ace4a35f022a6ebcd09fc39b033947205ffe

Observation c6928d55-1f75-46db-bb7c-805e90b5bf41 · outbound

This paper cites Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Residual Learning for Image Recognition 10.1109/cvpr.2016.90

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.324053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.324053Z digest=sha256:0e6ec782b3b18e7d8eda6666821aaf1cd0b85ad3f9b5c6d15c71e1a72cd36c3d

Observation 752329bc-34da-4c9c-97ac-02aaa5a0259e · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Tangent Kernel: Convergence and Generalization in Neural Networks https://proceedings.neurips.cc/paper/2018/hash/5a4be1fa34e62bb8a6ec6b91d2462f5a-Abstract.html

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.456856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.442933Z digest=sha256:5a32dbc6153ffab98929b7dd9e856eb3e99009efa40b961dbc0d21bf161ef5a0

Observation 16283734-4f5a-42f1-bc36-49195400c4fa · outbound

This paper cites an unresolved cited work.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:57.560687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:57.560687Z digest=sha256:fc4e782aea2b19cd56c31de9fc632dd311e61103e9563702e467695414101891

Observation f4c597be-e459-448e-9c63-6d5b0870e528 · outbound

This paper cites Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Mildly Overparameterized ReLU Networks Have a Favorable Loss Landscape https://openreview.net/forum?id=10WARaIwFn

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.283568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.708312Z digest=sha256:9bebab23c0fb0f140b77ca2c97c55d5e865da5960a8257b0fcbd8ac88aff7a67

Observation 46685940-4451-4b0e-8239-74ac682d9155 · outbound

This paper cites Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Learning without Poor Local Minima https://proceedings.neurips.cc/paper/2016/hash/f2fc990265c712c49d51a18a32b39f0c-Abstract.html

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:05.082243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.835214Z digest=sha256:6d34b3c57765b97614cfd341b2808634f337acb94aaa44d8b87bd2f5a44f3fa8

Observation eee97ce2-6b4e-4dad-983b-60b3b5db89fb · outbound

This paper cites Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Exploring The Loss Landscape Of Regularized Neural Networks Via Convex Duality https://openreview.net/forum?id=4xWQS2z77v

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.832441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.920365Z digest=sha256:45e2b4da6ee17c24d3e7c8fe9e95ce4dd1b5edbf75e6cba2bf67b3723b27c9e7

Observation 2dc68690-56d3-48db-8bf9-d4c0fb558842 · outbound

This paper cites Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Deep Linear Networks with Arbitrary Loss: All Local Minima Are Global http://proceedings.mlr.press/v80/laurent18a.html

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.610229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:57.982098Z digest=sha256:4e249fee35830f95f44408fe31199fc4b4022c9ac8641a8df9281a51573d8971

Observation e05247a5-00b8-498d-a4fc-57a01db631b3 · outbound

This paper cites Michaud, and Max Tegmark.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Michaud, and Max Tegmark

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.456563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.075285Z digest=sha256:9da02d4df3014141ea8f87b88fe3146e4279119e0e413cda2f19e7dbc0dbc151

Observation 79238af5-cf16-4216-898e-969faaaa50b1 · outbound

This paper cites Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias https://proceedings.neurips.cc/paper/2021/hash/6c351da15b5e8a743a21ee96a86e25df-Abstract.html

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.326855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.144580Z digest=sha256:b8d3f7a6ff1bc12558f94f93be6d4fd448d640f77fc2e294d8bec99a04bd5a89

Observation 7dcba60d-5130-4eb8-a06f-75bed81ef595 · outbound

This paper cites Lee, and Wei Hu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Lee, and Wei Hu

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:04.151036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.265403Z digest=sha256:a9332b981d8b0a7692f107f6643ebabb9d87840add5f65c60dbb26520e5828b4

Observation 7f4f4e97-1a83-4c3e-95aa-f8f1cb223d96 · outbound

This paper cites Gradient Descent Quantizes ReLU Network Features.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Gradient Descent Quantizes ReLU Network Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.348208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.348208Z digest=sha256:74253fc282efc7a501d2dcb80d024146612c99a4b02a058dcac5b3be35dd08c8

Observation bb55700f-f0c9-4f44-a56d-d5e518060c56 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A mean field view of the landscape of two-layer neural networks 10.1073/pnas.1806579115

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.426905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.426905Z digest=sha256:f30e293952126d251b7a0fc6109bffa6afca64402ccc5e05a8d2488511cdec78

Observation 349c5a1b-fa92-42ed-b34e-603a66389385 · outbound

This paper cites Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization https://openreview.net/forum?id=QibPzdVrRu

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.976474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.544465Z digest=sha256:108cfd0ee714cfd66259c230f4a59816c5a1200a7dcc7ccd3f3d5ca70b371a0e

Observation 279e937b-4b17-45e4-bad8-8d78436d72b9 · outbound

This paper cites Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Optimal Sets and Solution Paths of ReLU Networks https://proceedings.mlr.press/v202/mishkin23a.html

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.780828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.664378Z digest=sha256:4b034a4588f7d174b6332356489652839b6aac47898727cd62f87fc475189554

Observation 14ad19bc-d854-492b-9efd-5753fc7dc9e5 · outbound

This paper cites In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.738470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.738470Z digest=sha256:269ed58a57e82e227704868bb9354db5597e8af9d8b21926bc08ce4629b6468c

Observation 7b1bb096-e0cc-491c-b36e-c123b365e76b · outbound

This paper cites On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On Connected Sublevel Sets in Deep Learning http://proceedings.mlr.press/v97/nguyen19a.html

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.632268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:58.852854Z digest=sha256:034e87af4d0d1a21de6b11b84a62fd0d5262e5f1691a9e5842b220474c05fd88

Observation 9331022f-4f6d-48ce-b281-a344eee577ab · outbound

This paper cites A Note on Connectivity of Sublevel Sets in Deep Learning.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization A Note on Connectivity of Sublevel Sets in Deep Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:58.964229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:58.964229Z digest=sha256:ec89839f0ce2f3a388e8d630ed89be46bbbcc667f5b5eeea502120823abf1349

Observation 0d225fce-7953-4a55-ae08-b58b9f4cdc38 · outbound

This paper cites When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization When Are Solutions Connected in Deep Networks? https://proceedings.neurips.cc/paper/2021/hash/af5baf594e9197b43c9f26f17b205e5b-Abstract.html In NeurIPS, pages 20956--20969, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.463658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.027410Z digest=sha256:ef2642785231d99a5d405e36d5538ec1b50d4013be98671ba78b39cad060dc0d

Observation ded24de2-dbb3-4229-b5a3-351500f6c86a · outbound

This paper cites Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Banach space representer theorems for neural networks and ridge splines https://dl.acm.org/doi/10.5555/3546258.3546301

Reference 46

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:13:01.452769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.146372Z digest=sha256:aa30631109abbe7ef19c0666a5522fa9639b5f48e6e4b1d2e169d52d7f9f0fc5

Observation 027e40ea-cedb-4cb4-bbe5-9e2f036d8106 · outbound

This paper cites Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks http://proceedings.mlr.press/v119/pilanci20a.html

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.250869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.231078Z digest=sha256:593c3861ae0c5a89143fbb615b406eae2457fafd402e6305a3e10547f7b896f1

Observation af995fc2-0e66-43c6-827f-2fdf0ae520e5 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.316611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.316611Z digest=sha256:b9bc658f619b898b1ca09b44c1c9864579b816abe752d5c9d14c09e623b33476

Observation 147fe6d0-4954-400d-b601-98d8cdd4522c · outbound

This paper cites Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Trainability and accuracy of artificial neural networks: An interacting particle system approach 10.1002/cpa.22074

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.419770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.419770Z digest=sha256:da13e22ed5c5d8f18c3d762e6b45ff73626fc1e67a6f2f11570d37e55d867db6

Observation a7b54f09-d21b-4f87-9ddd-80e61a9a7a81 · outbound

This paper cites Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Spurious Local Minima are Common in Two-Layer ReLU Neural Networks http://proceedings.mlr.press/v80/safran18a.html

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:03.111915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.498281Z digest=sha256:8275cfbefc8d79c6744c124a2f5700fdf7f3645e75a00fc083247891effbd61c

Observation f60e7662-4104-42ba-b431-718fd99987c1 · outbound

This paper cites How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization How do infinite width bounded norm networks look in function space? http://proceedings.mlr.press/v99/savarese19a.html In COLT, pages 2667--2690, 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.939431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.586553Z digest=sha256:231d0b7466ab3a790c43e07ca9698a3d60d71b8db09686ee97f288abc4b0b3cd

Observation 52ee293f-9430-491e-915b-45413cc09286 · outbound

This paper cites Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Understanding machine learning: From theory to algorithms 10.1017/CBO9781107298019

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.653892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.653892Z digest=sha256:fe85d9c4118a50f0548d99a5d77f52d7459a466a7ff96185da869377abefc2cb

Observation 3a176d66-6b8b-45b8-83d3-617a6339a7b5 · outbound

This paper cites Jamaloddin Golestani.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Jamaloddin Golestani

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.705425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.707457Z digest=sha256:40c27e8d972a5adca0427321faf346c24a46102a9f867762c43622c185e01401

Observation 0855d4a2-ac27-4f9b-bfdc-55c07735c983 · outbound

This paper cites Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances http://proceedings.mlr.press/v139/simsek21a.html

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.467783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.787288Z digest=sha256:9a223d926d32c55e458055dc56089add84ced1388c382299583a0da2ccd5bd27

Observation 8fc4d611-2594-4d34-ab6c-1d366d687115 · outbound

This paper cites The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Global Landscape of Neural Networks: An Overview 10.1109/msp.2020.3004124

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:59.866232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:59.866232Z digest=sha256:9fdc3d567a4ff757ff94670ce2e11a42e390dae08bcaf8af7d76160033a22817

Observation 42386525-5e0c-41b5-b985-efdf1ddaf1e1 · outbound

This paper cites Bandeira, and Joan Bruna.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Bandeira, and Joan Bruna

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.276608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:12:59.972636Z digest=sha256:f58a932f34ff674c6966b43af42ce85458068280c7d1f1d2854d0ea754614662

Observation c461e396-90eb-48de-954c-3a24e8ad053d · outbound

This paper cites The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization The Hidden Convex Optimization Landscape of Regularized Two-Layer ReLU Networks: an Exact Characterization of Optimal Solutions https://openreview.net/forum?id=Z7Lk2cQEG8a

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:02.091759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.068517Z digest=sha256:3a4cfd49a821ba876bc82d1738b6d42d2f867aa9ce9e8d1cdf933810e9e366ce

Observation 064d0a4b-9a0b-4ce9-a5ac-6bb0e04fee23 · outbound

This paper cites On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:00.129590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:13:00.129590Z digest=sha256:b9f0df290ff3eabc9acf8ee1600c79cc6651c10ef3f6dd7a990c3aa0adf48a44

Observation 047842c8-cae0-4b6f-88e8-63aebf6436ef · outbound

This paper cites Woodworth, Suriya Gunasekar, Jason D.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Woodworth, Suriya Gunasekar, Jason D

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.882611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.191389Z digest=sha256:335e2fefaeeb9560ef266112642d54ac437474fd2db3cb84db40e18488efac6a

Observation a1e10686-1355-4d63-8293-7c13b534aa61 · outbound

This paper cites Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7.

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization Small nonlinearities in activation functions create bad local minima in neural networks https://openreview.net/forum?id=rke\_YiRct7

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:01.699375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:13:00.259466Z digest=sha256:fa749b5b3ff75b502a5a3d76fd360c7728fd77eb42824c394c9d8660c737cec5

Pith citing papers

No inbound Pith citation observations are available.