Pith. sign in

Paper Citation Record · LEDGER

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2506.14460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14460 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:01:49.253498Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:31:09.879734Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:08:43.653689Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57d727e2-4e18-4129-9a3b-a41cf7f67db6 · outbound

This paper cites write newline.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.103621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.103621Z digest=sha256:3e656d874a5b248d1aa6857ec2e0150502e4cb4fe2c87b53b4d1ba759fe5b98d

Observation 78c96441-21d5-4c95-8d03-907a6868648b · outbound

This paper cites Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.110677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.110677Z digest=sha256:20a5efadb61b0703700f0361da84fdbde1ad95ad3c3bde933242fd630ba6b828

Observation 9bdeea19-90e6-4467-8b29-39f4524675bd · outbound

This paper cites On the convergence of prior-guided zeroth-order optimization algorithms.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization On the convergence of prior-guided zeroth-order optimization algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.115459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.115459Z digest=sha256:385f2218cc98e8d1054f65ada20ebed6b600274be3d731924eba4e212ac4424e

Observation f3c73ee0-a4e7-4335-bc5c-a7c71fbd0049 · outbound

This paper cites T., and McMahan, H.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization T., and McMahan, H

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.661735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.120356Z digest=sha256:fd01bbeea9dde714eb3d63b3910e865a5849efc24eb34d22b283ff01dc40477b

Observation 91e4752f-9096-49fd-8916-592c0ae41fb5 · outbound

This paper cites D., Kalai, A.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization D., Kalai, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.647506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.125532Z digest=sha256:224a1367d6cdeb500e9b4e8804078653a2c529e3e9d84cfd96116ea079d89713

Observation c428b578-5273-4838-95bc-b8433d5e4e25 · outbound

This paper cites and Lan, G.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization and Lan, G

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.130309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.130309Z digest=sha256:8c261dd7f11d9f0bf21e79a5592602bac6b48c4340cd5109157830dfd75cd30f

Observation 23051c32-e271-426a-a365-3c669df1f530 · outbound

This paper cites Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.134920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.134920Z digest=sha256:578b39b72612e9558dc4007c3390013373d9a3cab5d9422081b879c065c3d7d2

Observation adf94ff1-c2b8-4ae2-95ad-53c0b146d432 · outbound

This paper cites Optimizing Large-Scale Hyperparameters via Automated Learning Algorithm.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Optimizing Large-Scale Hyperparameters via Automated Learning Algorithm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.139676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.139676Z digest=sha256:4403654e34df6f98f54758bb3889a9573a37ce4d960283ece9b97c54ab2759bb

Observation 5efbdb3e-32e9-4db1-8654-158543e4b8cd · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.614416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.144556Z digest=sha256:6bb74dcd39620311a20e543e49eff82477a7b093a6d163fda77b5c7ef55b5a4f

Observation a91d7ad1-a24d-411e-9a32-e67aa4ccc14d · outbound

This paper cites Zo-adamu optimizer: Adapting perturbation by the momentum and uncertainty in zeroth-order optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Zo-adamu optimizer: Adapting perturbation by the momentum and uncertainty in zeroth-order optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.148863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.148863Z digest=sha256:4cff61d36fe7d75f940be7a5c7574f0d945c73e3c4419485e09a6b250062752d

Observation 7768c882-46e1-414c-93a9-9144506a315b · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.153262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.153262Z digest=sha256:63d01717b081651c864f47580dc00689d4b62daf05a3c3c5916f9589031d0123

Observation aedd7ef9-44b2-4a5f-819f-58e8c45e8895 · outbound

This paper cites Gradient-based learning applied to document recognition.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Gradient-based learning applied to document recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.157669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.157669Z digest=sha256:e0383f52f859215e540f3491154777c6f39223f99015bca7eb52e6df72652187

Observation 0472ceae-f71e-41f3-9a7e-cd3b86dc32cf · outbound

This paper cites A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.161822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.161822Z digest=sha256:d1a1d1778b993aa97078b6a74890d0afb691125d75436aa90a75a4205e17b032

Observation deac08bd-3a84-4784-8c8b-a93ee00214fa · outbound

This paper cites Zeroth-order stochastic variance reduction for nonconvex optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Zeroth-order stochastic variance reduction for nonconvex optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.166029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.166029Z digest=sha256:ece753eb5ac35c456fdf5e4d0939fcf6578fc30053973701772005c0e9ae86a8

Observation 1056cb1b-52fa-40bd-b8b5-babb558ca1c2 · outbound

This paper cites D., and Amini, L.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization D., and Amini, L

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.170334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.170334Z digest=sha256:273a2ca56e63dcee7a8ca1df6d28703801665df28faece19cdf24c49685c6f18

Observation 807eeac3-8367-4b56-9896-599f63173108 · outbound

This paper cites D., Chen, D., and Arora, S.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization D., Chen, D., and Arora, S

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.174552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.174552Z digest=sha256:57c4f074443e5e71310e70200cf40d2226542aacfec4bcee70644c91f34e1144

Observation b782b7a7-0bc8-45ea-84d0-f66774cccf1f · outbound

This paper cites Adaptive First-and Zeroth-order Methods for Weakly Convex Stochastic Optimization Problems.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Adaptive First-and Zeroth-order Methods for Weakly Convex Stochastic Optimization Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.179179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.179179Z digest=sha256:bf0adbabd222dc9362be3cbe267f880fa453e3fa5fd237b72d869a45f1927dfc

Observation 660da4a9-e7c6-4086-9006-3e47a67ca1c7 · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.184001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.184001Z digest=sha256:89ec95a2605d172675391ae208850d62ab521aed0644eba5a8e2c66c8969993e

Observation 35362983-5f3e-4cee-875e-5acfd01aad81 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.188783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.188783Z digest=sha256:e76bbf3d128d1f179db373402840089e17bb5e82e826ee1d09098b35ac6dd84d

Observation 78b1331c-1c10-4d4b-acb5-95f6a36e53c3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.193922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.193922Z digest=sha256:ae7ee95169b49c58b1119861d941a7d20340f53a4eb93b7d2650227058d85b2c

Observation 7b6320ae-4819-4579-9929-5a605675ad8b · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.198336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.198336Z digest=sha256:4a8673460f2228ff63528384bdd92df52328f6fe160da515ea0d06cf7c0f43c0

Observation 383511d1-3dd7-49b8-98e1-353c28106c09 · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.202671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.202671Z digest=sha256:1ab8d4ae50ce24bcfad51f8aab50245b6b2d1503a7d8f2dbf0ac7c8a4f8f2827

Observation 625beb09-e378-4a5a-8fe4-d0dac29d17c6 · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.504641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.207074Z digest=sha256:02972e2dccd8bd114dca43ec2294df5e617b3a17f49ff01d50c6bedcf956bca2

Observation 8e88b0b9-3d8a-443e-9f17-9b11cddd42df · outbound

This paper cites Refining adaptive zeroth-order optimization at ease.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Refining adaptive zeroth-order optimization at ease

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.489376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.211212Z digest=sha256:99c1a26731798862fde4f4fdc460f7f9d58d9e7083abc2ce32dfb5b8f2695570

Observation 6e2eec0a-2f67-437f-9ed8-24cb5f8dc79f · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.475007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.215345Z digest=sha256:e91be73021a8bab07064a42020f7188dc37fa15d891be6742a06291fbc676cc5

Observation f23354b2-faee-4bb9-8801-13d82c64ea10 · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.460226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.219905Z digest=sha256:dc0e75f437c9c209f6e5a1e5184cb159f800546c8656c17048773826ecd448e3

Observation eb2008fe-d874-4fab-b4c8-ef20b93f322a · outbound

This paper cites S., McAllester, D.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization S., McAllester, D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.444048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.224012Z digest=sha256:a637927184070510893a4bf0a01c7e0a76ebd264eb7be550eadd3d3d5ea56cd1

Observation c38af52c-b1ec-4742-8de2-b3e7b4bf6ab3 · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.429372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.228519Z digest=sha256:92649fc9405bbc438c925576bff133cb245c85dd7043ab299b48556e279b593c

Observation cf16f00e-702c-4924-8ded-f7f29b680fbc · outbound

This paper cites Relizo: Sample reusable linear interpolation-based zeroth-order optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Relizo: Sample reusable linear interpolation-based zeroth-order optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.414797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.232722Z digest=sha256:795420bbe443b2c63afd9d7a20ac52d8e328f7df224f9f344d96e65639e57856

Observation 1bc51a29-afba-4a25-bd61-a026ff45eadb · outbound

This paper cites an unresolved cited work.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:01:49.401236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.236845Z digest=sha256:068d89f8c04665b73505e07de09442de9656191f6aa4cf8cb21eaeb065edc1a7

Observation 97591630-0fb9-467c-aa45-8281dc083863 · outbound

This paper cites Unlocking black-box prompt tuning efficiency via zeroth-order optimization.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Unlocking black-box prompt tuning efficiency via zeroth-order optimization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.387452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.241034Z digest=sha256:e6b2f8ab96226e33e43aad5c0cafc2a62250412c0f8c96f4a925c17c352c232e

Observation 005b513d-a0ed-44ea-b25d-87523409fdd1 · outbound

This paper cites V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:01:49.245350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:01:49.245350Z digest=sha256:1b6a2cf58591e28cbbf922eb90bc32d90cf07cc6181cd0788718ea16d35f021f

Observation cb4eac6f-8599-4272-b6b2-dab4fdefc14a · outbound

This paper cites D., Yin, W., Hong, M., Wang, Z., Liu, S., and Chen, T.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization D., Yin, W., Hong, M., Wang, Z., Liu, S., and Chen, T

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.363993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.249420Z digest=sha256:5c35e1768750ce5051fdc2c51ead21a53ba43450ed80bc176b7d1af5c731e72a

Observation c794e71c-7a65-4a33-83d1-bc8deeefc3f5 · outbound

This paper cites Quzo: Quantized zeroth-order fine-tuning for large language models, 2025.

Zeroth-Order Optimization is Secretly Single-Step Policy Optimization Quzo: Quantized zeroth-order fine-tuning for large language models, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:01:49.348712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T20:01:49.253498Z digest=sha256:d4a9aaa974b7f6e169f83b0b282dbf30a9e3d13c7756e6a8ef6b0179599cf2fd

Pith citing papers

Observation b8929d60-21e8-410a-ae31-8af0b5fe07b0 · inbound

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation cites this paper.

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:58.896606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:38:49.399967Z digest=sha256:2114134d1b9a70e0d62625cde60ce78df7a83bb2384e340a73b549d1911da0bd

Observation c82187c8-6ae4-403b-a20b-5814e2bd2ff7 · inbound

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered cites this paper.

Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:23:44.549328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T21:19:55.074853Z digest=sha256:e2f8d1db45cdd110a275a97040357d4b40204c4be7e1be42f055f19778efea8d

Observation f43334c6-d3dd-4385-92e5-79729074c6cc · inbound

Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens cites this paper.

Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.943345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T23:31:09.879734Z digest=sha256:f293254e1d90ca5aaaf576ba3798b28669a1962036ee130810e8c022c77e326c

Observation f531c490-514a-4ff1-a5c0-1023ef06467d · inbound

Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning cites this paper.

Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.655176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T04:33:10.554853Z digest=sha256:9f0954c4029138023f96ed576383c9690290d3f2aacb9327cc1e0959be8ec3ec