Pith. sign in

Paper Citation Record · LEDGER

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

As of 23 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.17852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17852 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:49.061779Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T15:25:21.814232Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T15:26:33.682527Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9085359-cf6e-40c8-9f34-c1ce24afa187 · outbound

This paper cites Demystify Mamba in Vision: A Linear Attention Perspective.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Demystify Mamba in Vision: A Linear Attention Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.162107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.162107Z digest=sha256:a138af571d3ee660fe59aa060899fda4b7cbb37acb9391d86cbd958224457ba7

Observation 452e639c-9a53-41c3-9be9-ab3616972d59 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.233553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.233553Z digest=sha256:c4d3019a824473f100cd03487be48e7e994a574b288ac54e8872d59d69fae8be

Observation b5fb0433-8b3e-4338-bd1a-a288115f0777 · outbound

This paper cites FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.516439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.516439Z digest=sha256:bdb8622cafb9ff4bca69626373a6feb6da9fce25de9767091b9439220ee848e4

Observation c591d36a-529e-4088-ba46-6240fde96dd9 · outbound

This paper cites Aditya Rawal and Risto Miikkulainen.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Aditya Rawal and Risto Miikkulainen

Reference 11

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:49.663860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:44:48.591763Z digest=sha256:be8a4ffbebf4bb3e2d5857b3ecd15cdaed6ee7999697775a4438c348ca02498d

Observation 9cc023b2-3ebc-42ce-bb83-c01f2f4541a5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.962796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.962796Z digest=sha256:e23e686d5eb37bb88c44bd7bc4fe2e3336fb9e4aa281273828a4fc53b00d640b

Observation d0da4f7d-486b-4e53-84fc-1612d11b0d46 · outbound

This paper cites Unbiased Online Recurrent Optimization.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Unbiased Online Recurrent Optimization

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.774666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.774666Z digest=sha256:6e3ff687b3374cf367f9d833494337ca17755a964de391a954d0fc83c337da02

Observation d32bcea1-64c7-49e2-9041-ea2875a5bec2 · outbound

This paper cites Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Suraj Srinivas Malladi, Xiang Wei, Josip Djolonga, and Dale Schuurmans

Reference 2005

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:44:50.012543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:44:48.329435Z digest=sha256:71c64ee73041bd0ad54cf9b2546d0c3b782727a6579c1d3219ea36b3ee22a637

Observation e5a19286-959e-4c74-8927-59b1935e1d25 · outbound

This paper cites doi: 10.1145/2908812.2908941.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization doi: 10.1145/2908812.2908941

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.689680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.689680Z digest=sha256:bde5aa104ee94a49cb2a25c73a113a131d93ed0895a15d9638f29ff5f6506119

Observation 4e49812c-ecd8-4cb3-817d-fde0fbffd9fd · outbound

This paper cites Efficient Transformers: A Survey.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Efficient Transformers: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.847581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.847581Z digest=sha256:b0b68e540bd80cf68b6767bc3b578c798e887eaa812ea4a3948bf296ae5d9dca

Observation f3cf41fa-b619-4722-bf73-eaa3d9399cda · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization OPT: Open Pre-trained Transformer Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.061779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.061779Z digest=sha256:994e9019670688d51fcac7a4c292b3a51d881cb779263bc75b2f7a8c2d49a469

Observation b04da9fd-e410-415c-8f3e-82eb4e7ac224 · outbound

This paper cites Koopman-informed recurrent neural networks.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Koopman-informed recurrent neural networks

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.740119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.740119Z digest=sha256:a5833aafbe4e41ae422277a5998d579d245f37a15185261fa41f93fa92b6a60c

Observation ccf275d2-8c12-401d-a93b-52688f3993e9 · outbound

This paper cites Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Simultaneous Computation and Memory Efficient Zeroth-Order Optimizer for Fine-Tuning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:49.001777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:49.001777Z digest=sha256:88dc41a5f90cc7c7b9b65271a100e5297b91cc2402c63beaa65429311b91ea70

Observation 8e3decb9-8c5a-4957-b82d-50682f695791 · outbound

This paper cites Textbooks Are All You Need.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Textbooks Are All You Need

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:48.421248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:48.421248Z digest=sha256:f7d3c8517a5dc7a23b85c808f30670d44a026f2187324b20af9a171538b588f6

Observation ec1e4f59-413a-46e9-8bfc-2cae6704f589 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.613549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.613549Z digest=sha256:c22ccc717fda7902ba9f085772d65338af03bf8a87e77b5dae120d14a16f77c6

Observation d958b103-bc00-4191-8d6c-a7e0973594d7 · outbound

This paper cites Alex Graves, Greg Wayne, and Ivo Danihelka.

Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization Alex Graves, Greg Wayne, and Ivo Danihelka

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.890058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.890058Z digest=sha256:f7b5b2004aa429509b86d3a742a16612bf9f5fd3fe84e87c98965eaa839cee18

Pith citing papers

Observation 4e24ab96-e287-4206-8acc-317294dcf2a4 · inbound

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers cites this paper.

Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.685182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T15:25:21.814232Z digest=sha256:6b7cbd021e3cd6787f2fdc544097b67ecd0d6b4b6d6854a8c574e460a599f4a3