Pith. sign in

Paper Citation Record · LEDGER

Approximating Two-Layer Feedforward Networks for Efficient Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2310.10837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.10837 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:17:41.909750Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T07:35:29.119466Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c26bf0d3-ad0a-47c4-97eb-d68aea22306b · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.846334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:015c8067cc6b393dce65c7acef1a00ef45d0ff46deaead6f129a46f1195c9004

Observation f1c6fb34-e731-41cc-aa3f-ff880a8298ca · inbound

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning cites this paper.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.909750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.909750Z digest=sha256:a44c9d83a92cf71450365f10702765e205ce7790cebecbc18fa5d5d5cf2350c2

Observation 56cc0234-7584-49e9-9f8c-2f2ec2a01af1 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:28.362411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:7b553fd463c5e4607942bd35da572088545050830c060819e0748541428fc08e

Observation 012e7808-59e0-460c-b772-3acd216a5213 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.121225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:bacf2840034fe6c8c81981d73c441b5bca435593655f015fb5514c57482abf76