Pith. sign in

Paper Citation Record · LEDGER

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2501.14713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14713 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:48.350745Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:52:28.170298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:52:28.290404Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71a7b657-c5c8-4a52-b94c-c8fce5c2cb8b · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.250451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.250451Z digest=sha256:ca27006c7e7e1b22e7f62cd6892fab33cc2aadc32d837230d5b347826fac81c1

Observation 9d050b94-7530-448f-85c4-29d3b92b4509 · outbound

This paper cites Head-wise Shareable Attention for Large Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Head-wise Shareable Attention for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.254511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.254511Z digest=sha256:92771ae864a9229f6ee42f4a2740226f2b12f0bf34ea2c471ed8a7da33278ed2

Observation b998c2a9-f8a3-46c5-b9b5-4bf5bdbb2583 · outbound

This paper cites Universal Transformers.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Universal Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.266501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.266501Z digest=sha256:ecbf9cf87c5f5e0afacc1762fca12b67fd12c8bae4ad411b84f020f98a516b68

Observation 03967287-d11a-4066-9a80-d4262baee57a · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.270004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.270004Z digest=sha256:2e46a48be0702b6a1e5834a90cc7eedd2aa3943ae73840c94b81e1a5c65d907c

Observation 096671c5-da39-4457-9027-a9c865b69c54 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.284156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.284156Z digest=sha256:8e4379cc0edf24f3ff43a6c616f77be7fe98234b4d59d1e0d377cdac1a620677

Observation 381fde8c-c300-43a5-a67b-859ea6a06d4c · outbound

This paper cites The MiniPile Challenge for Data-Efficient Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The MiniPile Challenge for Data-Efficient Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.287589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.287589Z digest=sha256:9a224bb8bd5d28caf6f7a4741e8774d3130b6795cf0fa4df690bcf4f90198c65

Observation c63f6358-52c3-4134-875a-b3b1b02e4097 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.291038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.291038Z digest=sha256:4a399797df4104e0f1c57083dc619d6003e45f3a39b8e7b0c2a90c8a032912f2

Observation 022155da-5826-47fe-9228-d151cb30d5ff · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.294454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.294454Z digest=sha256:1ea89304d9b8fb20eaf98e87f5b3a325e03136219f2943d5d3e05a80b2c7af2a

Observation b0ff9be9-71d9-456f-a40a-8d06e0fb6c8e · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.297273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.297273Z digest=sha256:864f4e0f40ca53472c9e4059577952be754d3279033bd2505b61c7fec2fbb6b7

Observation 7b52c128-d06d-4c9b-aaed-1b6e3fb89ff4 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.300193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.300193Z digest=sha256:0d5b7db0128a26c553fbfe9150d99a7ed4ad7b687b7e30f6a4adbb2c72d43842

Observation 5c53e008-1a12-4626-91d5-5a0d4f019fff · outbound

This paper cites Pointer Sentinel Mixture Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Pointer Sentinel Mixture Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.303078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.303078Z digest=sha256:bf325288ff4db1eb0e68d2e5f2fb725a88c390b2317ac9f0e423235acc1d3f2d

Observation c14f7341-46bb-480e-975a-ee48ea735a50 · outbound

This paper cites Large Language Models: A Survey.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Large Language Models: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.306014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.306014Z digest=sha256:c98c3f9f6d98ef1150c963e7c245f976fbc2d2c123a05202a7e428a7e747443b

Observation a1bce8de-2d85-4448-ab4b-cd73becccce5 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing A Comprehensive Overview of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.309409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.309409Z digest=sha256:bf602cba8a4247a2d331bd7be077c77400f1e6d1b7ce39de6b6493126f77d685

Observation cbf55e81-76b6-4a48-ac36-3cd3d18a12d4 · outbound

This paper cites EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing EE-Tuning: An Economical yet Scalable Solution for Tuning Early-Exit Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.313273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.313273Z digest=sha256:a8c459caf461cadade3bd988214d8f125563ccaa0345deb899ade58c7014c68e

Observation 9bcbf74b-1ddb-47e3-ad53-9f5b398d549c · outbound

This paper cites Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.316556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.316556Z digest=sha256:592c90feb8778f90b3bb493850f910c790004b3eaee3e12f046e01d9fb2860f3

Observation d04cfec8-1c97-4c4a-bb5d-bd871882b7ac · outbound

This paper cites Accelerating Transformer Inference for Translation via Parallel Decoding.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Accelerating Transformer Inference for Translation via Parallel Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.319981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.319981Z digest=sha256:a7134a47026aeee32b705c09538df08cb6dfa6526ff3839011e61855e0b2d9bd

Observation 38583a41-8ab5-4c1b-ae1e-2bf9da7864cb · outbound

This paper cites Lessons on Parameter Sharing across Layers in Transformers.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Lessons on Parameter Sharing across Layers in Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.323153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.323153Z digest=sha256:775258d8e4b7fee25bf0ade712de35fff90d70d106b5a8779f6326e477692082

Observation c6cb0186-e725-4e98-bdb6-8b6a6d063d1b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.326654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.326654Z digest=sha256:49647815a0da5a59822dea1a2877184ea37298bc996552b49942be9d6dbf82b9

Observation c69ce8be-58d2-4e11-bc9a-c3f7eba90d24 · outbound

This paper cites LaCo: Large Language Model Pruning via Layer Collapse.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing LaCo: Large Language Model Pruning via Layer Collapse

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.329840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.329840Z digest=sha256:e938eb2d8ebfd2febb20ca96f9f084571cfeef8a4d17d46f0156781dc144b66b

Observation 57c418fc-d422-42a0-9d67-ab4c55626e42 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing TinyLlama: An Open-Source Small Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.336936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.336936Z digest=sha256:4afe275614ba183e6bd991131a5db483219ce5f225111a9caa8d51fc97d74ebe

Observation 0cc733c4-c108-4279-958e-4d34f1faaaf5 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.340361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.340361Z digest=sha256:f45be72a0445ac929a02cd37b0442a75f4259974a1beff942673c6514d0db65b

Observation f1921a86-2b40-4166-ac1b-11252f00cfde · outbound

This paper cites For zero-shot performance evaluations, we used the ARC-e, ARC-c (Clark et al., 2018), PIQA (Bisk et al., 2020), WinoGrande (Sakaguchi et al., 2021), and HellaSwag (Zellers et al.,.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing For zero-shot performance evaluations, we used the ARC-e, ARC-c (Clark et al., 2018), PIQA (Bisk et al., 2020), WinoGrande (Sakaguchi et al., 2021), and HellaSwag (Zellers et al.,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:48.676276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T14:59:48.343955Z digest=sha256:fd9f90c09123cbfad13083de36cf26d6181062102fadea10541dacca3ac53b9f

Observation eb18ec19-2fb0-4e74-a975-2ea1477988df · outbound

This paper cites For perplexity perfor- mance evaluations, we used the validation MiniP- ile (Kaddour,.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing For perplexity perfor- mance evaluations, we used the validation MiniP- ile (Kaddour,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:59:48.666836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T14:59:48.347351Z digest=sha256:5b2eddb055ef2f7567e1b915ca242ea8a35bc235afb0a33038de4533bc2df1a8

Observation 697440f5-e2b3-4665-9975-ccd751f4a3d2 · outbound

This paper cites an unresolved cited work.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:59:48.655850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T14:59:48.350745Z digest=sha256:10948b376716437df71485cf4ceb550e0cd4709a4f3e415576b154e364e18434

Observation ac36b88b-c193-4622-8c60-352c07211b0c · outbound

This paper cites Layer Normalization.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Layer Normalization

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.246704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.246704Z digest=sha256:7453e4ae1c6e407d544d47970f084633510d32c128de846fd4ce5285c479a61e

Observation 3be2bc2e-8043-4b15-a740-b14a3c659549 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.262369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.262369Z digest=sha256:7632e7715ff04ccde8f2a3e3b9999a0e635db35823bd29aaecc8798772c2e370

Observation 765a187d-65ab-4d0a-b874-08c7216758a3 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.333160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.333160Z digest=sha256:eeda24f9e3122c0af7fdbfb0689cdc5a7551312c47787405301ede64270dde98

Observation 0be7abd8-2ab4-4ee8-af43-a850c2b04bd8 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.273551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.273551Z digest=sha256:259a92b5413da7283ba7c03f59d5ba5d45b5ab4a8ddf80ab36f709e6089ccb4d

Observation 836ca620-f868-4a1c-aaaf-dc40be6eb857 · outbound

This paper cites Towards a Unified View of Parameter-Efficient Transfer Learning.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Towards a Unified View of Parameter-Efficient Transfer Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.277091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.277091Z digest=sha256:fe78d2c6aeaef4842b275c302a875f2534f514d208558f3085108ee638139168

Observation 6909fb1c-f9a8-43b1-9b96-497726c1e759 · outbound

This paper cites Language model compression with weighted low-rank factorization.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing Language model compression with weighted low-rank factorization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.280607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.280607Z digest=sha256:4b6ebe1883610bfb0256edfeaa8ede5ff84b47a1fd8b0aa6c336b06d3dcbf453

Observation 6765384a-655a-40d2-b6ec-dd16d198f3b6 · outbound

This paper cites EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.258621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.258621Z digest=sha256:c114fe00459ec09b45ef5d674a88c6dd156d910a55626785e7be232148e98a5a

Observation 4c4bb74e-6d21-456f-86f3-308bbb75b210 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T14:59:48.242547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:59:48.242547Z digest=sha256:7360cadd79824b0ffa0bb72f07896dd4955719909c69cec7cbb4106687018741

Pith citing papers

Observation 9364cb06-1227-4de1-ad60-37f0966a8470 · inbound

CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression cites this paper.

CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:52:28.297343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T17:52:28.170298Z digest=sha256:c1fb0b6f83a930e1f626131b1d9202df84ab2cc7801be00d009a98e80ec421bd