Pith. sign in

Paper Citation Record · LEDGER

LLM generation novelty through the lens of semantic similarity

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2510.27313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.27313 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:22:09.743961Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7057e73c-ff38-4ddb-b10a-56800f576849 · outbound

This paper cites Towards tracing knowledge in language models back to the training data.

LLM generation novelty through the lens of semantic similarity Towards tracing knowledge in language models back to the training data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:01.796716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:01.796716Z digest=sha256:e1945bf2e74138a7dfdcb8f2594778b593d15c60e1eb794a5291e09b45d5147e

Observation 7f038324-75d4-4b73-b3b2-c76b90d0f3f1 · outbound

This paper cites Smollm-blazingly fast and remarkably powerful.

LLM generation novelty through the lens of semantic similarity Smollm-blazingly fast and remarkably powerful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:01.934159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:01.934159Z digest=sha256:cf1360a1b9ab4fd38017f313d59afbef2d9467bc08b456d780f34b89630e4067

Observation 35f6c891-fe8d-4d93-8836-898eaf4272b8 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

LLM generation novelty through the lens of semantic similarity SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.168632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.168632Z digest=sha256:2ffe658b7f5fc9c7c0983024c32239bc24d57cd1dd336646b63ca93395d00587

Observation 63354f86-9774-4da2-b903-bc4f1837d279 · outbound

This paper cites If influence functions are the answer, then what is the question? Advances in Neural Information Processing Systems, 35: 0 17953--17967, 2022.

LLM generation novelty through the lens of semantic similarity If influence functions are the answer, then what is the question? Advances in Neural Information Processing Systems, 35: 0 17953--17967, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.364747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.364747Z digest=sha256:bc795e94cdf61a27369918467b2327460aefbe99c2e7c659cf50aedd7c948c55

Observation e4685868-1c64-4cef-b683-d65c686b5aa9 · outbound

This paper cites Training data attribution via approximate unrolling.

LLM generation novelty through the lens of semantic similarity Training data attribution via approximate unrolling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.464999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.464999Z digest=sha256:1bd99586af26727ccc18d86e2336152c87f64459586b633df6cb7c6a2e3c86f0

Observation 20068405-b99d-4949-9a26-54f06608887b · outbound

This paper cites Influence functions in deep learning are fragile.

LLM generation novelty through the lens of semantic similarity Influence functions in deep learning are fragile

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.584535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.584535Z digest=sha256:da78bc045985aa5b5af6d8b983c688f934beeb7293afbeaa63e45a8533005394

Observation ec2e9b65-39c7-4484-aaa4-b816a7a102d5 · outbound

This paper cites Quantifying memorization across neural language models.

LLM generation novelty through the lens of semantic similarity Quantifying memorization across neural language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.744747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.744747Z digest=sha256:cdb9a8eb40e4feeb3dc038a81190a216aac6f789c96ed163cd9cb458e61d8fcb

Observation a96cdbc7-e339-4bdc-91a4-c458dc9c38b4 · outbound

This paper cites Scalable influence and fact tracing for large language model pretraining.

LLM generation novelty through the lens of semantic similarity Scalable influence and fact tracing for large language model pretraining

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.873525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.873525Z digest=sha256:084b3a52bce4b66fd5c32f9ff8065dd48687514dafedf16d75213f36059c07ae

Observation 040c75cf-a3d6-4bb2-9784-bd68fe2671aa · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

LLM generation novelty through the lens of semantic similarity What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:02.954855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:02.954855Z digest=sha256:5c74d030c25cbbe089405e8caf38e7c6095a378b38b777b0cfcab1b048c86758

Observation e4b58db7-455c-47c1-a886-aa56d6b8c92c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLM generation novelty through the lens of semantic similarity Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.067482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.067482Z digest=sha256:5e0ef1b1184adebaad68c8ff9719ef346d943ff13b0309804c023ef9b9cfb216

Observation 4c49077d-d6f4-4060-8ac8-b786c58b9285 · outbound

This paper cites an unresolved cited work.

LLM generation novelty through the lens of semantic similarity Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.204839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.204839Z digest=sha256:f3763266a02567411900fe66a3b93f93ba35ebc13eb7af3de6a397eea32723ce

Observation d389c756-da16-40ed-b2d5-81761da4db20 · outbound

This paper cites The Faiss library.

LLM generation novelty through the lens of semantic similarity The Faiss library

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.345841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.345841Z digest=sha256:5a235cb3deac8c313c30b790f975180dbba212d53ffbd4511fb83ef922467d25

Observation 76b5ece1-fbf5-479d-9e12-0ed11568203b · outbound

This paper cites Mmteb: Massive multilingual text embedding benchmark.

LLM generation novelty through the lens of semantic similarity Mmteb: Massive multilingual text embedding benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.512579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.512579Z digest=sha256:69173fc069be7cd341c54f46fdb99245d3e17679f29f1b8c2cca38ea24c42bdf

Observation 7f6512ee-ca37-45ff-9cdf-c6c86f179395 · outbound

This paper cites Revisiting the fragility of influence functions.

LLM generation novelty through the lens of semantic similarity Revisiting the fragility of influence functions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.617449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.617449Z digest=sha256:fa5d3e529f5911d36b1d7bb3d53f78f17260b5aadf3224f45d3356f5a2e30616

Observation a5007c6e-2973-4308-94e3-a09c20c1ad8c · outbound

This paper cites What neural networks memorize and why: Discovering the long tail via influence estimation.

LLM generation novelty through the lens of semantic similarity What neural networks memorize and why: Discovering the long tail via influence estimation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.688721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.688721Z digest=sha256:8aa1dd3cc508d04bb2d28fb154cb2a6ecbcbc66c016eea5fba29dbc0b373ca33

Observation 83933c59-d09f-4010-b430-5b7555497179 · outbound

This paper cites The language model evaluation harness, 07 2024.

LLM generation novelty through the lens of semantic similarity The language model evaluation harness, 07 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.757380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.757380Z digest=sha256:740358bec87a5e3ae231a6794efe311277d0f5774258a5dc2712c21847d53d81

Observation 99040118-967d-43d6-accb-6da2b4ee0637 · outbound

This paper cites A closer look at the limitations of instruction tuning.

LLM generation novelty through the lens of semantic similarity A closer look at the limitations of instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:03.854759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:03.854759Z digest=sha256:fbda06cbd514d847959ebdb62934db8134bed8abb789820804d81bbfccc6cb40

Observation 1bf76fea-cc1e-482c-a64d-c3574953aa30 · outbound

This paper cites LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations.

LLM generation novelty through the lens of semantic similarity LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.015766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.015766Z digest=sha256:1d5b85fff55402128ef4d59c122d7ed30f10b972af55574bd2ddb6547e3fcfa3

Observation 461b36d8-226a-4d1b-abb1-df2ce0f1774d · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

LLM generation novelty through the lens of semantic similarity Studying Large Language Model Generalization with Influence Functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.219818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.219818Z digest=sha256:0683812bb553e1a72c3c33819ea016dbeca9fee0aadf27bafa1f5d4eca80509c

Observation 131b9465-304e-4455-8292-cb760feb1186 · outbound

This paper cites Fastif: Scalable influence functions for efficient model interpretation and debugging.

LLM generation novelty through the lens of semantic similarity Fastif: Scalable influence functions for efficient model interpretation and debugging

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.369134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.369134Z digest=sha256:3e3bd09193e9d57328c1e3f262fc62bc93d4f08e523e13cbf0bdce0648efe448

Observation 53b20509-aeb8-4a65-bf96-5cd9742f8c8c · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

LLM generation novelty through the lens of semantic similarity Lighteval: A lightweight framework for llm evaluation, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.512506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.512506Z digest=sha256:7d382967b8419deb9630e8a0bd38a351c9bf45b796c697a2bb93aa3eda509bf8

Observation 6ac03696-b350-4c37-af00-40c497c1bca7 · outbound

This paper cites Training data influence analysis and estimation: a survey.

LLM generation novelty through the lens of semantic similarity Training data influence analysis and estimation: a survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.654817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.654817Z digest=sha256:8c18289e4a21fd4b18f05b671745898c28de3c18a572ff8aec245b187d407e08

Observation 02caca92-0bde-415e-9bce-7b53fcb0b00a · outbound

This paper cites The influence curve and its role in robust estimation.

LLM generation novelty through the lens of semantic similarity The influence curve and its role in robust estimation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:04.860327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:04.860327Z digest=sha256:4155049e36a9d98f320562955b9b63ab929a1e405de299ddc41c1ff635028094

Observation 4eb12f7c-1cc5-4817-934a-1a23760d773a · outbound

This paper cites Data cleansing for models trained with sgd.

LLM generation novelty through the lens of semantic similarity Data cleansing for models trained with sgd

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.019500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.019500Z digest=sha256:498461f5fc07968f49658fb436513b4b33b5d1bd88e061329d8725d0df754302

Observation 2bef80c5-9478-488e-a413-fe6fc6283ae2 · outbound

This paper cites Most influential subset selection: Challenges, promises, and beyond.

LLM generation novelty through the lens of semantic similarity Most influential subset selection: Challenges, promises, and beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.204029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.204029Z digest=sha256:80579ea218d63df3f0f0e2612da88dcbc59ee45deb24f97fca6a16a006b3de65

Observation 589b7f10-d0f3-4193-a0e8-645a5b32c470 · outbound

This paper cites MAGIC: Near-Optimal Data Attribution for Deep Learning.

LLM generation novelty through the lens of semantic similarity MAGIC: Near-Optimal Data Attribution for Deep Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.367970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.367970Z digest=sha256:7041b26414fb2f2a231025e6b7f528d600a1fd3a6ff1bd852ac80e5e99a45ae3

Observation 08c3f432-415e-4edb-860f-d4c4a5b54425 · outbound

This paper cites Understanding black-box predictions via influence functions.

LLM generation novelty through the lens of semantic similarity Understanding black-box predictions via influence functions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.518800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.518800Z digest=sha256:58ca55984802a77770555a401f00fdd7d79782b7b405beea199273b2e99a76eb

Observation 9a495442-1670-410f-8cd5-23663daaff3a · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

LLM generation novelty through the lens of semantic similarity Rouge: A package for automatic evaluation of summaries

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.695810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.695810Z digest=sha256:20d4b6262a8e93242a98eef86837c425ca73240486ac15fb470237c573ae5978

Observation 6391549b-e3f9-478d-9648-b27916ef2a8c · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

LLM generation novelty through the lens of semantic similarity Truthfulqa: Measuring how models mimic human falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:05.851123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:05.851123Z digest=sha256:8aa839ceaadec60c3b8e2337b2930c440c1db81cf3bb3a2a96a2d33d37958024

Observation bf831371-c2d7-4d56-9a17-20d0d8f0e699 · outbound

This paper cites OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens.

LLM generation novelty through the lens of semantic similarity OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.014745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.014745Z digest=sha256:27fe09ff2f0e3e00bb5b1572a9cf005fd50cb23a4e9089d04a7991e360987f07

Observation e0264989-7d06-43ef-9954-c1e87a805a77 · outbound

This paper cites Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens.

LLM generation novelty through the lens of semantic similarity Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.164775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.164775Z digest=sha256:64289a1e8b766b503468b72062e1d3b06e684997d3b7de234f010753eccc31fb

Observation 1745a09c-6c79-4b38-9c5f-01be15b4e9f6 · outbound

This paper cites How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven.

LLM generation novelty through the lens of semantic similarity How much do language models copy from their training data? evaluating linguistic novelty in text generation using raven

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.294755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.294755Z digest=sha256:1e5e79ec96360c68667bf8ad85a828006552f765ff21a1d8b15ff9a69406b1f9

Observation 593920ad-9bbf-4e20-a63a-cb03f5570159 · outbound

This paper cites Evaluating n-gram novelty of language models using rusty-dawg.

LLM generation novelty through the lens of semantic similarity Evaluating n-gram novelty of language models using rusty-dawg

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.464750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.464750Z digest=sha256:158cec4cb2aa662875cf2ea78b43f985c2cde2d6077ef01ba8e6a06e1a403072

Observation 4686c351-8a9e-458a-a114-c5a534eee020 · outbound

This paper cites Waka: Data attribution using k-nearest neighbors and membership privacy principles.

LLM generation novelty through the lens of semantic similarity Waka: Data attribution using k-nearest neighbors and membership privacy principles

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.613245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.613245Z digest=sha256:8754de86d0d124273f87bf2ebd9217ec06d13a5f114a859b572bd361de5f80f0

Observation 4684a11d-3f47-48e9-9de7-5743fb4f933a · outbound

This paper cites Mteb: Massive text embedding benchmark.

LLM generation novelty through the lens of semantic similarity Mteb: Massive text embedding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.804164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.804164Z digest=sha256:0cd9166efabca9c4efbaef26cddf87a66323059810acbf57a766e4870266c1f0

Observation 39e7ade9-47c0-4fe1-bfb9-68c94b734ff5 · outbound

This paper cites A bayesian approach to analysing training data attribution in deep learning.

LLM generation novelty through the lens of semantic similarity A bayesian approach to analysing training data attribution in deep learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.976803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.976803Z digest=sha256:d6fb8f589278ae4999a8b1cc0439eede1d498d358ca3774377c3fcf129479902

Observation bc2e9998-3794-4be8-9746-aacc08d230db · outbound

This paper cites Trak: Attributing model behavior at scale.

LLM generation novelty through the lens of semantic similarity Trak: Attributing model behavior at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.124909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.124909Z digest=sha256:1d2e9ba83f1812a24fecd226cf0902e0f6ecea252ddb619a0cc6fb192a225616

Observation 55c58e19-6020-463c-b76a-e345abcd93f1 · outbound

This paper cites Near-duplicate sequence search at scale for large language model memorization evaluation.

LLM generation novelty through the lens of semantic similarity Near-duplicate sequence search at scale for large language model memorization evaluation

Reference 38

Resolution
verified exact
doi, observed 2026-08-04T07:23:22.122647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T07:22:07.169383Z digest=sha256:2d803d52800abbfd00935f9a8cc37405f87f323645cbbc675310e9f6995789b3

Observation 31c86b42-3a1e-4292-a232-796566de75c2 · outbound

This paper cites Estimating training data influence by tracing gradient descent.

LLM generation novelty through the lens of semantic similarity Estimating training data influence by tracing gradient descent

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.235275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.235275Z digest=sha256:c15b2ccbac369544ba0e23930ac84def7d2a3b263141aac6e24d2b1d2c3c4081

Observation 5e49d205-5c62-4ffd-8030-b4b548b16189 · outbound

This paper cites Scaling up membership inference: When and how attacks succeed on large language models.

LLM generation novelty through the lens of semantic similarity Scaling up membership inference: When and how attacks succeed on large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.322577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.322577Z digest=sha256:597be8f17ed1a23e0e73ab860a0eef5aa26926b967d06de5b22cbffdc90b6304

Observation e76fdc55-e775-4fc9-aa64-aef91271b77d · outbound

This paper cites Learning or self-aligning? rethinking instruction fine-tuning.

LLM generation novelty through the lens of semantic similarity Learning or self-aligning? rethinking instruction fine-tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.408257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.408257Z digest=sha256:c487589e74b2495108f5ec4e2edbc13c640be21cf60d4c2338281d1bdbca9ab4

Observation 25f11077-c65a-43b1-a5e0-7bb54ae5b02d · outbound

This paper cites Colbertv2: Effective and efficient retrieval via lightweight late interaction.

LLM generation novelty through the lens of semantic similarity Colbertv2: Effective and efficient retrieval via lightweight late interaction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.480652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.480652Z digest=sha256:429257e0f9e4702e1386c554f2479a941b4a4dbfb8fb28862f824c0020fe12e0

Observation 9741e13f-1344-41e7-bd04-463cb7a45330 · outbound

This paper cites Scaling up influence functions.

LLM generation novelty through the lens of semantic similarity Scaling up influence functions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.663656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.663656Z digest=sha256:127c282c61e91d83e6124a87633fb18fe272aecad13bd6b099842c7f68cd2e61

Observation a5d6752d-0d45-4374-8dbc-c2103a1a3b04 · outbound

This paper cites Rewritelm: an instruction-tuned large language model for text rewriting.

LLM generation novelty through the lens of semantic similarity Rewritelm: an instruction-tuned large language model for text rewriting

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.721184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.721184Z digest=sha256:f3c82bcecef4fc37116ba000662612da00188ed44c14f34719a5f3f8bdac30aa

Observation eea290b2-8ed0-4287-9509-6ce53d81ee13 · outbound

This paper cites GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning.

LLM generation novelty through the lens of semantic similarity GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:07.934203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:07.934203Z digest=sha256:a97359647c92e73df8861a00137794f03780d4e9766808d4885a1c00f1cc26eb

Observation 59bdb7ce-799f-4fd4-b849-f0b11da4061e · outbound

This paper cites pes2o (pretraining efficiently on s2orc) dataset.

LLM generation novelty through the lens of semantic similarity pes2o (pretraining efficiently on s2orc) dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.063837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.063837Z digest=sha256:0014ac61a078ecd112d9f542db894aff6c3d283a2892bb2ecf20f52b6a2b0f3c

Observation 77038309-dd55-4453-9c4a-f328728db090 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

LLM generation novelty through the lens of semantic similarity Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.190319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.190319Z digest=sha256:ba34d39313ce881035509e06488e192350e025b5b7a7bfea79363de180c3be77

Observation 284b8b2a-8dea-4483-844d-f84cacc9eeab · outbound

This paper cites Enhancing training data attribution with representational optimization, 2025.

LLM generation novelty through the lens of semantic similarity Enhancing training data attribution with representational optimization, 2025

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.307626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.307626Z digest=sha256:c3bb5af7b8c48b639363f233d39284ddf177eeaf4cb27281895158685c76efcc

Observation 374aeb14-73c7-4a0f-82e9-2db62305992b · outbound

This paper cites Better Training Data Attribution via Better Inverse Hessian-Vector Products.

LLM generation novelty through the lens of semantic similarity Better Training Data Attribution via Better Inverse Hessian-Vector Products

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.451533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.451533Z digest=sha256:4e586a06715edf4d1b2cdc2756b8b0a1d9e8a6fa8892fe071c920aff2977fda5

Observation e855cc4b-71e1-4927-bd4b-fba3e83088e8 · outbound

This paper cites Capturing the temporal dependence of training data influence.

LLM generation novelty through the lens of semantic similarity Capturing the temporal dependence of training data influence

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.570579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.570579Z digest=sha256:d95f59a788c9f2d04a60115c2c04779abb857b3994390ff9f4c08c3c596340db

Observation 1cc1e034-90b1-4c14-8af9-df8117f01c3c · outbound

This paper cites Generalization vs memorization: Tracing language models’ capabilities back to pretraining data.

LLM generation novelty through the lens of semantic similarity Generalization vs memorization: Tracing language models’ capabilities back to pretraining data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.765688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.765688Z digest=sha256:4a55024b93112dfcd8bf68b579f907b5359d46ef4e091e72c8137b5fac58a120

Observation e2b8d048-e9ec-45d0-82cd-50c2b27a36e5 · outbound

This paper cites Memhunter: Automated and verifiable memorization detection at dataset-scale in llms, 2025.

LLM generation novelty through the lens of semantic similarity Memhunter: Automated and verifiable memorization detection at dataset-scale in llms, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.905551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.905551Z digest=sha256:e68d3f6eb5942b41a19f823087a7fc69217f164adc07eb6fb3212c72d8bef2a8

Observation 3dd82ca8-7d3a-4e97-ad8a-e71168aea809 · outbound

This paper cites Less: Selecting influential data for targeted instruction tuning.

LLM generation novelty through the lens of semantic similarity Less: Selecting influential data for targeted instruction tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.007497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.007497Z digest=sha256:904e80da4cefdd5e65f24ec0abf68434632922c1256b101d0c72fa543d55a1ea

Observation 555818bc-d506-49e7-b6c2-064d26891fd8 · outbound

This paper cites Representer point selection for explaining deep neural networks.

LLM generation novelty through the lens of semantic similarity Representer point selection for explaining deep neural networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.134693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.134693Z digest=sha256:1da94a936f600498c8ef4b0c97dc37a4f2deeb50666a8eece7d9f78122fcb342

Observation 36584a79-e6c1-4784-b819-5ccbfe15c863 · outbound

This paper cites Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models.

LLM generation novelty through the lens of semantic similarity Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.311136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.311136Z digest=sha256:559f9d1d89ca932a2ea5c6c23af19c06415cd2b6ee6f3dbe7d673948bf844fe0

Observation b4ebc04b-00d3-41ad-b172-acf127c4636d · outbound

This paper cites Pretraining data detection for large language models: A divergence-based calibration method.

LLM generation novelty through the lens of semantic similarity Pretraining data detection for large language models: A divergence-based calibration method

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.391718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.391718Z digest=sha256:5a0d1c3a518033e39c906a4316742ba6007727eca0d413cfb91ff2a1d4ad2854

Observation e4a0e93a-eff6-4032-8e72-7ba3b99d4679 · outbound

This paper cites Dense text retrieval based on pretrained language models: A survey.

LLM generation novelty through the lens of semantic similarity Dense text retrieval based on pretrained language models: A survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.446894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.446894Z digest=sha256:191e4a59d6193791abadec7bceb403a480d316edbfbf9aa7598de87eda607692

Observation 0d7d3ae1-df83-4c7e-9a8f-38dd8c5992b3 · outbound

This paper cites @esa (Ref.

LLM generation novelty through the lens of semantic similarity @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.559811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.559811Z digest=sha256:3f7e3ec73853a0743eae45c6575bb5cb1735e8315000a67784e4a25f2782b0d5

Observation 57f96633-8b27-4546-b6ed-29af55462054 · outbound

This paper cites an unresolved cited work.

LLM generation novelty through the lens of semantic similarity Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.650722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.650722Z digest=sha256:2d3f8b69137fc0d4e7ae77ba282e0e1e4754d699f99e04bb4329a95939830388

Observation 8502a74d-6fba-49ea-a891-f303784f564b · outbound

This paper cites Most training‑data attribution (TDA) methods ask which training examples causally influence a given output, often using leave‑one‑out tests.

LLM generation novelty through the lens of semantic similarity Most training‑data attribution (TDA) methods ask which training examples causally influence a given output, often using leave‑one‑out tests

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:09.743961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:09.743961Z digest=sha256:dd5825b923937ee66aaf9722a566f1495a0eb5b400d9a033840c0564eb921ecd

Pith citing papers

No inbound Pith citation observations are available.