Pith. sign in

Paper Citation Record · LEDGER

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.05246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05246 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:17:53.524604Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbe5b08b-a1a0-4570-84a2-5ea7da5335df · outbound

This paper cites Learning to reason for multi-step retrieval of personal context in personalized question answering.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Learning to reason for multi-step retrieval of personal context in personalized question answering

Reference 1

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:54.748244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.272051Z digest=sha256:d043195af876a2608776238ba53f7cbecf346bc858aa1b32083a0bb0f84a911a

Observation 79ad8bfc-f4b9-429c-8877-4842ba5b5e35 · outbound

This paper cites Large language models empowered personalized web agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Large language models empowered personalized web agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.277727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.277727Z digest=sha256:b0a10a32c47158ff320bd74e99b417f5ebe5b94a5abb7323bceb79bf79413533

Observation 6b6bc56a-3184-44fe-a44e-f73234dc6201 · outbound

This paper cites Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.283149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.283149Z digest=sha256:ee6657ef349126684f868c6c268a28377b0f5340cfbff580e8acb97efaaabff3

Observation 05ac2751-17c0-4bf0-90b6-984eefa4c23c · outbound

This paper cites KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.289378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.289378Z digest=sha256:6e288289b68ce1c4cfc7195e2c092e298b3b6b6ecc4fe86a39efa16d9645c9ba

Observation 647d158d-7195-422d-87bc-e353944c7d53 · outbound

This paper cites POPI: Personalizing LLMs via Optimized Natural Language Preference Inference.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs POPI: Personalizing LLMs via Optimized Natural Language Preference Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.294909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.294909Z digest=sha256:f373e8aaf303b9f6428128702f6fe9a669b6c4b83c0fc65123d506ca943eb7ff

Observation 5fa22ae1-bd35-49d5-a584-db57a9fdf609 · outbound

This paper cites Lifebench: A benchmark for long-horizon multi-source memory, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lifebench: A benchmark for long-horizon multi-source memory, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.301349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.301349Z digest=sha256:dc98f2fa78eb40d9a357f51c59000c6cdacbad7bd5adab0be71c07c6b8c889d3

Observation d566f43c-52a3-489b-b7d7-a44398374f6c · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.306961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.306961Z digest=sha256:1305bf5436abb1404fd58f678be8738846753589c328069b98690261371b7d37

Observation 1f13e3a0-abd7-4c66-8078-1ab509693913 · outbound

This paper cites Lifesim: Long-horizon user life simulator for personalized assistant evaluation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lifesim: Long-horizon user life simulator for personalized assistant evaluation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.125871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.312072Z digest=sha256:455d02695fd52e54f05b37987e07f99328522c0bd19b4149708ae8bfc19417af

Observation 015c68cd-1c79-4cd7-87f4-2dc8ef03ff31 · outbound

This paper cites A survey on personalized alignment—the missing piece for large language models in real-world applications.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs A survey on personalized alignment—the missing piece for large language models in real-world applications

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.104016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.316730Z digest=sha256:73de48e5466b6ea805d0c9fe3b5c7b28981cc02eae9263d5407c9d9c23a43f9f

Observation 2f3e4e44-ae26-4184-bf7f-b33dbd09a03a · outbound

This paper cites Towards realistic personalization: Evaluating long-horizon preference following in personalized user-llm interactions.arXiv preprint arXiv:2603.04191, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Towards realistic personalization: Evaluating long-horizon preference following in personalized user-llm interactions.arXiv preprint arXiv:2603.04191, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.321693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.321693Z digest=sha256:de7f5f95e3f9054df845eb5f252f4e3173a95543c97d05fd6a44a0c61ce6e6c9

Observation 5a833f00-6104-4aff-8872-1a955379b6ea · outbound

This paper cites Computing inter-rater reliability and its variance in the presence of high agreement.British Journal of Mathematical and Statistical Psychology, 61(1):29–48, 2008.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Computing inter-rater reliability and its variance in the presence of high agreement.British Journal of Mathematical and Statistical Psychology, 61(1):29–48, 2008

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.327020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.327020Z digest=sha256:27387946cbdda69a1949d7b71b39e667c17f7b703564474613e7d680ce951788

Observation 4e65dcc8-0313-48e6-ba2f-143f54d7cea1 · outbound

This paper cites Rap: Retrieval-augmented personal- ization for multimodal large language models.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Rap: Retrieval-augmented personal- ization for multimodal large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.074174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.331856Z digest=sha256:626c016c08cc9dd6f94ca0b9180d5ca1615ebdbfad5c5695345a048c4e3bcd70

Observation 56da7c87-ee3c-4df9-b752-e35cd45242dc · outbound

This paper cites Asking the Right Questions: Improving Reasoning with Generated Stepping Stones.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Asking the Right Questions: Improving Reasoning with Generated Stepping Stones

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.336566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.336566Z digest=sha256:4d24a89d51fb12fcf24049d694b598677b3bc11555a2ad3838b9e1a6f0d3e6d6

Observation 587f4df2-c4fc-4f95-8a4f-e51722811664 · outbound

This paper cites Op-bench: Benchmarking over-personalization for memory-augmented personalized conversational agents.arXiv preprint arXiv:2601.13722, 2026.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Op-bench: Benchmarking over-personalization for memory-augmented personalized conversational agents.arXiv preprint arXiv:2601.13722, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.342011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.342011Z digest=sha256:a5df55ab9aaa33d4527dfb5a1c1f654f30e65ca78e706951675d08dbb954b54d

Observation 26d5f4ea-1045-4c42-a35a-452f625912d9 · outbound

This paper cites Mem-pal: Towards memory-based personalized dialogue assistants for long-term user-agent interaction.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Mem-pal: Towards memory-based personalized dialogue assistants for long-term user-agent interaction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.056133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.346834Z digest=sha256:31c4de0cd34557a73e9a9d9afaf76c861a968d66f89e516de142b266fb43009a

Observation fc44df1e-9e30-44e3-8dea-6c4ea85bddbd · outbound

This paper cites Taylor, and Dan Roth.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Taylor, and Dan Roth

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.352101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.352101Z digest=sha256:da4ecc05dfe7bc9c58b80880988fff7cf2d457a6a8a9cb8b81570447e6b5fefc

Observation 09fed038-012f-4d9a-a4e7-d0b7da1e84b5 · outbound

This paper cites Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.356534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.356534Z digest=sha256:240af5a1ad77d10b87389aa82eb6e33b794056c0a7ff80306f2ff184699da7d4

Observation 1761f197-9cf5-44b7-9579-ff4770ea77d5 · outbound

This paper cites Humanllm: Towards personalized understanding and simulation of human nature.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Humanllm: Towards personalized understanding and simulation of human nature

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:54.163090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.361281Z digest=sha256:4eedc33b63aa0b4fcec17259ccb09c3e11aa6a2f4ddf8654aa02e45aced0f3df

Observation 3df11aa5-64a0-4e50-9c98-4ab4a69a979f · outbound

This paper cites Retrieval-augmented genera- tion for knowledge-intensive nlp tasks.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Retrieval-augmented genera- tion for knowledge-intensive nlp tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.039109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.366074Z digest=sha256:3d9274839c53f3494d49aa34533d20099e9e301c3e3f414563bf2333eab9d240

Observation 0ec4ee73-f0bd-48d8-b716-81e5dd09304c · outbound

This paper cites Can llm agents simulate multi-turn human behavior? evidence from real online customer behavior data.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Can llm agents simulate multi-turn human behavior? evidence from real online customer behavior data

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.021118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.370929Z digest=sha256:1adc76daa2b29245735acd7c1990efeb5297346743ff66a94f2b99cc92e0fdc1

Observation 1f0c53b9-6e91-4f50-94c1-229004c058e5 · outbound

This paper cites Exploring the potential of LLMs as person- alized assistants: Dataset, evaluation, and analysis.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Exploring the potential of LLMs as person- alized assistants: Dataset, evaluation, and analysis

Reference 21

Resolution
verified exact
doi, observed 2026-08-08T17:17:53.595425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.375836Z digest=sha256:1b7f92a04fdbf8aeda847282d2bdc0bdab7689b27ab89a5a4ba9827a0a3f4ff9

Observation d126780d-3f88-4567-a0a9-0fda6c6b59fc · outbound

This paper cites Privacybench: A conversational benchmark for evaluating privacy in personalized ai, 2025.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Privacybench: A conversational benchmark for evaluating privacy in personalized ai, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.380675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.380675Z digest=sha256:ff99b3d25d9c0d691ea25c06b670bef321145b39884951a154ddc2fd40d61902

Observation ad1f9ba3-504a-4ab0-8a64-429516c6ef1e · outbound

This paper cites PersonaVLM: Long-Term Personalized Multimodal LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaVLM: Long-Term Personalized Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.385274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.385274Z digest=sha256:f4ffa9574360509154ad77c5fb78154ade8308dbdf9e662226579da8df24af43

Observation cb7d46fd-ed79-4425-ae21-9c923116454c · outbound

This paper cites On Memory Construction and Retrieval for Personalized Conversational Agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs On Memory Construction and Retrieval for Personalized Conversational Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.390227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.390227Z digest=sha256:ac2e948a7c859a4e509ab24d3d2a739a5afb5548946169cd23f5d0119ad1aded

Observation a2e0a7b9-8a7c-4286-a73a-a35006f2af53 · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.395160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.395160Z digest=sha256:ffa1b73df909582cf9c88837e68cd69922936e6e19b14e92983f74a3dd2b2175

Observation c4df0f79-a0e7-40ed-aa7c-a6b0fc3768e0 · outbound

This paper cites Lamp-qa: A benchmark for personalized long-form question answering.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lamp-qa: A benchmark for personalized long-form question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:55.003847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.400822Z digest=sha256:0cef6f2ef35e67c415600ab4a9c8c37d002734112997ff8e54276dc9ca7b89bd

Observation b7cb8012-50dc-419e-bf7d-a489e1e3d0e5 · outbound

This paper cites Optimization methods for personalizing large language models through retrieval augmentation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Optimization methods for personalizing large language models through retrieval augmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.405758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.405758Z digest=sha256:3a48270e6493c70e0e506f15b9fd33f407a8666dc1b7fd96372eb47d45037230

Observation 636e4262-22b5-469c-be69-20b8decfc0c2 · outbound

This paper cites Lamp: When large language models meet personalization.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lamp: When large language models meet personalization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.987973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.410357Z digest=sha256:a818eecc57ea2192b8f94f7b00c0e7660f112a69fd96c26e3a14f948b688c23a

Observation 63cdc21d-4945-41e9-b83a-2d50e16f610b · outbound

This paper cites PersonaBench: Evaluating AI models on understanding personal information through accessing (synthetic) private user data.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaBench: Evaluating AI models on understanding personal information through accessing (synthetic) private user data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.414867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.414867Z digest=sha256:8dd583569c7f037e502adad6366c02781499d580bd536e75674ac9bc2ea7bf23

Observation ad19aebd-878a-49c0-a267-9cf724497249 · outbound

This paper cites Democratizing large language models via personalized parameter-efficient fine-tuning.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Democratizing large language models via personalized parameter-efficient fine-tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.961165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.419312Z digest=sha256:6a40aa91687ed634bc4bf7213c2fcc61674b2cd1427dd9dbdda1b0a6f4705550

Observation 5e34d33a-f1a6-4e94-88bc-38449f3a8511 · outbound

This paper cites In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs In prospect and retrospect: Reflective memory management for long-term personalized dialogue agents

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.945799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.424373Z digest=sha256:89525435f82f66578e8ea0b1db4ebc229b4e04c8b6465b798a36a5ce25ba116b

Observation cba440d3-9b69-45e2-a3f3-0aa57e8bb554 · outbound

This paper cites PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.434087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.434087Z digest=sha256:7b0a8bf67b188ba52970168396162c4e65359a117ac82870230b4d8f74da1fd2

Observation d24e8727-09ee-4441-ab20-6213f38d17de · outbound

This paper cites OPeRA: A dataset of observation, persona, rationale, and action for evaluating LLMs on human online shopping behavior simulation.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs OPeRA: A dataset of observation, persona, rationale, and action for evaluating LLMs on human online shopping behavior simulation

Reference 33

Resolution
verified exact
doi, observed 2026-08-08T17:17:54.930558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.438950Z digest=sha256:eb45b8eb030a7c500ed26683c028240e10d9b54719fbf0aabb8cd7728741c1fc

Observation d7cabc12-5fd3-4ad5-9e9c-693f2d03cd75 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.443925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.443925Z digest=sha256:5cb0cafd2e559e2665e98a10ccca854baacca7f63108633d7544cddb77505cd1

Observation 1636cc8f-8755-4d64-a7a7-2d33cc89459b · outbound

This paper cites DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.448806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.448806Z digest=sha256:2553337000e0fdcec7028e9d70bd7785e9277e9df9713222cb3470f213626b06

Observation 294eab45-e44f-4b36-b70a-debc93e08f27 · outbound

This paper cites Lauvrak, Jon Atle Gulla, and Heri Ramampiaro.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Lauvrak, Jon Atle Gulla, and Heri Ramampiaro

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.913126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.453682Z digest=sha256:0cddd35c6a8a940afaabdf426c58d29861e083a86cd5cd1fa25657cb4ee0c9f5

Observation d87058c1-ad0c-498c-956d-3ab710d15acf · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.459192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.459192Z digest=sha256:58d22dc99e4113ec38b6a84af71043c3259456409731b409367a529a11569982

Observation 19ce0ba7-0d2c-4154-8a0b-359a578a26b8 · outbound

This paper cites Promax: Exploring the potential of llm-derived profiles with distribution shaping for recommender systems.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Promax: Exploring the potential of llm-derived profiles with distribution shaping for recommender systems

Reference 38

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T17:17:53.739143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.464402Z digest=sha256:2faad0c944569141394c126815d8d7e88cc13e1e4aff7a62b5281db3a71f3e11

Observation adffac51-bba2-4278-aec2-65d48b3e81bd · outbound

This paper cites Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.469368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.469368Z digest=sha256:fa507700217cea5a0df5b8c55e43c4102ad95f24c65019d67031f078624b8d56

Observation d0a33fd7-19bc-4d37-b060-191a6d177e0e · outbound

This paper cites Cohen, and Emine Yilmaz.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Cohen, and Emine Yilmaz

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.474220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.474220Z digest=sha256:078ea60c8f42b34e6ae63ef8091cc3d8642a334ae1d50aacf6b37400f788ce0c

Observation 3bef6ec4-4c44-4604-9663-e53fa6a84e70 · outbound

This paper cites Cantonese cuisine.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Cantonese cuisine

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.479258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.479258Z digest=sha256:261575c77d17f9868f9dbf78818dab1bb8f0cd203deb70058db4dcdd40dfc824

Observation 46159ad9-2549-44f9-819d-091c4950fd99 · outbound

This paper cites You tend toward budget-conscious choices.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs You tend toward budget-conscious choices

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.896381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.485499Z digest=sha256:a7849fc704a2041d336d53d33762b7bd09b2fa8675eb34cccc1627bd71c20290

Observation cbacb1d9-02bc-4d7d-911f-7b2fd1fbb502 · outbound

This paper cites Note that cross-domain references serving the response are legitimate; offense occurs only when references are purely demonstrative.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Note that cross-domain references serving the response are legitimate; offense occurs only when references are purely demonstrative

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.880274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.490402Z digest=sha256:c6c7258598b718d9a67f709f71d559639c520e9d9b11eefb6c63636d07e807c9

Observation 91dd12f4-3f53-478f-99a7-0b60caec758d · outbound

This paper cites educate" the user, such as.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs educate" the user, such as

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.863931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.495608Z digest=sha256:8c595e0aeb51cf85c7e10cdece42f7e5a14ad4695bcc2b0242381cc1d9024f07

Observation 8d064032-1750-4f6c-9512-85b3a0f62214 · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.848030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.500517Z digest=sha256:440d49d944695c25db87d63330c204cf348922ff176f19916d1170c16ef54956

Observation ad1790ac-5891-4b0f-ad40-1fab75938661 · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.830944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.505050Z digest=sha256:1e1d247432d8e15b85c5a3066c3ce2e030dfcfc2617cbf0f21850f23a3334e87

Observation 177d19fa-f24d-4b90-bd51-6c8034977286 · outbound

This paper cites retrieval.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.815062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.509773Z digest=sha256:e5c48993eaa0fa194a9f5dc0f04fdd7eaed1dce858aa53d79eba1159f743c5b8

Observation c704c8f8-0641-4b77-8a85-e02176aefcae · outbound

This paper cites swap test.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs swap test

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.799387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.515287Z digest=sha256:caa76224378fe1f798c5a73a849a26f32adde975bca2560005529cd7a00089c9

Observation f44d3304-2982-4065-862d-6d738a01164c · outbound

This paper cites an unresolved cited work.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:17:54.781914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.519873Z digest=sha256:5b49c0320309d9e6ea770e79de5e53ffa7ecbb8b152fcc7b5302b9ce7e10810a

Observation 7f5c301d-97ed-471f-a7b4-055c5b9993f3 · outbound

This paper cites retrieval.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:17:54.765834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:17:53.524604Z digest=sha256:e8e97a5ad9a1f5997f5cf4f130c81feafa34b8327cb541467bc315bb377ed230

Observation 0e0385c0-3997-4f02-8d7f-73c7f45ff40a · outbound

This paper cites ISBN 979-8-89176-251-0.

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs ISBN 979-8-89176-251-0

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T17:17:53.429103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:17:53.429103Z digest=sha256:9da8f1fcf0c35e4925d706206f291846c90f2e20ed699c390034d17b3ecbcbc9

Pith citing papers

No inbound Pith citation observations are available.