Pith. sign in

Paper Citation Record · LEDGER

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

As of 13 August 2026, this Paper Citation Record lists 100 of 218 outbound references and 0 inbound Pith citation observations for arXiv:2608.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10692 v1

Coverage vector

measured 100 of 218 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.257857Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 218 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d2b53-cf36-4553-99ae-37d9bec5e39a · outbound

This paper cites Langley , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Langley , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.624053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.624053Z digest=sha256:5623c3d8a60df3b791a11ff5a6addb8867b22fa1d4cee62fea92b33de0017f9b

Observation 2e475699-9ac0-4f50-ab64-29a8dff91f87 · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AutoGLM: Autonomous Foundation Agents for GUIs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.630665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.630665Z digest=sha256:6c79007fdbf9c43088ad250785482ecc984a1c0cbe158930aaee26ea2234d637

Observation 2cbf283e-58e2-4f43-a012-a5ff84ee17f4 · outbound

This paper cites Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =

Reference 3

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.637393Z digest=sha256:da4f915c084a9f563a5925eedf796f95065ae0bde0d4e299ce6a7a371a167bda

Observation 262e6737-2251-4969-8689-dc1ad8c3da03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.644209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.644209Z digest=sha256:bad73982c786de08f119f200e2c00a27c66967a47e2d6f3d7db68ba907a040e1

Observation e5bd2610-ebc0-4921-beae-035751c68f56 · outbound

This paper cites Gaia2: Benchmarking.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Gaia2: Benchmarking

Reference 5

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.824200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.650845Z digest=sha256:626f62fd1ebd1c2ab4085eb29dfeccfb36db5d255414622df8552f91563b6b5d

Observation 5f2f6079-8e0c-423e-9418-3d4d740cc9cf · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.656643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.656643Z digest=sha256:1f5332e500f064eaec4dcf7db98525f21b00df145281daa307f9861d43d64b8b

Observation 2eed80e8-4fda-4798-b393-f5c232c1374f · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.662759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.662759Z digest=sha256:52a0b9fe3ec4fe54f97dfde98830b898a72df778a4edcace1a138be7d47a27e4

Observation aab3ff45-1e69-4b58-8a2b-06ba5ec7c50b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.671216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.671216Z digest=sha256:289988e634692c501f838c30dc368fd1d53aa9c4382fc5bb17da7c569dcc1e0c

Observation f6fa389a-5855-47e3-91a9-d72036d08e70 · outbound

This paper cites Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.642993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.676550Z digest=sha256:9a9514020d902c7526d5852ab43feb339c07cc4d36cc74bd4ca8821498eeb3a0

Observation e2c12d1e-c277-49ae-9edd-0c9fd3d326f0 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.682339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.682339Z digest=sha256:25eb72d8cbd1fa870f2dd7c94cf9eec778c45956a136f6f57db4a2d188e0f8c7

Observation 1690fc54-620e-43d7-8ebc-5973c210db23 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.687685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.687685Z digest=sha256:af918fabcb291b65f9457b5db7ced1dbf506f45bf2c210e0413a32d6c318542b

Observation e7fbe459-18b3-4e79-85f5-3ccb31a31ff5 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AppAgent: Multimodal Agents as Smartphone Users , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.693441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.693441Z digest=sha256:d939e173f07d9e6087e285b1f8b05332ef100645921a7e0b836dad8a7f7905fe

Observation 76666871-2092-490d-aa4e-e780c00513fe · outbound

This paper cites DeepSeek-V3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.698799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.698799Z digest=sha256:864e6e2403b789b39b11a25e6311a43da3625d28e27df58fad5cf81f82bca62f

Observation d908d904-43c2-43cd-8707-8af95370b703 · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.704359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.704359Z digest=sha256:1720324181287e66e92fd28352095bc49cd9bcc7601c2afdede39b76b81de077

Observation 3a3361d7-5ae0-455f-9895-ca98c1e57cfa · outbound

This paper cites AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.709584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.709584Z digest=sha256:1d0273a8a24874cbf2653b422ca297a3b2c4b3ab4d3440dfa468aa535bc02036

Observation 3dda646a-c5d9-4723-85f3-7de59453ace0 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.715967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.715967Z digest=sha256:a2ab5f05f327bd29f75ed8cb972d9a12431ab32d8da62da9d5f3188fbac27d1c

Observation a46d83d4-d7ba-4869-ae16-463c5eba8be4 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.721936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.721936Z digest=sha256:1bc68733379f5b535dcb811b21b0cb01df270db34a82a94242704d72e52358b0

Observation 8999ef49-cbed-404e-9226-2314b38e0b14 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Twelfth International Conference on Learning Representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.727753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.727753Z digest=sha256:e6c2ccec9672ab1386071aa89a926255758775a250ae46b7656fb4317fed8f21

Observation 376c67ef-44bd-4260-9b91-fcdda83d8c95 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.733603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.733603Z digest=sha256:8a4e83a8ba3e4d82ca85b16893abf317256668d658367f66af331de4bb796a63

Observation ab0b0730-1558-44c2-9e1e-c3ae8030677a · outbound

This paper cites FanOutQA:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FanOutQA:

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.739412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.739412Z digest=sha256:e3e2fde107e69dede2ccd3b64abee846f76161fcfc2e2b3cdcce65f118d3ece2

Observation 98778e32-1a63-4115-b78c-204c22101d1d · outbound

This paper cites Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.746854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.746854Z digest=sha256:2eab7b1da137650f2628ed303d50ad2ef3496da7822d67e793473c8fbd98a923

Observation ef4c0d86-c4b4-4050-bc2c-3b701dc087d9 · outbound

This paper cites Cohen and Ruslan Salakhutdinov and Christopher D.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Cohen and Ruslan Salakhutdinov and Christopher D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.753970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.753970Z digest=sha256:c6969ce9dd6450ba7174cf6af8cb55683018c55de82e153665618b42d6491ebd

Observation 713569f4-462e-41df-8ac5-7ab811aef820 · outbound

This paper cites Generalizing Verifiable Instruction Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Generalizing Verifiable Instruction Following

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.763030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.763030Z digest=sha256:a1eecae982f178fc05d2cc44b23d150bf2b482447faa44837da502e756b38d86

Observation f677e6b5-217f-4c6d-aa0a-463420e36a22 · outbound

This paper cites On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.770692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.770692Z digest=sha256:db952d7705e1dc8740550a2df8f072bbad08ccba95d8474e2a6cccce58db6e08

Observation 17883964-782b-490b-a663-6947cb8de8d1 · outbound

This paper cites Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =

Reference 25

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.778032Z digest=sha256:496f2621edfc50743b08ef95ab522d6e226323e663e9da9cbdcf9e557343babf

Observation 9b496bff-8d9a-478d-8414-54874be83004 · outbound

This paper cites TL-Training:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TL-Training:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.783860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.783860Z digest=sha256:b0ca45887b55eec49792b7db2e1e65dc1229ecbb60469f0bafe003d293071cf7

Observation a707554c-dccf-4d71-8ff1-4964b98cfabc · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.790762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.790762Z digest=sha256:4519fb379f16f743d8852afa1009ab50b644e1fa22b215593adfa357c000f931

Observation aacb158e-0f82-445c-a4aa-d300ef1634b9 · outbound

This paper cites First-Person Fairness in Chatbots.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information First-Person Fairness in Chatbots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.796199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.796199Z digest=sha256:7807a3ac686d597be48375cc5c3061f242475c615ac673a11c113afce70beee2

Observation 504da331-f6d2-4bc9-8721-fc5ce57263aa · outbound

This paper cites Tool learning with large language models: a survey , journal =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Tool learning with large language models: a survey , journal =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.802775Z digest=sha256:b7b8f6a100d0f3180e10c58489fa85f42f5c813f7b75cbad369b3a9ea2563f0c

Observation e8a15dec-dc31-4a63-9077-949d426406d9 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.808303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.808303Z digest=sha256:9d15b4660b0ecedb362b3013d5c20b58e127697d801ab6f33b35cc1707314aed

Observation 60fe019a-b7e9-434c-9d3c-66d7cf60f3aa · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 31

Resolution
parse uncertain
no resolver link, observed 2026-08-12T19:21:58.814046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.814046Z digest=sha256:e962cc0f4930a0ef1abc8c7b82a93ebda525722edd8ca0a44cae7d0b1d2b4ca5

Observation 870b9e98-03eb-4190-9e6f-4636f8289b2e · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.819426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.819426Z digest=sha256:21b88f7aca3aeddc6aafc3f1dbba0b60332e42d214ce36c3b98fb65d539bc01d

Observation 0f664a46-f911-4e3c-8dec-6054188e3f03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.825117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.825117Z digest=sha256:268f541742815fc5c367e1881220c5e8c4a702728b2cfa9aa5dc59fbc456a81c

Observation 28bd6c66-ec43-40bb-a772-66eaeebf9336 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.830362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.830362Z digest=sha256:2e640c454d3828374595295d3805efa9be9b3f2236bb4142cf27da73a624377c

Observation 84bf6d1d-6af2-4e55-bd46-320b798ddfd3 · outbound

This paper cites Newell and P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Newell and P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.836583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.836583Z digest=sha256:eb05b2c3ff1f7fb0253dc77994b7c1ec4b75a875ed56c0d3d471f2898e3ca404

Observation cc89091e-2067-42c9-906d-0a82fbe3de2c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.842032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.842032Z digest=sha256:996207a556bb7cf2aacebe9bcaf86fed07a18601afaa64183dde519c2880d4b7

Observation a5ecb3c6-38d4-400a-9903-416c564c4d3a · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.850170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.850170Z digest=sha256:797cda7579d1734df1774b7b59efc0e0b5c32f24a95bc470579dc496dc44382f

Observation 6e94bcfa-ef29-4123-8fa6-6c6146d3122b · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information and Stoica, Ion and Xing, Eric P

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.855533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.855533Z digest=sha256:f0c4e5f0286fdfc706d2f4c24956c6487fa9bcfa222e29e104e570817fb52dbe

Observation eead0c0a-dcdf-469c-a06c-a3382e82d138 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.862026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.862026Z digest=sha256:6b671f2afa2828027a420d4fc6d2ac67cfd9bf108917467f4fbbaf9f26249595

Observation f503f6ea-2063-4c3e-bdf7-4b9d802ef501 · outbound

This paper cites Scaling Laws for Neural Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Laws for Neural Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.868620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.868620Z digest=sha256:8312c240d058d316907819f4648e056f0bdf0c3394af3b893f0a8ecc809d2f55

Observation 409a4d5f-eb31-4393-b95c-0ab7fa7dbbe1 · outbound

This paper cites 2011 , publisher=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2011 , publisher=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.875174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.875174Z digest=sha256:8ddb7d6ef03788719aceb004c720c864d8fece17011078388d6eecfd58f5ffae

Observation d1bf6ce0-cf16-42e1-801e-9a1d8e4da587 · outbound

This paper cites RestGPT: Connecting Large Language Models with Real-World RESTful APIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RestGPT: Connecting Large Language Models with Real-World RESTful APIs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.880808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.880808Z digest=sha256:5e49d4ef66d5db126d793f42efc1de263c689235612d66396ca9377d493a5de3

Observation 677d6db5-626c-49a2-8c2c-1758ecae4f89 · outbound

This paper cites Improving Language Models by Retrieving from Trillions of Tokens , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Improving Language Models by Retrieving from Trillions of Tokens , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.886530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.886530Z digest=sha256:3b5e34b926b4969efa5d7310a2612032c3491faa7703bfb34a4005932f23a8a6

Observation 22ec1c83-0ad8-411f-8276-1f1df8abbe85 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.891620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.891620Z digest=sha256:021b4b168875bff304ddfceed5e74ad4bdf3cf303b5c11e9d119c6b1a6b86736

Observation 559368ef-7b8c-4ab8-a315-6d5ec915e2f3 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.897223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.897223Z digest=sha256:afd5b0becc13b56a3d4c229f19119506986ea2b3b609459297de34e506ea53c1

Observation 76062a80-5d32-46be-b837-e9e9ab1b69fa · outbound

This paper cites TALM: Tool Augmented Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TALM: Tool Augmented Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.903319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.903319Z digest=sha256:d2b0c013ab510207934d1d695cb104424f66c2e2e70b2d6f3998a185cb8fe22d

Observation 16d20607-24fc-42bd-beca-d3ee04d9c111 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.908726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.908726Z digest=sha256:50db7b9eb1c87f86e731b60774fb4797551790a82879476ab6080bc85358849a

Observation 6416dcc2-8d0e-40de-be70-bd8c56b2e984 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.914291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.914291Z digest=sha256:78ddbfb569f38702d06f34ad6f8248485b971d7f42e27654980729d7e98755de

Observation 0d8f8a41-0d7f-467f-9700-deacd409d182 · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.919706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.919706Z digest=sha256:55cca0e77862f1f7b9d7590df32dca6c571449e31543b242c83cca0755166892

Observation 243aed30-d86c-4f25-98f4-3ca26e64c1db · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.925867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.925867Z digest=sha256:da3385d116de45a6522c2df8720c0c7882f91a5d072b4cc3a70d8c5e0ca70a72

Observation 02c142c0-e3f2-4826-9212-702502787eff · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.932398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.932398Z digest=sha256:1555d3921f1f558ad58c7dc05778780c1249defca037d04ac835a7b0cd4ad94f

Observation 6a088319-384f-4d81-8edc-c2a88aa94bc5 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information WebGPT: Browser-assisted question-answering with human feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.938065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.938065Z digest=sha256:796fba49583fc45cef4ce33916797abc3feb1eb716623b6ca07e25c2eb1eba7f

Observation 92a932cd-81a9-49f3-abcb-5181cffe8eaf · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Instruction-Finetuned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.944883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.944883Z digest=sha256:b69a039d2a01639d6e9f4efbaa63747a3697764db4c96f3b419c6ea884c903b3

Observation d0e7730e-4561-43e5-b1b9-12d4d2d64d5c · outbound

This paper cites Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.950479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.950479Z digest=sha256:7800cabdc719839a7aa6ca8f0ce4da8b6483445c9827d2082e8b51f447aadca3

Observation ffc3ef02-0713-4f13-9505-a6ef09f776e8 · outbound

This paper cites Narasimhan and Yuan Cao , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Narasimhan and Yuan Cao , title =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.956016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.956016Z digest=sha256:fb2fefc159b1c7f7a4cfb425141668297abdf06b1832685c14fe08a301f619bf

Observation b290845d-1959-4e9d-aee2-e1cef7ef86fd · outbound

This paper cites Chi and Quoc V.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Quoc V

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.961736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.961736Z digest=sha256:36e1a402a7161d2b6b89afdb141c6fd9d190c50f4432175069bd9cae42bbb42e

Observation 501493c6-4cda-46d7-af8f-2689efd2876d · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.967514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.967514Z digest=sha256:dd8e3ae8fa6691188ab9d879358ef473d973f6cdcad62da885716320885d13be

Observation 18b28abb-e018-486c-b4ec-84c4705b8b48 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.973424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.973424Z digest=sha256:5ee647ee9756ea8f22d13c4e2ef94d9b17674d45d9eeb1cf3d444d4a4aea0740

Observation 43770c68-76d8-4d26-97d8-0aa098cde640 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.978942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.978942Z digest=sha256:f30c350e0d7274ad505a2c9857584823b9f71ddb41fa1275c6c2c18a2c0c12f1

Observation 67a41b55-0915-4817-b76f-f48acc6c4949 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.984416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.984416Z digest=sha256:a4dd4cc7c8bd96875675b02b80a9193f14dce980a5bf6ff25689e0d09a838ed8

Observation aca258d9-31f2-48d8-8c41-62638fd2de37 · outbound

This paper cites , author=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information , author=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.990406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.990406Z digest=sha256:ce3951123f5e77e7f8150c0e24561582683d1f23adf530e0073d8df975fea2b5

Observation cd50d75b-b662-4e67-a8d2-d9eb00453e05 · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.996026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.996026Z digest=sha256:5735a42add4e0647d969feb6a37be76c113b3a46a26ced78c96ddf242760a186

Observation 444ba3f7-2efe-410e-9ed9-ad98eb8d3466 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.001037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.001037Z digest=sha256:8be03f82423e08c415edea76dc8941a3140d1f8a4221191d0ecd8c1b8843421e

Observation d61a495b-167e-4bd3-9e9c-9f728b833094 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.008901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.008901Z digest=sha256:bc9438d790d4b2cf24fd5a8a90265835f51a85b438f153f54e12374c6b42d782

Observation 06fcce3a-1348-4f1a-8bbc-40eb06386d6c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.014624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.014624Z digest=sha256:abed086fab33e4aae481a87e07e8e46ee672fab865528a1956743a42faf6e7e7

Observation 7d6ee9cc-40a1-4fe8-ba2c-7d3bfa9bf6e4 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.019788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.019788Z digest=sha256:c0cdac36a43debea290e8c22772ca3841af140bd032428a30b663f4f29307de6

Observation 7a84ff0a-1049-410a-ade2-7db6586bc985 · outbound

This paper cites Program Synthesis with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Program Synthesis with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.028908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.028908Z digest=sha256:ab071c7694056dd942405abf486387a70be7196e29f7da387fcda22097459abf

Observation 5e6bd6cc-eee7-45d3-8a77-6595fb417be9 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Measuring Mathematical Problem Solving With the

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.034957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.034957Z digest=sha256:ea77d7a584076b709898078b4203e96d30436e26fb3368689ef618421895acfa

Observation c5f4779d-be23-4e95-a7ed-c12dbb3cf7a2 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.043085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.043085Z digest=sha256:196b8de73df79caab4870ecfc1a509443058a2174b25e3925aa01cae70b80ae9

Observation 64e51fcf-4af9-40e3-abb9-65f4ea18934e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.049354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.049354Z digest=sha256:54a53ab96eb89ad703a418078c42606097de01325b2ae44320140f2dd97f1e47

Observation 3fec37c2-3645-4121-acad-f36ed77b903b · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.055383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.055383Z digest=sha256:d1c8dfe7e8a805370743b5a7ad65b6775f8296ec4e22f13b0c27013b6fd78a2a

Observation 43eecf78-d042-48cf-93b1-2fc5599aec3b · outbound

This paper cites IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.062541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.062541Z digest=sha256:de494e898b25cdc5857dd2498207da3d37be31b571c6650f3e5609500d833412

Observation c381b3c8-f251-46ff-8408-0aeceb4b2af3 · outbound

This paper cites CIF-Bench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CIF-Bench:

Reference 73

Resolution
verified exact
doi, observed 2026-08-12T19:22:01.708390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:59.069885Z digest=sha256:16ce2af4cb22cba7dfcdf9a37064fd5590730661ecea96c46cb0d48a74faa86b

Observation 219506e4-e879-48af-b797-41ff5f411b71 · outbound

This paper cites Manning and Stefano Ermon and Chelsea Finn , editor =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Manning and Stefano Ermon and Chelsea Finn , editor =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.077006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.077006Z digest=sha256:e75833932517c2789862f10a5992d5aeb8ce7f0e404a6700461a1c0d35090795

Observation e1fda741-9ab2-465b-8824-813dcb2ae246 · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.082243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.082243Z digest=sha256:7b40dff8ed356d1f9925d92d2bc0b6f3021e19215cac70e0543e3accf1fe026a

Observation d7131c1c-784a-45c8-924e-ac75120185b5 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.087672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.087672Z digest=sha256:727bf8fc68f5a6f8140a9424ec1825838e9db567d30dcf57cbb44f26acd3de10

Observation 37c7d433-dc3e-4ac7-a205-acaa6e91b16b · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.093215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.093215Z digest=sha256:8a68abf23917c5e55c3d0d942c2b6cd4b6189efa40256819d5f4d585d3f1b331

Observation 89c0deb6-04e4-46e9-a823-fd598e39e6a4 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.098498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.098498Z digest=sha256:5a5406d5bf1630eeb84d26001f9a2dacd238324ee1b400592e1397bc6b2d10fd

Observation bb77a871-757a-4ad2-842f-b800da01771a · outbound

This paper cites SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.104812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.104812Z digest=sha256:0659eb221862b90c64d887333b328979aee3d409d1490cd15fa6cda551a9e109

Observation 7cb0ccac-6e5e-4765-8b12-d6536c089f2e · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.112477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.112477Z digest=sha256:c1966d573084d1d48eaff2870f075e458314fcf0b9d4dd2154185f87b7039dea

Observation b21da20d-4005-4b86-bb06-3eee258a38e9 · outbound

This paper cites FollowBench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowBench:

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.118087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.118087Z digest=sha256:e4ca825f62e3673a8c9cb9b035739a1f0145c58a300faf73da663172255e3918

Observation 8dea1378-933a-4d8f-856b-792e97303e34 · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Needle in the Haystack for Memory Based Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.129521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.129521Z digest=sha256:36c95885124336f0996feeedc60b45bf36668795d681634103f32b675b87ad29

Observation f76cd0a4-0968-40aa-a153-c79214cd5fc1 · outbound

This paper cites Qwen3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.135717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.135717Z digest=sha256:2532b3cf83ef9caad1d28a8bb71a84c4cd4274506b0217ab1ab44538d07bca2d

Observation 259dd41d-7df7-40b7-9726-859a10de69a5 · outbound

This paper cites Findings of the Association for Computational Linguistics,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Findings of the Association for Computational Linguistics,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.143832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.143832Z digest=sha256:894293bcb78e4e43dbe33faa11d0542ae7bcc2989f7a964b1faaac5c9d80a593

Observation dc5eecb0-ae1b-4ca8-b5ad-b73c7525b84c · outbound

This paper cites InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.152621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.152621Z digest=sha256:dd5f273cce3556a26a7b83044807225342f4a08ae387cf91c89bd44fd1cf39dd

Observation e6d96620-9d51-4b89-af16-2dd3598b649a · outbound

This paper cites Long Context vs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Long Context vs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.160015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.160015Z digest=sha256:45e60460c430e66b500e6f264ec6f93d27d294154f6dd5672ba16bdd5cd062a3

Observation c183f780-9786-431c-b428-3cd84bbbf6d9 · outbound

This paper cites MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.166468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.166468Z digest=sha256:1451216cb211d7be501b01c79d41d76bf0f65d9782891fa8aacc1de38e9aed5e

Observation f43e3356-ddf2-45e9-889f-ebe1c070661a · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.172188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.172188Z digest=sha256:ad0546201770b24e31a7fbafaa8c362614fe1a9414ea4ec7a1405a16c3c07ad8

Observation 6edc0562-3a95-46b9-8612-4b61d88ba15b · outbound

This paper cites Hanjie and Runzhe Yang and Karthik R.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Hanjie and Runzhe Yang and Karthik R

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.179579Z digest=sha256:620478bcef60c562503cbe5556acb88566b641c7b48e3464aca9293dddcd0143

Observation adc69345-6af4-49d4-be9c-d20a6219a7af · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-OCR: Contexts Optical Compression

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.189384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.189384Z digest=sha256:af8fed9bc2cbc1c28a10a00a7e453f906583a96ad71db92726132daa16264767

Observation 0a24fd01-5bb2-49dc-9b7f-205d925dea79 · outbound

This paper cites Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.195991Z digest=sha256:c112d9654c1422c84b6220fc709bbf89f26c87a8c9486cc7d05cbc595e9299c0

Observation 2c123124-ae6f-400b-ad14-e6cb13f5a139 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.203762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.203762Z digest=sha256:5013524b670c40eaf869fb48a96df162232e0da59d32a9dd0338b6b6a49781e2

Observation 10870325-622c-473c-b126-934be6591ef0 · outbound

This paper cites Qwen2.5 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen2.5 Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.213050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.213050Z digest=sha256:0ff8a6d736f2e4cc94be99f6dfbf198496d79ed1a0b5b15cc8e27f82a50c9518

Observation b581982a-1e51-4025-b3d5-e08152b243f8 · outbound

This paper cites 2024 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2024 , url=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.219235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.219235Z digest=sha256:4a4a736987225d89e5498f3b8c63ad94cde4034bcd1909f2a8062b1ac24d46c3

Observation 478158d4-3e28-4940-9f30-97e5f055682d · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.225271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.225271Z digest=sha256:6b6bf54807636267cbd74335534acc894a2aff878184206161d4f4c2c57b3547

Observation 24e30cd2-de56-4bd5-8032-d86f246dcfc0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Code Llama: Open Foundation Models for Code

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.232907Z digest=sha256:0105c5b6a6560c68e4b49c1ea196d548e86b8a7df96e2f436514a25c181e19a9

Observation 3a52fc40-6cf4-4664-a5e6-d87f60e13b2b · outbound

This paper cites ToolHop:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolHop:

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.238846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.238846Z digest=sha256:7199914912edb2c47c040570e75f89baa43ad7d0e70a1f83b2a5ddb30674c05d

Observation bd455ddb-5ec9-46e5-8f49-1a32cc7c7b27 · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.245764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.245764Z digest=sha256:5d32e8bb1c84d425c8641834193c5d156a702e394540ca8f5e9b30c052c66891

Observation 44819e06-aebc-406c-9097-bb6d2c07442b · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.252278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.252278Z digest=sha256:75d3945dadaaa431e002ea5af801f6d3ac1a44caebc91d05fff9e72545483dad

Observation 4081df13-c0af-4874-95c5-8689dd87b6c1 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.257857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.257857Z digest=sha256:e71da2f456b5240c6205528d91b27dbf2ed9a38abc046d777e67c6927a006b4d

Pith citing papers

No inbound Pith citation observations are available.