Pith. sign in

Paper Citation Record · LEDGER

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

As of 13 August 2026, this Paper Citation Record lists 100 of 218 outbound references and 0 inbound Pith citation observations for arXiv:2608.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10692 v1

Coverage vector

measured 100 of 218 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.257857Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 218 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d2b53-cf36-4553-99ae-37d9bec5e39a · outbound

This paper cites Langley , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Langley , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.624053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.624053Z digest=sha256:cb29f489dec6b40ea3b84fee8cbe0afdc12f1a2e79983c7751d4720d129ec7e6

Observation 2e475699-9ac0-4f50-ab64-29a8dff91f87 · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AutoGLM: Autonomous Foundation Agents for GUIs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.630665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.630665Z digest=sha256:b24035ad250b516f1bff61b0c2d9b4d1d8490903adc777e7beb8b53e2488b639

Observation 2cbf283e-58e2-4f43-a012-a5ff84ee17f4 · outbound

This paper cites Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =

Reference 3

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.637393Z digest=sha256:e12fd6b856da345ece21ed373d1b25f8c8696b7fed8c33b22ce09f8620a5c6e6

Observation 262e6737-2251-4969-8689-dc1ad8c3da03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.644209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.644209Z digest=sha256:b4421dd154f75c831149e2d1a21a2d77dd26930463926fc704613fb9df6d131b

Observation e5bd2610-ebc0-4921-beae-035751c68f56 · outbound

This paper cites Gaia2: Benchmarking.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Gaia2: Benchmarking

Reference 5

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.824200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.650845Z digest=sha256:10ae206487c6b17533fd310d3d003cae42132dcfa19a2805da8556765aaf7ff2

Observation 5f2f6079-8e0c-423e-9418-3d4d740cc9cf · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.656643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.656643Z digest=sha256:261f26d54b02f5fa923523efbacb1386b587602919d9ce1b2082f7cbf5cced69

Observation 2eed80e8-4fda-4798-b393-f5c232c1374f · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.662759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.662759Z digest=sha256:1a626531c50923539293a69608acb88c595587a25aada795b36db45f064247db

Observation aab3ff45-1e69-4b58-8a2b-06ba5ec7c50b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.671216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.671216Z digest=sha256:ee8dcbd495d5f191d9cf4b383085d7b0fb92cb796759a473077b2b6f63a0b05e

Observation f6fa389a-5855-47e3-91a9-d72036d08e70 · outbound

This paper cites Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.642993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.676550Z digest=sha256:566951c678e97819c6aea964a026a8107ebb7307ec17f39d0c9af57d5627f49e

Observation e2c12d1e-c277-49ae-9edd-0c9fd3d326f0 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.682339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.682339Z digest=sha256:e2a4db12c4b3839fd16a97da5e622a13037c1c1f7c4c3c5966cf87ca3f078a89

Observation 1690fc54-620e-43d7-8ebc-5973c210db23 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.687685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.687685Z digest=sha256:c46a317a1fff70cee7b0d658fa5b02c2bc1537592f5e7dc335c0ffd2329423e9

Observation e7fbe459-18b3-4e79-85f5-3ccb31a31ff5 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AppAgent: Multimodal Agents as Smartphone Users , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.693441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.693441Z digest=sha256:93dbb36b5b798311d8eb774721045bc3dd21df3911666160694a938c24fab705

Observation 76666871-2092-490d-aa4e-e780c00513fe · outbound

This paper cites DeepSeek-V3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.698799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.698799Z digest=sha256:c7120467532be8345b840452de0f5e6f746ea25c8f62f5303d185897f63a677b

Observation d908d904-43c2-43cd-8707-8af95370b703 · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.704359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.704359Z digest=sha256:0c5e8e2eb6eed079165eef3127290df1e12fce4a11302b0e76a5f3bbe51fc970

Observation 3a3361d7-5ae0-455f-9895-ca98c1e57cfa · outbound

This paper cites AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.709584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.709584Z digest=sha256:5a38af7e520788d1cd13b12c3708a72c6d5afcda13a7ee5e0bcd20017476f023

Observation 3dda646a-c5d9-4723-85f3-7de59453ace0 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.715967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.715967Z digest=sha256:d0e5ef615e620487e4834c1bb99d8cf77612aa858b24ff92eaece0ec59412495

Observation a46d83d4-d7ba-4869-ae16-463c5eba8be4 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.721936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.721936Z digest=sha256:c90ee5ee27014009782b7c40482ba4fd521965a53b8c6f748831e72f624f176a

Observation 8999ef49-cbed-404e-9226-2314b38e0b14 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Twelfth International Conference on Learning Representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.727753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.727753Z digest=sha256:48dacddb8f286ee8c02153453424be053ab37544966a7d9a1be903497f75b4cc

Observation 376c67ef-44bd-4260-9b91-fcdda83d8c95 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.733603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.733603Z digest=sha256:99663b85901820a0d13dd92ea943b41d70ec9455611ccfb7e1dd56e688bf36d4

Observation ab0b0730-1558-44c2-9e1e-c3ae8030677a · outbound

This paper cites FanOutQA:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FanOutQA:

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.739412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.739412Z digest=sha256:46c95d4535d63650ad56f134cdc51b2c465d8502f29d2d04e0c29788fed10955

Observation 98778e32-1a63-4115-b78c-204c22101d1d · outbound

This paper cites Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.746854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.746854Z digest=sha256:d5a5a953bdde46602181d9769bc0512054699ec1b4ff4013bf394581e372959e

Observation ef4c0d86-c4b4-4050-bc2c-3b701dc087d9 · outbound

This paper cites Cohen and Ruslan Salakhutdinov and Christopher D.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Cohen and Ruslan Salakhutdinov and Christopher D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.753970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.753970Z digest=sha256:8e8456626eec7c8a54f3d60e5b796625be517046a94a086c7b922d1fbd64cdd2

Observation 713569f4-462e-41df-8ac5-7ab811aef820 · outbound

This paper cites Generalizing Verifiable Instruction Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Generalizing Verifiable Instruction Following

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.763030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.763030Z digest=sha256:e75b5cead60940ee6c973f4ad02a0d7d0c034ceaa708221115b737ba108db158

Observation f677e6b5-217f-4c6d-aa0a-463420e36a22 · outbound

This paper cites On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.770692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.770692Z digest=sha256:46def7857f59b11386202786333c0b8c072310e89ec3b4673dd30080f58b791e

Observation 17883964-782b-490b-a663-6947cb8de8d1 · outbound

This paper cites Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =

Reference 25

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.778032Z digest=sha256:063582b1b1f39804f35c09cdc9c105b7ef3130cc3c3ea70ef9de8934771f01c1

Observation 9b496bff-8d9a-478d-8414-54874be83004 · outbound

This paper cites TL-Training:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TL-Training:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.783860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.783860Z digest=sha256:da9294d99382afe9c523b203af3fe128818818827a4170603bd545d7969c1081

Observation a707554c-dccf-4d71-8ff1-4964b98cfabc · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.790762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.790762Z digest=sha256:409e03d6940388db13eb20e012c7a850aa663685368aa9b823a5c6acb07ca7d6

Observation aacb158e-0f82-445c-a4aa-d300ef1634b9 · outbound

This paper cites First-Person Fairness in Chatbots.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information First-Person Fairness in Chatbots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.796199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.796199Z digest=sha256:3b26e13ac3446fc9bc0b96d61dc530421bedb70e507a6fb23b598f0ad199f17a

Observation 504da331-f6d2-4bc9-8721-fc5ce57263aa · outbound

This paper cites Tool learning with large language models: a survey , journal =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Tool learning with large language models: a survey , journal =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.802775Z digest=sha256:c70fb8b42c50265d75093e3ab27c07f40b7d427dbe538665e98dd6cd3548bf55

Observation e8a15dec-dc31-4a63-9077-949d426406d9 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.808303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.808303Z digest=sha256:b2ac14aa18716509a60f3085f0786952257cefc4e7c0bb513301b4e970c2c3ed

Observation 60fe019a-b7e9-434c-9d3c-66d7cf60f3aa · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 31

Resolution
parse uncertain
no resolver link, observed 2026-08-12T19:21:58.814046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.814046Z digest=sha256:949d5bfa9b9f383cf57347c3e48335bd410cb694760f335972d12189ccf5e74b

Observation 870b9e98-03eb-4190-9e6f-4636f8289b2e · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.819426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.819426Z digest=sha256:6c8461f9e22fe20b08316db2c675f30f399ec3567611525ea5b145b9420e3a4a

Observation 0f664a46-f911-4e3c-8dec-6054188e3f03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.825117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.825117Z digest=sha256:f64a1f3a64ef0758b508036639389ec4244d1c3ad928d2d798cd47b78f84ee8a

Observation 28bd6c66-ec43-40bb-a772-66eaeebf9336 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.830362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.830362Z digest=sha256:502a519635a348599e7de545d05aad68dd83a4e4592a2d0003c360c2bb253ce4

Observation 84bf6d1d-6af2-4e55-bd46-320b798ddfd3 · outbound

This paper cites Newell and P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Newell and P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.836583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.836583Z digest=sha256:3d165bb76a678a822fbe0bec5b759021e117c28c6bb082191869e50b0ed1479b

Observation cc89091e-2067-42c9-906d-0a82fbe3de2c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.842032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.842032Z digest=sha256:8ed5c906a12858bbdd165c5dea9eb0995ceb2fb44215aeece65404697af3e6d2

Observation a5ecb3c6-38d4-400a-9903-416c564c4d3a · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.850170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.850170Z digest=sha256:bc9f8d273ba460572d0488edb1ab94f8ff51c2b360ea0d4a6aa6a45912d3945c

Observation 6e94bcfa-ef29-4123-8fa6-6c6146d3122b · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information and Stoica, Ion and Xing, Eric P

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.855533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.855533Z digest=sha256:5da99760e76ec760207e739fa45179bc02d925c8cc06b885c759f5d371f9dd4e

Observation eead0c0a-dcdf-469c-a06c-a3382e82d138 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.862026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.862026Z digest=sha256:25f8b1de050fac89466907e1cc6f33f52fb00d0a65d5e20b146984550e090231

Observation f503f6ea-2063-4c3e-bdf7-4b9d802ef501 · outbound

This paper cites Scaling Laws for Neural Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Laws for Neural Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.868620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.868620Z digest=sha256:db699363b92ae692fa66e7fb386363375561d02f5a5777551a5b45adaef28638

Observation 409a4d5f-eb31-4393-b95c-0ab7fa7dbbe1 · outbound

This paper cites 2011 , publisher=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2011 , publisher=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.875174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.875174Z digest=sha256:833042b28cffa017716c1b416e8133aab96a661e8b8a27c7f7c98bbab3e9c1a2

Observation d1bf6ce0-cf16-42e1-801e-9a1d8e4da587 · outbound

This paper cites RestGPT: Connecting Large Language Models with Real-World RESTful APIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RestGPT: Connecting Large Language Models with Real-World RESTful APIs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.880808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.880808Z digest=sha256:e48ad117c3bb0d2d307babc7ab522184f3bdbc662d800333ba14c4d56813c1c6

Observation 677d6db5-626c-49a2-8c2c-1758ecae4f89 · outbound

This paper cites Improving Language Models by Retrieving from Trillions of Tokens , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Improving Language Models by Retrieving from Trillions of Tokens , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.886530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.886530Z digest=sha256:8ba3f21b6a3f986791b83728bd007ee62d3f3a3cd99392f751bf090be54a75bf

Observation 22ec1c83-0ad8-411f-8276-1f1df8abbe85 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.891620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.891620Z digest=sha256:232d04f2d13bf3416e83ec6c1698e634b09b9eb5a21b3ea9505f5b8b4829fc39

Observation 559368ef-7b8c-4ab8-a315-6d5ec915e2f3 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.897223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.897223Z digest=sha256:e35e158be35ce782a5381aa9ca45fad1c0a42b67d6b3202973116c32a15b0084

Observation 76062a80-5d32-46be-b837-e9e9ab1b69fa · outbound

This paper cites TALM: Tool Augmented Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TALM: Tool Augmented Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.903319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.903319Z digest=sha256:697e8e2949bce4b7718090cfad1770d0d2c9f9d4b59d42f2c5e1af0984dc5d7f

Observation 16d20607-24fc-42bd-beca-d3ee04d9c111 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.908726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.908726Z digest=sha256:45e566084b7b4f9affafa5dd6abcbc3d039aa83d5e2fcbddfb8db76fb31f6b76

Observation 6416dcc2-8d0e-40de-be70-bd8c56b2e984 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.914291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.914291Z digest=sha256:84ed58d7c6b54ba6f51b770a9a96698e10bf245a2d0e0c48b1c8f043c9884cac

Observation 0d8f8a41-0d7f-467f-9700-deacd409d182 · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.919706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.919706Z digest=sha256:31de0a338b423865f890239a76771b9ecc90823ece79bd99106bc8a0ea05ccd9

Observation 243aed30-d86c-4f25-98f4-3ca26e64c1db · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.925867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.925867Z digest=sha256:44a9af5b668da8ee324798e474481063385b938669ac3c560c7b9c72db5b112d

Observation 02c142c0-e3f2-4826-9212-702502787eff · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.932398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.932398Z digest=sha256:b030e77b7194c1c6031625db8e0d089e08d942fba1da209efd8bcfd9e725417c

Observation 6a088319-384f-4d81-8edc-c2a88aa94bc5 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information WebGPT: Browser-assisted question-answering with human feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.938065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.938065Z digest=sha256:86723dbd796e2c4ba84ecdc8ac94bf534e937b45c6009d98779c69cd60a84eb9

Observation 92a932cd-81a9-49f3-abcb-5181cffe8eaf · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Instruction-Finetuned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.944883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.944883Z digest=sha256:42ba4d982cb519ce5d8feaa88876dadce6416ddc521ac2da328562bb63bf6d57

Observation d0e7730e-4561-43e5-b1b9-12d4d2d64d5c · outbound

This paper cites Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.950479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.950479Z digest=sha256:1f7e2611fcf121ed9140ea2f7bcc08187a625ce3674d959a73f251249e41ab1a

Observation ffc3ef02-0713-4f13-9505-a6ef09f776e8 · outbound

This paper cites Narasimhan and Yuan Cao , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Narasimhan and Yuan Cao , title =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.956016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.956016Z digest=sha256:030a8a729ecca959d3d58ac713e330d32d9021ab69f997bd10d62d65c344d97f

Observation b290845d-1959-4e9d-aee2-e1cef7ef86fd · outbound

This paper cites Chi and Quoc V.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Quoc V

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.961736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.961736Z digest=sha256:420517c5547a84de14b20dc02bca4d70e01a57b2b53450989829c0cbf5181879

Observation 501493c6-4cda-46d7-af8f-2689efd2876d · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.967514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.967514Z digest=sha256:8d4f31bc0fe04259abaddd1f4ccab70e7f778055f504758b0ba295bea978b1f8

Observation 18b28abb-e018-486c-b4ec-84c4705b8b48 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.973424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.973424Z digest=sha256:541eb2ec497d72bb7dd1e05bf43c6d9a02971ed97b6925674acfbbcc0bd4d21c

Observation 43770c68-76d8-4d26-97d8-0aa098cde640 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.978942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.978942Z digest=sha256:494e63ce43499caa14eeaf0d20e28ba46343620c196324829276662ad62312bb

Observation 67a41b55-0915-4817-b76f-f48acc6c4949 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.984416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.984416Z digest=sha256:837646b13543cb9176bff69d15e13acca2243cefc4c7b2ee51d6555ccc7bc73b

Observation aca258d9-31f2-48d8-8c41-62638fd2de37 · outbound

This paper cites , author=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information , author=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.990406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.990406Z digest=sha256:59a0563c2ec6c37c1df2f6680028804138567d8eaa3a05275a67ff4e40205759

Observation cd50d75b-b662-4e67-a8d2-d9eb00453e05 · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.996026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.996026Z digest=sha256:64865624fb0a182f16d06060f9f3b339fe949784d63018b348335fc836978459

Observation 444ba3f7-2efe-410e-9ed9-ad98eb8d3466 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.001037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.001037Z digest=sha256:b5cd67f59e081176d06a4a40db53161851ff4b74ccd61d049007d11d9205762b

Observation d61a495b-167e-4bd3-9e9c-9f728b833094 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.008901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.008901Z digest=sha256:c2dce5af3de764bcd9e5d102abd70c4a940a30f0990b9f06200c8ac7122ca424

Observation 06fcce3a-1348-4f1a-8bbc-40eb06386d6c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.014624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.014624Z digest=sha256:60c8263cfed019d92010da31f983729aee45d02cf4de13ba4bfbb9165308f3af

Observation 7d6ee9cc-40a1-4fe8-ba2c-7d3bfa9bf6e4 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.019788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.019788Z digest=sha256:259645b8109fcc552a045d4febe15dbc28961ce53e714b07a0035cdfb314576d

Observation 7a84ff0a-1049-410a-ade2-7db6586bc985 · outbound

This paper cites Program Synthesis with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Program Synthesis with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.028908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.028908Z digest=sha256:0505f7d767c709d2b7e6f2828061d8c80079395f8217d64f63292b4879951b7f

Observation 5e6bd6cc-eee7-45d3-8a77-6595fb417be9 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Measuring Mathematical Problem Solving With the

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.034957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.034957Z digest=sha256:8df082273925635557789c0200b48819bf4d78c13266e9ca32cc5e57ae803aec

Observation c5f4779d-be23-4e95-a7ed-c12dbb3cf7a2 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.043085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.043085Z digest=sha256:7641208df3042c8b399906c0ad8854075b606d78958b8673e26208b88e327a52

Observation 64e51fcf-4af9-40e3-abb9-65f4ea18934e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.049354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.049354Z digest=sha256:ebe1b996fb56e8f36758955c8675c4556db0edbf391a22a4ba9a4bd2fc03c096

Observation 3fec37c2-3645-4121-acad-f36ed77b903b · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.055383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.055383Z digest=sha256:37a04692fb5904ae5444effe290134e28f8319005768d41b83f19b26afd28459

Observation 43eecf78-d042-48cf-93b1-2fc5599aec3b · outbound

This paper cites IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.062541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.062541Z digest=sha256:35c064590fdccabfb67a275c4c8358ac373438b3a7fa23c6460be79362c4f65c

Observation c381b3c8-f251-46ff-8408-0aeceb4b2af3 · outbound

This paper cites CIF-Bench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CIF-Bench:

Reference 73

Resolution
verified exact
doi, observed 2026-08-12T19:22:01.708390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:59.069885Z digest=sha256:6ffcf66f5f052ec35c9e990c7646bd50bb87ab940188fc02f1071cdfdd794b8a

Observation 219506e4-e879-48af-b797-41ff5f411b71 · outbound

This paper cites Manning and Stefano Ermon and Chelsea Finn , editor =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Manning and Stefano Ermon and Chelsea Finn , editor =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.077006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.077006Z digest=sha256:f2cf91529c95ed2ac6fc85ee09764b7ee47697de7dab55676864a37b4efdcae9

Observation e1fda741-9ab2-465b-8824-813dcb2ae246 · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.082243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.082243Z digest=sha256:8eee6189129cf1a6f78d2e7bde48656688f31c9aadd237379b9353c2c6b30452

Observation d7131c1c-784a-45c8-924e-ac75120185b5 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.087672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.087672Z digest=sha256:758f6922ef5e4eae4091c6a159709584c5fd83165df7453ebe0f660cbcef490f

Observation 37c7d433-dc3e-4ac7-a205-acaa6e91b16b · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.093215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.093215Z digest=sha256:607e54b190c81a3251933525e9afb4201f635c96bac94f6b4755e08aa9a3fffb

Observation 89c0deb6-04e4-46e9-a823-fd598e39e6a4 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.098498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.098498Z digest=sha256:db66493a2bd80a3b16d44161e7df16627761b054b65a920c259f9cf8108830f9

Observation bb77a871-757a-4ad2-842f-b800da01771a · outbound

This paper cites SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.104812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.104812Z digest=sha256:7a9d707d604058b74ca134817fbbc299411e983d1c19de99818f91b442730d63

Observation 7cb0ccac-6e5e-4765-8b12-d6536c089f2e · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.112477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.112477Z digest=sha256:be19421fa40ea139f0861eeaead716636e1c97c47b9a45c1b1b86d53f04c523b

Observation b21da20d-4005-4b86-bb06-3eee258a38e9 · outbound

This paper cites FollowBench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowBench:

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.118087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.118087Z digest=sha256:d9eafcfbf03eea2930e8f0747008c74756845bf2def9ecb247f5cc12734c4eed

Observation 8dea1378-933a-4d8f-856b-792e97303e34 · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Needle in the Haystack for Memory Based Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.129521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.129521Z digest=sha256:a24450a1f55075b95030b9e4e1c3782b6abfd7633194fb9f818c4b3a6fa0f74d

Observation f76cd0a4-0968-40aa-a153-c79214cd5fc1 · outbound

This paper cites Qwen3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.135717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.135717Z digest=sha256:a6b897fa2eecc48efd8de3e3f8364a17fc70a7d42ceda147d018414ecbc19cf3

Observation 259dd41d-7df7-40b7-9726-859a10de69a5 · outbound

This paper cites Findings of the Association for Computational Linguistics,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Findings of the Association for Computational Linguistics,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.143832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.143832Z digest=sha256:1cbb97a1e74a7420cd9f9cbfa0131e20de9b42c8c2acdff2e5feae9387a4a4c0

Observation dc5eecb0-ae1b-4ca8-b5ad-b73c7525b84c · outbound

This paper cites InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.152621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.152621Z digest=sha256:1e458d72a2a89cf6636c06af3cd8d6839124647043135a61771e061e18399ae6

Observation e6d96620-9d51-4b89-af16-2dd3598b649a · outbound

This paper cites Long Context vs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Long Context vs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.160015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.160015Z digest=sha256:728f8b778b4849233d994fe2cfa46f7741c1da80bdb0ec2f02e4212fa21f2ab8

Observation c183f780-9786-431c-b428-3cd84bbbf6d9 · outbound

This paper cites MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.166468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.166468Z digest=sha256:15338cfa15f3c793e6e141be3ffb3dfb331134096f38f9070335382f0859d102

Observation f43e3356-ddf2-45e9-889f-ebe1c070661a · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.172188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.172188Z digest=sha256:897132716233df4e62252e6c775425d2f6199f670cd6a3942bcbb5594092d976

Observation 6edc0562-3a95-46b9-8612-4b61d88ba15b · outbound

This paper cites Hanjie and Runzhe Yang and Karthik R.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Hanjie and Runzhe Yang and Karthik R

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.179579Z digest=sha256:2bac2046964feb303e6125d04b5253a4e731ed538f40c905aac3a17989728ac3

Observation adc69345-6af4-49d4-be9c-d20a6219a7af · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-OCR: Contexts Optical Compression

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.189384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.189384Z digest=sha256:e618a222bafc8e2ae846449c1f6adb0a280950e3f73fba8a0468093650244a02

Observation 0a24fd01-5bb2-49dc-9b7f-205d925dea79 · outbound

This paper cites Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.195991Z digest=sha256:307bce04e79bedbce28d3b36bc249fe17ab86ed66571895f1fe586eff95a8d41

Observation 2c123124-ae6f-400b-ad14-e6cb13f5a139 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.203762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.203762Z digest=sha256:3fd18e503a1c3c2f8f684040d2905c7289a6b48bcb3fe7f65047e9cfebf6f742

Observation 10870325-622c-473c-b126-934be6591ef0 · outbound

This paper cites Qwen2.5 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen2.5 Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.213050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.213050Z digest=sha256:10d415f3985863bc4562e01c0984e6828fc1b8110b3c3ed498cfc84c57f24e4e

Observation b581982a-1e51-4025-b3d5-e08152b243f8 · outbound

This paper cites 2024 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2024 , url=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.219235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.219235Z digest=sha256:0634a0c1bdb7571049fc7462e4d6f326ad9ac435846dc3e8288403fcd40ee760

Observation 478158d4-3e28-4940-9f30-97e5f055682d · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.225271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.225271Z digest=sha256:4e018cebc5c812b2096faa9d4b43dac64e6096ba810055088a6ddc068812d5bc

Observation 24e30cd2-de56-4bd5-8032-d86f246dcfc0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Code Llama: Open Foundation Models for Code

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.232907Z digest=sha256:a04b6d94fc27c396ec2499ea91db199bde517ab64f7236cf455d00ba8a572d49

Observation 3a52fc40-6cf4-4664-a5e6-d87f60e13b2b · outbound

This paper cites ToolHop:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolHop:

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.238846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.238846Z digest=sha256:e568b4fe122881e40fd49fb9c09fbc684597084bd5d74cb83d18c71a5771e1ae

Observation bd455ddb-5ec9-46e5-8f49-1a32cc7c7b27 · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.245764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.245764Z digest=sha256:593aa09649fc35ba7e28e7a37574c2c7c2f02ee01187786198b6633c234ffbe0

Observation 44819e06-aebc-406c-9097-bb6d2c07442b · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.252278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.252278Z digest=sha256:15bec3eca8427cbe91c3abc420d69380f38732a9e6e55dcfd3cb3d2565e4ddee

Observation 4081df13-c0af-4874-95c5-8689dd87b6c1 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.257857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.257857Z digest=sha256:9fc62cd485d5b25e226452c8b5c63ecb12ce5c88e5579bc57b63103cfe780e13

Pith citing papers

No inbound Pith citation observations are available.