Pith. sign in

Paper Citation Record · LEDGER

ARIA: Training Language Agents with Intention-Driven Reward Aggregation

As of 22 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2506.00539.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00539 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:04.438910Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:42:59.909229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:48:12.945063Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80ad7cea-d893-42d4-8fb8-8d1dc96e9c22 · outbound

This paper cites A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.443310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.443310Z digest=sha256:3db685d8960546270175f0f8c2b030345b8643f6548a1f0467d9f95d82fd7b3d

Observation 9620f307-85c9-438c-ac40-08993cec4f70 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.562789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.562789Z digest=sha256:642c18690baec55ed4adef2312fdaaad5e50cfecf406160d4b2d21abd7026683

Observation 08dc02f7-5467-446a-b50b-0c95824aed5b · outbound

This paper cites Cognitive architec- tures for language agents.Transactions on Machine Learning Research, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Cognitive architec- tures for language agents.Transactions on Machine Learning Research, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.722060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.722060Z digest=sha256:efc9fb553666058fa6f0c47f8738e09b46c13455423eedb2a7ecb4b67ed1a7ff

Observation 26122e5f-d53f-416d-afdd-296c44456f03 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:58.885758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:58.885758Z digest=sha256:4f2720d6440e1f18f52f0ebb1ea82d52ca753377c15d55d1bf22e144e83932ab

Observation c161d582-e1ed-4dfe-80a8-3d1af98e887f · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.066328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.066328Z digest=sha256:c56f3bf85e02fd78b179834846ea52c15313de19b2681aaf8bb0abd9a1d14902

Observation 8b3b0c7d-8cac-47ed-820c-658eff6bda9e · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.239397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.239397Z digest=sha256:d7d01c046f601769a450a5e56c9f2bf02fdf2054d85cd547c9d0fecd4d3445e2

Observation 5cbe73be-554d-47b1-a061-01188b55cced · outbound

This paper cites TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:05.090357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:08:59.403455Z digest=sha256:bb94cd1ceb711b9be68312e8fe8db8fdbc0c6dfb507ff87abf5bcf0c861e82ff

Observation a042abfe-c539-4c27-a6c1-e58591cbdc71 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.541246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.541246Z digest=sha256:392f99dc29ab66721ea0e9816377951effd8807214f54a3068e09e45466f3598

Observation b75ecb96-8246-4666-9bf4-f38c158d1f4f · outbound

This paper cites Evaluating language model agency through negotiations.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Evaluating language model agency through negotiations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.754365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.754365Z digest=sha256:dc1f30dbf4638bf974ed919e9224b78360c898aa2f71a1a1dc4865bc8bd94438

Observation ad64771c-0f22-46dd-96ec-5627d4093b27 · outbound

This paper cites How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:59.948196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:59.948196Z digest=sha256:aeb362dcb79dbabe760f7307320d929c8b06f41eb3eae7e7e6b14ee60f917f86

Observation 46b68e55-dc53-4d56-9ebf-81317dd519d4 · outbound

This paper cites CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.107183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.107183Z digest=sha256:612f27bfc4406777fae201ba880f1e7bb15234072c125a5afa53249d33ef0068

Observation 4a26fead-1325-49cd-8a15-53c0823dec83 · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.264859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.264859Z digest=sha256:b80d47c218d421e1d57b6738f123273f6471ce7071875e5a44fc348f0ce0ba25

Observation b0282fc0-9feb-4ac3-a30a-0f6c441ec3e6 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.382912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.382912Z digest=sha256:97b6d4394e8c58237874b9c4d4700b5897482304259a2da4e1465d6da6192c21

Observation 64dda7ae-cab9-4447-b0e6-144f1c6b5289 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.454738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.454738Z digest=sha256:fe97445357d9b1c01e875e7f8418f406871893b4f652eb57d62c6a9aad8d4dda

Observation acfa2ab7-d82e-4c26-a9d1-1892540ca541 · outbound

This paper cites Selfgoal: Your language agents already know how to achieve high-level goals.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Selfgoal: Your language agents already know how to achieve high-level goals

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:00.548686Z digest=sha256:4725dafd0fa61297f664819a9868fe64cabe8d5d325cf7085c5698f184a0045b

Observation 60fadc7d-c7a7-4280-a143-c016ad995460 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.612952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.612952Z digest=sha256:fe10e726d882636b6a989659e4cc8b3f19f90d1fcf5126efe25a93792e701f9d

Observation f8f102a7-b5ef-4855-98d4-bbc843c64d73 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Webshop: Towards scalable real-world web interaction with grounded language agents, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.701429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.701429Z digest=sha256:6185360f786d595ad4f5a3e5810939ee882615f716f30bd6004d6347ea06c5a6

Observation 8385a51d-7781-4cde-b779-bdf200c5f991 · outbound

This paper cites Deal or No Deal? End-to-End Learning for Negotiation Dialogues.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Deal or No Deal? End-to-End Learning for Negotiation Dialogues

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.787910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.787910Z digest=sha256:7e3d6f8dfbe6172f0f65648c41e6e519271b3be5b51dca418a4b70946a6c5172

Observation 05755954-fad1-44ff-9e02-6556f7349060 · outbound

This paper cites Glee: A unified framework and benchmark for language-based economic environments.arXiv preprint arXiv:2410.05254, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments.arXiv preprint arXiv:2410.05254, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.857814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.857814Z digest=sha256:e2e6e8257883d15b82ba28870bfaf2bd7b603443678040321d3afff40fbb47f1

Observation b389c14e-e26e-4d02-9c71-a9b0eadee153 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:00.985006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:00.985006Z digest=sha256:8b778ccc9ba79448fbe810b3e09bf9ac8aa82a93822c7941e79595679ad46c15

Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.135410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.135410Z digest=sha256:2584b168c7c9995ed18c4247d47018f7a135f5f90890f708b89074404cbc59ed

Observation 7b7ff850-89b1-49fd-98db-56015d0e1fab · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.250499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.250499Z digest=sha256:ccef8787cea7001e80f959fb651b59f2766beaf43716766a1e9147f1221cee59

Observation 11531a67-1dfa-4331-ad07-68535afe723d · outbound

This paper cites Williams.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Williams

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.348373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.348373Z digest=sha256:fe1025ecc2f49c0519cff2bdd7e9cd99e33b1b566eb7fe32326ed5aaa23369d6

Observation 2f549245-1f64-40f6-a229-01c395381c2a · outbound

This paper cites Hierarchical grouping to optimize an objective function.Journal of the American statistical association, 58(301):236–244, 1963.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Hierarchical grouping to optimize an objective function.Journal of the American statistical association, 58(301):236–244, 1963

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.453841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.453841Z digest=sha256:a57c907f17ee04c9b64041edff21453281c7ab7fc8c8d9c91c54875bdc0de02c

Observation 23c052aa-133b-453d-be09-6464489841e1 · outbound

This paper cites SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.633394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.633394Z digest=sha256:3675d26db674a19486377107bb22e9a75c0a80cb2e6f257ab3a620784ced22c6

Observation 29b6e2d0-369f-4eb5-a5e5-42983e1c28e9 · outbound

This paper cites Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.819815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.819815Z digest=sha256:cb54ecc5bf50fa400e64c331d7097b80c045be5799aa23139995d6ad1d0fe948

Observation f7e1c9e0-5ab3-441d-a0f4-7b0c98d53ea7 · outbound

This paper cites AvalonBench: Evaluating LLMs Playing the Game of Avalon.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation AvalonBench: Evaluating LLMs Playing the Game of Avalon

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.061227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.061227Z digest=sha256:b873adc667d597a4f16ea7ace9b30c21008aff2e032892c8be9f7bb3dbcfb658

Observation d2858b76-b913-4b2c-b2fd-2b24a6e29b11 · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Self-playing adversarial language game enhances llm reasoning.Advances in Neural Information Processing Systems, 37:126515–126543, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.356809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:02.194205Z digest=sha256:33e2f20e23a4c0f7a92d9c17a52fd26996326bbfe0bdd3c230310a2f7833c054

Observation bbd1ebb1-33ab-4190-9d6e-d788077e060c · outbound

This paper cites GameEval: Evaluating LLMs on Conversational Games.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation GameEval: Evaluating LLMs on Conversational Games

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.360560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.360560Z digest=sha256:53786aff1b87dcb363db6c0d2f50988a5e1e957bfd57590a66d9d1db767e7b7d

Observation 0c288d72-2f51-4edd-9cb1-e6f6d4a8f2c4 · outbound

This paper cites Least squares quantization in pcm.IEEE transactions on information theory, 28(2):129–137, 1982.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Least squares quantization in pcm.IEEE transactions on information theory, 28(2):129–137, 1982

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.463723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.463723Z digest=sha256:0055665d7d61ea2f14a3b679b3e725b7e4620167510e42d371a5037aaee70cad

Observation bb4b8319-ba75-4af0-8f3e-c49a25c75d6b · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:08.105891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:02.575342Z digest=sha256:3f39ab7e69128d25cd95026ab030f6aa4bc4c61f069c3cd7baedc27deddd506c

Observation dc0febd7-dfd3-4257-b303-d6d35b337bc1 · outbound

This paper cites PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.673624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.673624Z digest=sha256:41890fdc9ef6f099fdb7921f02baffaeff30980edca3de3e8cc83e2c6309d473

Observation 5d6ad255-2e81-4393-a49a-7658872f2539 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Direct preference optimization: Your language model is secretly a reward model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.783375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.783375Z digest=sha256:4ddb021f457cee483de8dc0a936ca52cb16386d8a47d4e030d4f052ae954fb36

Observation 25cb6302-961a-4c3d-8c72-d1278b151cd9 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation KTO: Model Alignment as Prospect Theoretic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.856184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.856184Z digest=sha256:867b9b1dce798939471e1b3616c111be131e874273d56cf5ce350f3e49f427da

Observation 6833c2ab-c860-44eb-bb1d-7b5757caf38d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:02.951993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:02.951993Z digest=sha256:73fffeb4434d6855d5eb4d7d4fcb585230fd340a6f5f377a1c091d4f5481d67a

Observation c1d10df5-5527-4490-9578-e69f4bc0dc85 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.054149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.054149Z digest=sha256:46b469e73f3416f082842b5c12065008329f397d16e40a1d1e5d6eb52b3976ff

Observation d1674156-629d-4af9-9f73-e7dcc6888ff6 · outbound

This paper cites Methods of Hierarchical Clustering.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Methods of Hierarchical Clustering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.131933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.131933Z digest=sha256:12fe869d5b99486cd4d3914ccb300040108f380369304a8b69e06eb9b57f0b0e

Observation 00333b33-ec03-47de-b745-e3d1bd363535 · outbound

This paper cites Silhouettes: a graphical aid to the interpretation and validation of cluster analysis.Journal of computational and applied mathematics, 20:53–65, 1987.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Silhouettes: a graphical aid to the interpretation and validation of cluster analysis.Journal of computational and applied mathematics, 20:53–65, 1987

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.225444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.225444Z digest=sha256:669e86e4946120f283e738ada433b6c0eaaba920b3258eb3675bb34aa74e7ca7

Observation aa50cd92-6e53-479a-ae57-4bd157a20113 · outbound

This paper cites A dendrite method for cluster analysis.Communications in Statistics-theory and Methods, 3(1):1–27, 1974.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A dendrite method for cluster analysis.Communications in Statistics-theory and Methods, 3(1):1–27, 1974

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.304018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.304018Z digest=sha256:8b2b0f3b0c745d7da39fc59354d84a9f99b55c224f4aa0c5840cccc6cb3dd81a

Observation bca82fdc-9a2f-471e-bf1b-6e772530ab90 · outbound

This paper cites A cluster separation measure.IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 2009.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation A cluster separation measure.IEEE transactions on pattern analysis and machine intelligence, (2):224–227, 2009

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.852623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:03.440574Z digest=sha256:13f8389947161158952f847f6c82f22c4f9918bdf9a7d209a2889391bccaf6dc

Observation 3285c6a9-3962-4f20-a7e0-e68f3db952a8 · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:07.565203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:03.516997Z digest=sha256:230f26ab7d9157987188fb38d14e8f0b557949044117fa2dea759141a9ad5d3f

Observation 3cee0ba5-e240-456c-a8e8-3aaa2c5bf7a5 · outbound

This paper cites Training agents by reinforcing reasoning, 2025.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Training agents by reinforcing reasoning, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.345382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:03.619683Z digest=sha256:3bba7967af92facf320a7e065324fc07f36d0758eda67318216220f244df0bfa

Observation f040559d-b41b-48ae-92a5-ffc0336ca356 · outbound

This paper cites Glee: A unified framework and benchmark for language-based economic environments, 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Glee: A unified framework and benchmark for language-based economic environments, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:07.099292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:03.705946Z digest=sha256:8b16754d9288e70d3f6c5e043e0e1c4722820130788bb7270fb9b891fc38720b

Observation 3865c1e6-fa12-4026-a4bd-f08b25df12c6 · outbound

This paper cites Llama 3 model card.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Llama 3 model card

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.818197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:03.782318Z digest=sha256:f6dbcede351e9c9021522671c43ff0553c0bde4ab083766055338fe726d18dcb

Observation 459696a0-c5ae-4e1e-9e10-1c9e624a0749 · outbound

This paper cites Text and code embeddings by contrastive pre-training, 2022.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Text and code embeddings by contrastive pre-training, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:03.880519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:03.880519Z digest=sha256:94b08cc472840c21fb2786c93aa1d1b870ed5fc77c99d259196e169526181086

Observation 2ee2879c-ef7a-4d54-9b9e-abc74446ea9f · outbound

This paper cites Gpt-4 technical report, 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Gpt-4 technical report, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:04.011359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:04.011359Z digest=sha256:b84c41692d94430d91ecbc955793ddca25c468e95a74046a7862e252227eac1a

Observation aae6b85f-644d-4cc7-a598-cc367b24b11c · outbound

This paper cites Introducing claude 2.1, Nov 2023.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Introducing claude 2.1, Nov 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.679584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:04.104372Z digest=sha256:a617300b7e8278d3f309ed8df4f8e1a77b6b702a3b8b64e0fa2148dc97897528

Observation 869581f2-df74-4623-807b-0e43b3fb49e7 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:06.305099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:04.178352Z digest=sha256:cf2ae332ad26b696f00a032bb253316d1c2157dfbc85cf0180aba2679f73b330

Observation 7612ea95-a64f-4ab3-8ca7-c8bef3c4fd24 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Qwen2.5: A party of foundation models, September 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:05.910016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:04.271744Z digest=sha256:78c8df8311b5339464a15d053ee79ee0fccadaa64de617680345f5f1a54b82ba

Observation 91bc5c2e-1b1a-40d5-9ef2-e5bd3ba6696e · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:05.625633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:04.344497Z digest=sha256:73e8562cc31289e01928debc3b07a699ee7a0357ecd28d1a8a208f9299c46ddf

Observation fddb8384-c1c3-43e1-8a42-afdff67bd4ba · outbound

This paper cites an unresolved cited work.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:05.376816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:09:04.438910Z digest=sha256:6d65ef3c9818d3ac4c558191141fa7bf8ebd6eb6a6d29d18e9f5b50edaf4aa01

Pith citing papers

Observation c3128c10-4099-4755-8cb2-2203ceb52e11 · inbound

Unsupervised Learning for the Elementary Shortest Path Problem cites this paper.

Unsupervised Learning for the Elementary Shortest Path Problem ARIA: Training Language Agents with Intention-Driven Reward Aggregation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:42:59.909229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:42:59.909229Z digest=sha256:23acb926f0da3fead3dfffef1c1e510a771cfa7988d69e54cbee52a339c9136c

Observation 6d425eb3-e2b6-43f7-87c1-d24c0a9c6d16 · inbound

Latent Action Reparameterization for Efficient Agent Inference cites this paper.

Latent Action Reparameterization for Efficient Agent Inference ARIA: Training Language Agents with Intention-Driven Reward Aggregation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:48:12.946615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:45:20.306945Z digest=sha256:3ae85dded90fc6c937a9a5f8c53a1e6996734d05290b409fb986c15a34a85bc5