Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:05:25.567581Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2605.08703.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:05:25.567581Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fdc92b74-d49e-46b7-92c1-4a8580a600a9 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Blip3o-next: Next frontier of native image generation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation efad21b1-b6e2-485e-93fd-5e4f8421cbb1 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9eb5df50-3df0-4f3c-adac-ad2123dcd94b · outbound
RewardHarness: Self-Evolving Agentic Post-Training ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46d4b809-2fb4-49b2-b61e-0b093d3aa756 · outbound
RewardHarness: Self-Evolving Agentic Post-Training OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 375a63eb-7ebc-4918-b512-785e9187dd42 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c64d9751-e8f7-4ef4-bf28-cc6469e781aa · outbound
RewardHarness: Self-Evolving Agentic Post-Training Rise: reasoning enhancement via iterative self-exploration in multi-hop question answering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5e6ca73-60a9-421a-bcec-c4a18cabf745 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation efea3658-4c47-4a44-b1d9-9e0b3c00777b · outbound
RewardHarness: Self-Evolving Agentic Post-Training Huynh-Thu, Q
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85d1ef97-27f3-468f-a439-dd6efda9b83b · outbound
RewardHarness: Self-Evolving Agentic Post-Training GenAI Arena: An Open Evaluation Platform for Generative Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 469c712e-747a-4417-b62b-71ad55dc4cb1 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Verltool: Towards holistic agentic reinforcement learning with tool use
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b3117d9-a646-4e91-801d-fb2c7b88082f · outbound
RewardHarness: Self-Evolving Agentic Post-Training Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36: 36652–36663
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82ccfb0d-475a-4125-b6d3-f7b9b40bdd5a · outbound
RewardHarness: Self-Evolving Agentic Post-Training Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a2fce87-eebd-4fae-b304-d56f27dcfb1e · outbound
RewardHarness: Self-Evolving Agentic Post-Training Rich human feedback for text-to-image generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c323698-ace6-4928-a4cc-4d5ca961af12 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5247429-535d-4695-b079-fd10bf6fc08b · outbound
RewardHarness: Self-Evolving Agentic Post-Training SimpleMem: Efficient Lifelong Memory for LLM Agents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f983c53-06db-4c18-9dee-c08c069ada61 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Improving Video Generation with Human Feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e176e5dd-0549-4167-b6cb-6f8c9e8c2af2 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Editscore: Unlocking online rl for image editing via high-fidelity reward modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9671900a-b8a7-4e38-8724-c982f3b473a4 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Gorilla: Large Language Model Connected with Massive APIs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9eaf8da1-b333-4104-ab6a-1ee984bd3a60 · outbound
RewardHarness: Self-Evolving Agentic Post-Training SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7bd8e4fd-ee45-4f0f-bfba-ed497ac78968 · outbound
RewardHarness: Self-Evolving Agentic Post-Training ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c779478-0a76-4424-ab4a-d8dcf389bc7e · outbound
RewardHarness: Self-Evolving Agentic Post-Training Evolvecoder: Evolving test cases via adversarial verification for code reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 473738f4-cca5-4cf2-8fc7-43b629d12b27 · outbound
RewardHarness: Self-Evolving Agentic Post-Training ImagenWorld: Stress-testing image generation models with explainable human evaluation on open-ended real-world tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1607fb23-ac9c-492e-8dec-95ab42a8861b · outbound
RewardHarness: Self-Evolving Agentic Post-Training Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc7a09bc-5203-48df-9820-0d6ac9963a4c · outbound
RewardHarness: Self-Evolving Agentic Post-Training Cognitive architectures for language agents.Transactions on Machine Learning Research
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2619ad0c-fc3c-45c9-8374-4ed6ab95b580 · outbound
RewardHarness: Self-Evolving Agentic Post-Training WorldPM: Scaling Human Preference Modeling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e164b9ab-aaae-4b90-89bf-0807cbf678fa · outbound
RewardHarness: Self-Evolving Agentic Post-Training Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 78ba788e-e526-4e7f-9f72-7dd17105962f · outbound
RewardHarness: Self-Evolving Agentic Post-Training Unified Reward Model for Multimodal Understanding and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 733391c0-38ea-4275-bf5a-a43a158a1a9f · outbound
RewardHarness: Self-Evolving Agentic Post-Training SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 651fe9d3-dd24-4206-9a4d-d1a24773df22 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Editreward: A human- aligned reward model for instruction-guided image editing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c8a0a68-b4a7-4909-8c5f-4c13558d6a15 · outbound
RewardHarness: Self-Evolving Agentic Post-Training EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64120d82-032c-4be4-a6fd-92d4c2d5ef61 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 826638e1-4f91-4922-9305-2621d8e0bdfc · outbound
RewardHarness: Self-Evolving Agentic Post-Training SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 39106a55-240e-44d6-bf21-91eef73abc8a · outbound
RewardHarness: Self-Evolving Agentic Post-Training Imagereward: Learning and evaluating human preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:15903–15935
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecea3a61-c1a1-4ca3-9da6-94af92d5a38f · outbound
RewardHarness: Self-Evolving Agentic Post-Training VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e147f5c-8042-4785-9923-2dee8dcbfac1 · outbound
RewardHarness: Self-Evolving Agentic Post-Training DanceGRPO: Unleashing GRPO on Visual Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e4b89f5-f1c8-4211-859d-8a1cb9134090 · outbound
RewardHarness: Self-Evolving Agentic Post-Training React: Synergizing reasoning and acting in language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c361ef97-08c0-4627-ad28-a3a3e2580100 · outbound
RewardHarness: Self-Evolving Agentic Post-Training ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 950bd797-e66c-4a5c-8c42-27d009a53289 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Self-rewarding language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97123500-3014-4aa4-9b44-8674963a7ef9 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18bb8e0e-d8ee-46d4-917c-c6b9f9807694 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9e05f455-34e0-4611-8690-d7e79faaf16c · outbound
RewardHarness: Self-Evolving Agentic Post-Training Watch Before You Answer: Learning from Visually Grounded Post-Training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c653dce5-864f-4656-a50c-2729c3645630 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Expel: Llm agents are experiential learners
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 98001f8c-bc21-4200-bfbb-ef193c9cd20c · outbound
RewardHarness: Self-Evolving Agentic Post-Training DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1697992f-216d-4bb8-a5b2-5e06b100509a · outbound
RewardHarness: Self-Evolving Agentic Post-Training Cartoonish or heavily stylized outputs score 1–2
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9e669c7f-324d-43ac-b377-2465f808010d · outbound
RewardHarness: Self-Evolving Agentic Post-Training Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f93334e-fdbe-4100-b9e5-b0a21e2527e0 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Skill: realism-and-artifact-penalties (iter 69, refined) description: Guidance on penalizing artifacts while allowing conceptual unrealism if requested by the prompt
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fefec41-e4a3-4227-8725-04105578e2f1 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Artifacts: If the prompt requests a surreal/ impossible scenario (e.g., ‘polar bears in a savannah’ ), DO NOT penalize for being unrealistic
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 644e4c52-5f91-4144-85c4-d8742c016656 · outbound
RewardHarness: Self-Evolving Agentic Post-Training Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3cb5738-efe3-4545-9734-4d80e939a8cb · outbound
RewardHarness: Self-Evolving Agentic Post-Training polar bears in a grassy savannah
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb33423b-7090-47f2-8bbb-3735a4ab1402 · outbound
RewardHarness: Self-Evolving Agentic Post-Training query”: “Is this image completely black or corrupted?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 334a64ac-76f1-4208-b7e5-615823b7818d · outbound
RewardHarness: Self-Evolving Agentic Post-Training Use text-and-ocr-analyzer to read the exact spelling before judging
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 865297ac-d243-4743-9d87-beb7d18fb5d2 · outbound
RewardHarness: Self-Evolving Agentic Post-Training a clear plastic bottle with a nipple
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.