Pith. sign in

Paper Citation Record · LEDGER

RewardHarness: Self-Evolving Agentic Post-Training

As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2605.08703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08703 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:05:25.567581Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact30
  • verified fuzzy19
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fdc92b74-d49e-46b7-92c1-4a8580a600a9 · outbound

This paper cites Blip3o-next: Next frontier of native image generation.

RewardHarness: Self-Evolving Agentic Post-Training Blip3o-next: Next frontier of native image generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.128142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:079e809e048d78fd7996ffe9ce9fd0fe15fa788fc725ed1112eb1826e9f69f7a

Observation efad21b1-b6e2-485e-93fd-5e4f8421cbb1 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

RewardHarness: Self-Evolving Agentic Post-Training Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:22.024625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:e71ead7f0f9097b1da4e20b7b805398cc5509d7155c23c1011560c765027fbb5

Observation 9eb5df50-3df0-4f3c-adac-ad2123dcd94b · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

RewardHarness: Self-Evolving Agentic Post-Training ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:42:39.160125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:8ff915ff5c55f8bc6d29cf68d0e8da079f2a975982c1e2a08283f3f3ae53544d

Observation 46d4b809-2fb4-49b2-b61e-0b093d3aa756 · outbound

This paper cites OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning.

RewardHarness: Self-Evolving Agentic Post-Training OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.121541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:ca5bd3058346cf815276b12b00778f0b1ec9eee19dc1eb684209b09f77f42cf9

Observation 375a63eb-7ebc-4918-b512-785e9187dd42 · outbound

This paper cites Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894.

RewardHarness: Self-Evolving Agentic Post-Training Toolkengpt: Augmenting frozen language models with massive tools via tool embeddings.Advances in neural information processing systems, 36:45870–45894

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.166374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:d91538b8e456006ff45bf9ca719377b290f5d5e18d59de013bb0208cd09a91f9

Observation c64d9751-e8f7-4ef4-bf28-cc6469e781aa · outbound

This paper cites Rise: reasoning enhancement via iterative self-exploration in multi-hop question answering.

RewardHarness: Self-Evolving Agentic Post-Training Rise: reasoning enhancement via iterative self-exploration in multi-hop question answering

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.171525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:8be03bf908eb7feef93d3a1dd158f488d439cd1aabfcee0b0c7e529d26f1936f

Observation d5e6ca73-60a9-421a-bcec-c4a18cabf745 · outbound

This paper cites Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation.

RewardHarness: Self-Evolving Agentic Post-Training Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.168933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:14692af0f94cc4b58e36f9eb820ef34fb20171c5ac3eb2ffd16e32e24b0ff913

Observation efea3658-4c47-4a44-b1d9-9e0b3c00777b · outbound

This paper cites Huynh-Thu, Q.

RewardHarness: Self-Evolving Agentic Post-Training Huynh-Thu, Q

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:31:24.100619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:1918de1055d1898f6821fcc52080912d641f8ca8f01c307e00aba59f543ae0a1

Observation 85d1ef97-27f3-468f-a439-dd6efda9b83b · outbound

This paper cites GenAI Arena: An Open Evaluation Platform for Generative Models.

RewardHarness: Self-Evolving Agentic Post-Training GenAI Arena: An Open Evaluation Platform for Generative Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.989227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:1fadbf9b16fa2f7bacf10a668a65e2f94b4957eeb98cc2989646d5cb1d4bf330

Observation 469c712e-747a-4417-b62b-71ad55dc4cb1 · outbound

This paper cites Verltool: Towards holistic agentic reinforcement learning with tool use.

RewardHarness: Self-Evolving Agentic Post-Training Verltool: Towards holistic agentic reinforcement learning with tool use

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.089429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:8b2132ee54050a28021c49b5f5c4209d522eb0de315bf081c107491483a1803d

Observation 0b3117d9-a646-4e91-801d-fb2c7b88082f · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36: 36652–36663.

RewardHarness: Self-Evolving Agentic Post-Training Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36: 36652–36663

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.892144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:7bf38212c0d77800f3fa7929d9a25dbf8472dacad1177f594e27f8f319c903b3

Observation 82ccfb0d-475a-4125-b6d3-f7b9b40bdd5a · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

RewardHarness: Self-Evolving Agentic Post-Training Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:01:19.940440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:368fdb3f5caea45edc711936c784d74e41b3172e6cddf581d4d0cf4668f7b805

Observation 5a2fce87-eebd-4fae-b304-d56f27dcfb1e · outbound

This paper cites Rich human feedback for text-to-image generation.

RewardHarness: Self-Evolving Agentic Post-Training Rich human feedback for text-to-image generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.900906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:fa1181be715342cd0ea7fc0352b2200b899fbd6568cb6d59b0a7c20afe7eab5c

Observation 3c323698-ace6-4928-a4cc-4d5ca961af12 · outbound

This paper cites Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning.

RewardHarness: Self-Evolving Agentic Post-Training Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.071180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:ddb4a29f1eb60478932ddb9f52dd10d0d3f958689ec38235f8e7bdae8d854a1b

Observation d5247429-535d-4695-b079-fd10bf6fc08b · outbound

This paper cites SimpleMem: Efficient Lifelong Memory for LLM Agents.

RewardHarness: Self-Evolving Agentic Post-Training SimpleMem: Efficient Lifelong Memory for LLM Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:10:45.412758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:b89c474b93d8548da433752fb3376731fbbe93454068681837689a8638e2d931

Observation 2f983c53-06db-4c18-9dee-c08c069ada61 · outbound

This paper cites Improving Video Generation with Human Feedback.

RewardHarness: Self-Evolving Agentic Post-Training Improving Video Generation with Human Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:03.116566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:a70986a44da62f26ed5efc0836b7548ebf7070106950d42ebeeb5d5e40e51fcf

Observation e176e5dd-0549-4167-b6cb-6f8c9e8c2af2 · outbound

This paper cites Editscore: Unlocking online rl for image editing via high-fidelity reward modeling.

RewardHarness: Self-Evolving Agentic Post-Training Editscore: Unlocking online rl for image editing via high-fidelity reward modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.026974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:c9310ed5e9646eca70b33fcfab6026d5fcc8b00a8055b19bb840ad950c56af6a

Observation 9671900a-b8a7-4e38-8724-c982f3b473a4 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

RewardHarness: Self-Evolving Agentic Post-Training Gorilla: Large Language Model Connected with Massive APIs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:24.048941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:5f8dce186e722c4c48a964bab69492c50ed5e37257167320310417f5fabef63a

Observation 9eaf8da1-b333-4104-ab6a-1ee984bd3a60 · outbound

This paper cites SCOPE: Prompt Evolution for Enhancing Agent Effectiveness.

RewardHarness: Self-Evolving Agentic Post-Training SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:59.937936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:3480ce9ca7b6d09cad15af70e2469cdb9e67d377ea13845216ab7595aafb1a71

Observation 7bd8e4fd-ee45-4f0f-bfba-ed497ac78968 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

RewardHarness: Self-Evolving Agentic Post-Training ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:23.948530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:a47af794252bdca2b1219626d9b279412defcd621c6ad1de2fcc7f8319f51df9

Observation 7c779478-0a76-4424-ab4a-d8dcf389bc7e · outbound

This paper cites Evolvecoder: Evolving test cases via adversarial verification for code reinforcement learning.

RewardHarness: Self-Evolving Agentic Post-Training Evolvecoder: Evolving test cases via adversarial verification for code reinforcement learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.977490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:a237858e245185c40e51506217372223c845c1f5db1abaa0fa3b4f1905a3e7c8

Observation 473738f4-cca5-4cf2-8fc7-43b629d12b27 · outbound

This paper cites ImagenWorld: Stress-testing image generation models with explainable human evaluation on open-ended real-world tasks.

RewardHarness: Self-Evolving Agentic Post-Training ImagenWorld: Stress-testing image generation models with explainable human evaluation on open-ended real-world tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.014537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:5d14080b45156d27fef199ab6cc28eb371494b19974084f8e2a8747d3030b803

Observation 1607fb23-ac9c-492e-8dec-95ab42a8861b · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652.

RewardHarness: Self-Evolving Agentic Post-Training Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.903050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:dfe4230a331c806cbe91890f993841a9417e26f03cf9d1a06fcebdca87bf0d45

Observation cc7a09bc-5203-48df-9820-0d6ac9963a4c · outbound

This paper cites Cognitive architectures for language agents.Transactions on Machine Learning Research.

RewardHarness: Self-Evolving Agentic Post-Training Cognitive architectures for language agents.Transactions on Machine Learning Research

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.913389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:5c75569bfb41627672a01f29654aeee802713f769fa145930555bc110e604c17

Observation 2619ad0c-fc3c-45c9-8374-4ed6ab95b580 · outbound

This paper cites WorldPM: Scaling Human Preference Modeling.

RewardHarness: Self-Evolving Agentic Post-Training WorldPM: Scaling Human Preference Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.995255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:0e77fa342beeb6f6a71eff73d6c6ba070d3e77f1ee16afe254f16d40f0476678

Observation e164b9ab-aaae-4b90-89bf-0807cbf678fa · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

RewardHarness: Self-Evolving Agentic Post-Training Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:24.043729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:77d77ef3f12f35859c72e857c71dced0c139a868cec8720ffc9ea9860fc49957

Observation 78ba788e-e526-4e7f-9f72-7dd17105962f · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

RewardHarness: Self-Evolving Agentic Post-Training Unified Reward Model for Multimodal Understanding and Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:44:31.060111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:89933c54c6e7334eee3a7aa994312b1bb304302ce45aca586c4234cb9c17328c

Observation 733391c0-38ea-4275-bf5a-a43a158a1a9f · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

RewardHarness: Self-Evolving Agentic Post-Training SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:27:56.991478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:42d318ffb0b48696e9f9cbdc6faecbf34dbf4aaeb13ca034e46977f009637963

Observation 651fe9d3-dd24-4206-9a4d-d1a24773df22 · outbound

This paper cites Editreward: A human- aligned reward model for instruction-guided image editing.

RewardHarness: Self-Evolving Agentic Post-Training Editreward: A human- aligned reward model for instruction-guided image editing

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:24.054975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:32ddff47e51d8bb7477ec93b4b9316594940afd2d0bc8a083e66f0b1b70fb872

Observation 8c8a0a68-b4a7-4909-8c5f-4c13558d6a15 · outbound

This paper cites EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle.

RewardHarness: Self-Evolving Agentic Post-Training EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:23.961457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:799bfb6fee42b826ef78fd02aebc300507ebc1b6e89ef0844bedee21d505328c

Observation 64120d82-032c-4be4-a6fd-92d4c2d5ef61 · outbound

This paper cites Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning.

RewardHarness: Self-Evolving Agentic Post-Training Agent0: Unleashing self-evolving agents from zero data via tool-integrated reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.955791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:369fc7a0597e18bb35fe786e0993de516939718d396c8345c419d2a866dd033f

Observation 826638e1-4f91-4922-9305-2621d8e0bdfc · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

RewardHarness: Self-Evolving Agentic Post-Training SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:39:12.115659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:9de93ffc0e4508ec36959fb4f9da2dce805988ee54f4af9054bf7f32beed2794

Observation 39106a55-240e-44d6-bf21-91eef73abc8a · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:15903–15935.

RewardHarness: Self-Evolving Agentic Post-Training Imagereward: Learning and evaluating human preferences for text-to-image generation.Advances in Neural Information Processing Systems, 36:15903–15935

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.881097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:e678a363c775bf39fba75b87a98c666cb07fb31b1ef2f439103733fe006262aa

Observation ecea3a61-c1a1-4ca3-9da6-94af92d5a38f · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

RewardHarness: Self-Evolving Agentic Post-Training VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:49:14.981353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:77fd51bab63fe82c36c454dad8ed1da71e12f594bdce9a851180056bd87cae24

Observation 0e147f5c-8042-4785-9923-2dee8dcbfac1 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

RewardHarness: Self-Evolving Agentic Post-Training DanceGRPO: Unleashing GRPO on Visual Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:24.020704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:0009653cb94dd09ea1da528ad8a68f9128bf5950c26322c6b25413ab5df7d2c2

Observation 6e4b89f5-f1c8-4211-859d-8a1cb9134090 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RewardHarness: Self-Evolving Agentic Post-Training React: Synergizing reasoning and acting in language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.908688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:be10599d067f0b5c4edb666250cd1308f192279ec856a57a4ac421825cd6f3f1

Observation c361ef97-08c0-4627-ad28-a3a3e2580100 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

RewardHarness: Self-Evolving Agentic Post-Training ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.693018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:6425a5e795ad1abb0265593bfbd6d9b5ff6fc229bfeb56a5776f848ebf9b8a5f

Observation 950bd797-e66c-4a5c-8c42-27d009a53289 · outbound

This paper cites Self-rewarding language models.

RewardHarness: Self-Evolving Agentic Post-Training Self-rewarding language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.874306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:9dd0ea3f897969cabbe03cf776177d78dbcbbb1e336209e6e2f382404bf88ddf

Observation 97123500-3014-4aa4-9b44-8674963a7ef9 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488.

RewardHarness: Self-Evolving Agentic Post-Training Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.860394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:8381afbac3c838f7f1ef97cb223a31f8a42b0e98f61aec53a314da5d81c7606c

Observation 18bb8e0e-d8ee-46d4-917c-c6b9f9807694 · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

RewardHarness: Self-Evolving Agentic Post-Training Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:44:08.223014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:a701049b21083ed36e2eab9250186fafe14cf751b72bd99e8ad1a294f4997ed0

Observation 9e05f455-34e0-4611-8690-d7e79faaf16c · outbound

This paper cites Watch Before You Answer: Learning from Visually Grounded Post-Training.

RewardHarness: Self-Evolving Agentic Post-Training Watch Before You Answer: Learning from Visually Grounded Post-Training

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:31:24.008246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:e17458b80838ef58fd171f449b90cfe57bb240ca59a8ffdda08a1f9cf865990d

Observation c653dce5-864f-4656-a50c-2729c3645630 · outbound

This paper cites Expel: Llm agents are experiential learners.

RewardHarness: Self-Evolving Agentic Post-Training Expel: Llm agents are experiential learners

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.176254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:9a66f52cb3f6d2420ced1061956f8ff0f08d4b616ab0fffc5463c217cd5a89d9

Observation 98001f8c-bc21-4200-bfbb-ef193c9cd20c · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

RewardHarness: Self-Evolving Agentic Post-Training DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:54:31.055857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:38c7fce0b53d98f454e3bf357e6bda8d207e910b03afb6e42026ea9613960201

Observation 1697992f-216d-4bb8-a5b2-5e06b100509a · outbound

This paper cites Cartoonish or heavily stylized outputs score 1–2.

RewardHarness: Self-Evolving Agentic Post-Training Cartoonish or heavily stylized outputs score 1–2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.868013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:a579952dbe8e2ac9fb7357ef739edffbd9818814e5f2cfcd44f969523d812927

Observation 9e669c7f-324d-43ac-b377-2465f808010d · outbound

This paper cites an unresolved cited work.

RewardHarness: Self-Evolving Agentic Post-Training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:46:37.884643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:56e910ddc3ae62adb25a004b99f9b63c669ea191a60c3a271e6bbf12ce6a7f12

Observation 0f93334e-fdbe-4100-b9e5-b0a21e2527e0 · outbound

This paper cites Skill: realism-and-artifact-penalties (iter 69, refined) description: Guidance on penalizing artifacts while allowing conceptual unrealism if requested by the prompt.

RewardHarness: Self-Evolving Agentic Post-Training Skill: realism-and-artifact-penalties (iter 69, refined) description: Guidance on penalizing artifacts while allowing conceptual unrealism if requested by the prompt

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.895968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:11fffe6c3db500bb8caac157be976eca4c703fae9c99c32f9956eb0329885b2d

Observation 2fefec41-e4a3-4227-8725-04105578e2f1 · outbound

This paper cites Artifacts: If the prompt requests a surreal/ impossible scenario (e.g., ‘polar bears in a savannah’ ), DO NOT penalize for being unrealistic.

RewardHarness: Self-Evolving Agentic Post-Training Artifacts: If the prompt requests a surreal/ impossible scenario (e.g., ‘polar bears in a savannah’ ), DO NOT penalize for being unrealistic

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.889360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:91420bc05c819dff2080d692cd4bf2d54a9cf758df9da571d38d2aae5cfa539c

Observation 644e4c52-5f91-4144-85c4-d8742c016656 · outbound

This paper cites an unresolved cited work.

RewardHarness: Self-Evolving Agentic Post-Training Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:46:37.840402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:bc7bf28c052213a22c787e8c669c76c9f750e46435eaed01ede19bac363a8044

Observation d3cb5738-efe3-4545-9734-4d80e939a8cb · outbound

This paper cites polar bears in a grassy savannah.

RewardHarness: Self-Evolving Agentic Post-Training polar bears in a grassy savannah

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.905436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:2afc5cf9c223599f54f973e2b99f5e3f0a947dc911527db2d5d4b638d296b768

Observation cb33423b-7090-47f2-8bbb-3735a4ab1402 · outbound

This paper cites query”: “Is this image completely black or corrupted?.

RewardHarness: Self-Evolving Agentic Post-Training query”: “Is this image completely black or corrupted?

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.899067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:11fdf1e489f0feaaa29af9bd609dba35ffabf61f267fca4b5e84caad3b7e31d5

Observation 334a64ac-76f1-4208-b7e5-615823b7818d · outbound

This paper cites Use text-and-ocr-analyzer to read the exact spelling before judging.

RewardHarness: Self-Evolving Agentic Post-Training Use text-and-ocr-analyzer to read the exact spelling before judging

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.911197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:33f5eda20103fe9de1ce61d7dfee1c183a9ea916193f2f02a9a97fd4a7bdceb3

Observation 865297ac-d243-4743-9d87-beb7d18fb5d2 · outbound

This paper cites a clear plastic bottle with a nipple.

RewardHarness: Self-Evolving Agentic Post-Training a clear plastic bottle with a nipple

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:46:37.174049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:ba683546977efb3dcd6b8f27f7ff1706366d45ee0cd26f4bca7f3ee5a6f1645c

Pith citing papers

No inbound Pith citation observations are available.