Pith. sign in

Paper Citation Record · LEDGER

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 8 inbound Pith citation observations for arXiv:2506.15651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15651 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:55:19.365804Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:42.225595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.705912Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d924e4a-6e1c-4a66-a136-cfe0b695ed8a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.878456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:16.078806Z digest=sha256:f61803ecd5d5647759377e049aa365ea81ce5f810bbd74c4dd90276b07241d16

Observation 6c02a401-4af8-4245-ac7f-1578f504b921 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback,.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a helpful and harmless assistant with reinforcement learning from human feedback,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.627779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:16.146217Z digest=sha256:d61eddf2a973021841ad1fa7350c997ffbf9ad1302195938fa2c2ee73b8ba2fb

Observation d409543e-9789-4a80-a823-91960e13630e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.515092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.515092Z digest=sha256:6f78edc4826f933f2aab9326aeb6353b281027b2ead3a5631a1ff161e86031ed

Observation bf3ebc87-f99e-47ec-bca2-45f2ea52e711 · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Odin: disentangled reward mitigates hacking in rlhf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.516900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:16.653707Z digest=sha256:eb5056c30f7ca74e679173b1c4055633404fcf1a31bc50df283ea4c6dc58fdf7

Observation 8e933ef1-de09-41c8-b638-2b2aa1613da6 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:16.808290Z digest=sha256:f7d3ba3d152a7a3357e96cadc46e15990f6721e0c01d83b4ca8a769d3af29635

Observation dc5c5bfc-8105-49c2-a982-d9c7d62efdea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.919221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.919221Z digest=sha256:e7aea337a023037facaa4560bae8932b74f9deef0358b1556225d56dcabbee3c

Observation 6c38fd67-dc42-46cd-b62f-8d1ad1879e81 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.952432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.098160Z digest=sha256:2ca620449e211f26d4f63dab601917228e4205c2e92198393d38cdc83f7829e3

Observation 1632e14a-4c0e-4f8b-a419-19fc043d74dc · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf, 2025.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Reward shaping to mitigate reward hacking in rlhf, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.220976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.220976Z digest=sha256:a2231e362274211b6a357a913e1719e78f9ef0a811568afbdd2e7a10e57c5cdf

Observation a17db659-2111-4ca2-a188-08f38718066a · outbound

This paper cites Scaling laws for reward model overoptimization.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Scaling laws for reward model overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.687833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.370411Z digest=sha256:8e4f7eac82f7ffbd34ccb21e1081042e69c06d99b57c541d90e3fe002971d40c

Observation 3b3496ff-6536-4ee3-8b78-27e842cb7a96 · outbound

This paper cites S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.395484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.475196Z digest=sha256:8aba8ded4aa5f18c7dcb47f8c7b69cd7aa451ffd4f4fa63518e86dd104e2f1dd

Observation 33486952-d28f-4202-89ff-01085059ccad · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.559872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.559872Z digest=sha256:58a800c90c22b1f09035c71f2b01b2c91dca7f30d5ff7409dd22186ea94a0eb2

Observation e46063c1-7856-4808-89d6-2be2185331d0 · outbound

This paper cites Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.089056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.694338Z digest=sha256:8da0e2b8cde372c69148c121ddcf2b537d0fb8f3e1d3310c327ec3dda0c5768b

Observation 2c941af4-2b1b-46cc-81cb-7852c8327d7c · outbound

This paper cites The Llama 3 Herd of Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.770085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.770085Z digest=sha256:3a7ded1ede530c98bb0c0d2fda6cb10cfd57195f0312badddfeb20d4f3f9e425

Observation 168800b7-45b3-482e-883a-7fb8975cd15d · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.776544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.841397Z digest=sha256:9e9c58a09603ceefb535a561573f8e5e3af76da9631f33cb32cf4e0d7f5e698a

Observation 96037b73-3393-4f8f-9b99-617ffb2e20f2 · outbound

This paper cites Rule based rewards for language model safety.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Rule based rewards for language model safety

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.471082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:17.907897Z digest=sha256:61a85dd24ac2356f45c3c705ec5fb8dace1e4ecb5325eb7e9d5dc3d97fe0287a

Observation 3be88102-f3ab-457f-a736-06037b856f0e · outbound

This paper cites GPT-4 Technical Report.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.001414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.001414Z digest=sha256:5832a0c3fe8f1251d2757c4cf8725753ce2d1defc236c48e4583028fe385f85b

Observation 6baabdaf-39d8-49c0-9a0a-198d02f2d057 · outbound

This paper cites F., Leike, J., and Lowe, R.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning F., Leike, J., and Lowe, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.197182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:18.056217Z digest=sha256:5a71517b5e7e96fb6d8d0c39f96d2a33a45cd3b79b70217dccd9faa1545bff5f

Observation f4682ebe-c9bd-4415-aa2b-018b7cde02f4 · outbound

This paper cites D., Ermon, S., and Finn, C.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning D., Ermon, S., and Finn, C

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.034443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:18.139566Z digest=sha256:47c953379871b77b45084ae3284890b411658f8138dd8ad146b4ba941ef280d9

Observation aef2970b-9987-4889-a699-3eeb389b82eb · outbound

This paper cites Warm: on the benefits of weight averaged reward models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Warm: on the benefits of weight averaged reward models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.818159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:18.263261Z digest=sha256:fc91d000d1b400cbaad0113f877c7be0ee7c90f07be69745a4fac08a111921dc

Observation 91244e0a-5e96-480d-92cc-fd7a69542366 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.379988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.379988Z digest=sha256:4219212696841cbc292858d1c20c382c0617a38bba5b693dcd276d0f6a5245ee

Observation 1a06f961-8d08-4459-8a29-ff3f0708672a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.568219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.568219Z digest=sha256:d9cf115d52f9345e732a05596a0389ed345cd58f86c8274e4688947e735f9acd

Observation 38a3a96d-3812-48de-9d11-7143f4d24650 · outbound

This paper cites A long way to go: Investigating length correlations in RLHF, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning A long way to go: Investigating length correlations in RLHF, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.621509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:18.684210Z digest=sha256:2a596c3ccc1bdb2d91b2cc8b2abce10d80f9bb7924bae43c4d922ec1295b27f0

Observation f1bce08e-2765-4650-ad99-e26236de101d · outbound

This paper cites Learning to summarize from human feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Learning to summarize from human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.885733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.885733Z digest=sha256:008c0727b1decea64aed7008e88ee79de19be761c511b2da9a4fafd24acf7350

Observation fe93116b-d111-474a-a295-271f3d6f21d5 · outbound

This paper cites Interpretable preferences via multi- objective reward modeling and mixture-of-experts.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Interpretable preferences via multi- objective reward modeling and mixture-of-experts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.993349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.993349Z digest=sha256:491fe0a421f924c577724970db27a1f4aa985120836f0d5ecb11a9de8ed9bbf7

Observation 5f3038dd-c30c-4500-b988-cc7392c35913 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Transforming and combining rewards for aligning large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.466949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:19.143851Z digest=sha256:23297ac9018470a23194f553396ab4e6d65c93840f21b9dcf291ca0476b3b3d3

Observation 2bfa2f02-2fad-454c-b3c1-0db4f0351f45 · outbound

This paper cites E., and Stoica, I.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning E., and Stoica, I

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.274208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:19.259824Z digest=sha256:95c56ba6f828b1b8364af5e56edc0bd5806c3cdb10c4c35ec50a1cc50aaee62d

Observation 3e40a4df-025b-4e51-bfef-1be957b8cf7f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.329410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.329410Z digest=sha256:da9ab951dbbf10d38b36392ac68f2b9932612cd1bc65c51cb493c748857dac08

Observation 29deb25a-cbbc-4ee4-ba28-98eb74ad27e1 · outbound

This paper cites confidence.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning confidence

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:19.995163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:55:19.365804Z digest=sha256:5aec43d3418209b39f0bbfcf59f9600a94c1b3e29b35f6316dd3fb7221626c03

Pith citing papers

Observation eb9dc29b-af21-4a33-837f-297378dc94bc · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 262

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.294691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:c09ca55b9858963b4432ed4305ce130f4b5c85475f31914294dfc83310c75f60

Observation ada27720-584f-40ab-a0a2-a0c9f39baa49 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.225595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.225595Z digest=sha256:2f8efe12c59eece17e143c5ad21ff51f2558b9bcd566732e2cb6b507e6917a5f

Observation 0ca00ec7-4d2a-4e55-9fcd-e12388adae42 · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.451741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:f9ec7afc3c8a90d8bf5ae5da81bd591dcfe64f42a0a6b548746cb6eca5042798

Observation bfb0290e-40fb-4bfc-a48a-58a088745556 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:05.330212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:166f54ffe27f8fe0ee1116844761ba962e87262435a90457c454ceea76cd692c

Observation abdb9b9b-5c93-48c9-94f5-4a5ab4fc5248 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.112770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:f2d73016212300a8b61091c4358fcc11b2cab64829a00f437d4466d431f09bb8

Observation cee7903f-9cbe-4bc3-b9c4-dbedb0d01671 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.439793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:f6eea338338002715baed5509d7889ae8770e8cacac00ff5f459324df549c504

Observation e2dc734c-b2b6-479d-9a4c-e9bd28f1bad1 · inbound

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge cites this paper.

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.618105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T07:21:37.937934Z digest=sha256:42fd4382cd9ca3d7d26a4037f84c7fce98834fa656c26a86a63b207d1d7688be

Observation 2107dac5-c1d6-48f2-8801-6c2dd0e50e20 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.707284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:f51dbbb5be934ab0df62df3a8be025c963e5d64b56fd978d41d1b493c6444f99