Pith. sign in

Paper Citation Record · LEDGER

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

As of 15 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 8 inbound Pith citation observations for arXiv:2506.15651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15651 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:55:19.365804Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:42.225595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.705912Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d924e4a-6e1c-4a66-a136-cfe0b695ed8a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.878456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:16.078806Z digest=sha256:9136bb15356a9a0a9c7b0264d0bb3339ee5e1d5798731d8f1142ff3d3cc3013d

Observation 6c02a401-4af8-4245-ac7f-1578f504b921 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback,.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a helpful and harmless assistant with reinforcement learning from human feedback,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.627779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:16.146217Z digest=sha256:807ece2c2d3fc48f0240d3c5eae9d2e99d43a0d12f703898f552f1f1901ec3dd

Observation d409543e-9789-4a80-a823-91960e13630e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.515092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.515092Z digest=sha256:2f541e715f08368426c74de00ab76c77015cffa91f4332386310ce9cd5e24592

Observation bf3ebc87-f99e-47ec-bca2-45f2ea52e711 · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Odin: disentangled reward mitigates hacking in rlhf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.516900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:16.653707Z digest=sha256:92e64b2cd855defb16a0af3874936444f139a604c94334c1c405cf2335460d90

Observation 8e933ef1-de09-41c8-b638-2b2aa1613da6 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:16.808290Z digest=sha256:638659fdcb67f3fbd9235eda223ccee7deca3e5cc1e1bf2a0d59935f7a687687

Observation dc5c5bfc-8105-49c2-a982-d9c7d62efdea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.919221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.919221Z digest=sha256:14d66825eb46969795c6fce020c77b61eaf6477122e98b0a9d7db1778eceb60c

Observation 6c38fd67-dc42-46cd-b62f-8d1ad1879e81 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.952432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.098160Z digest=sha256:ad8fa611bb9aa300cac5b9ecea5a8e27a0969428ef26c91c2a5d55b0e1a1954b

Observation 1632e14a-4c0e-4f8b-a419-19fc043d74dc · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf, 2025.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Reward shaping to mitigate reward hacking in rlhf, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.220976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.220976Z digest=sha256:448544c7b0ccf2ee41b614c77337e6780ac12be12f3a3d27711270d927323e3c

Observation a17db659-2111-4ca2-a188-08f38718066a · outbound

This paper cites Scaling laws for reward model overoptimization.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Scaling laws for reward model overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.687833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.370411Z digest=sha256:ba75be7f9ca513a955583f8c976fef3d4d34fa52c73ff9a6f622ebcae058c872

Observation 3b3496ff-6536-4ee3-8b78-27e842cb7a96 · outbound

This paper cites S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.395484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.475196Z digest=sha256:1a0ba0bde722808019654baf666b0155c5eaaf89b857a5565436c1e21db483eb

Observation 33486952-d28f-4202-89ff-01085059ccad · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.559872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.559872Z digest=sha256:717b93fcaecc068b3703c6d12467d9db8926b0e67a1e337a86b2078057100c08

Observation e46063c1-7856-4808-89d6-2be2185331d0 · outbound

This paper cites Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.089056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.694338Z digest=sha256:ce9e154a2858dc5fc75aca9294fc739d2caa28ca0e77aaec6f7c403a9e6d0f67

Observation 2c941af4-2b1b-46cc-81cb-7852c8327d7c · outbound

This paper cites The Llama 3 Herd of Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.770085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.770085Z digest=sha256:0c06fe34ec43d448114163cba75e690ff85de1b21c18093b292d28aaaf9781f6

Observation 168800b7-45b3-482e-883a-7fb8975cd15d · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.776544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.841397Z digest=sha256:2ae2862a8923f486a8a6a6acd7bffb141a5a15e6a8593124575b64245e4d0eb4

Observation 96037b73-3393-4f8f-9b99-617ffb2e20f2 · outbound

This paper cites Rule based rewards for language model safety.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Rule based rewards for language model safety

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.471082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:17.907897Z digest=sha256:8110d2cdba8b3a6d2601e58d4dcdee1a4c2580dbf8086b56f1d4ad5ac4436e6d

Observation 3be88102-f3ab-457f-a736-06037b856f0e · outbound

This paper cites GPT-4 Technical Report.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.001414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.001414Z digest=sha256:eb7f3587050fd7fce512bbaaa89070406962237d4005070f8b6a3b376f73fb37

Observation 6baabdaf-39d8-49c0-9a0a-198d02f2d057 · outbound

This paper cites F., Leike, J., and Lowe, R.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning F., Leike, J., and Lowe, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.197182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:18.056217Z digest=sha256:3881489c41db3f8600dcf3a4d39794df9c03662a28a059811393108f6f961256

Observation f4682ebe-c9bd-4415-aa2b-018b7cde02f4 · outbound

This paper cites D., Ermon, S., and Finn, C.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning D., Ermon, S., and Finn, C

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.034443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:18.139566Z digest=sha256:0f9a7ce0b55cd85372b2b042c78cada31ec5bc5a1e4a75253aeb114412a39e58

Observation aef2970b-9987-4889-a699-3eeb389b82eb · outbound

This paper cites Warm: on the benefits of weight averaged reward models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Warm: on the benefits of weight averaged reward models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.818159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:18.263261Z digest=sha256:445ced8a6ced956e30c6c286908cf88bc97b6273db3890676354a08313e7f696

Observation 91244e0a-5e96-480d-92cc-fd7a69542366 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.379988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.379988Z digest=sha256:9deb3ec135f7e1022283a2c57927be182e62d6c4b2e8bd31a345ba9be62f00af

Observation 1a06f961-8d08-4459-8a29-ff3f0708672a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.568219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.568219Z digest=sha256:1d564639999a4d95f0f8118bab6be654637bee1acb57f46aef10074fe79a2de8

Observation 38a3a96d-3812-48de-9d11-7143f4d24650 · outbound

This paper cites A long way to go: Investigating length correlations in RLHF, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning A long way to go: Investigating length correlations in RLHF, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.621509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:18.684210Z digest=sha256:4fb25351a8fb7a541909a09aa762a12d1865d12e7f3ee4389bfab9e87efaf869

Observation f1bce08e-2765-4650-ad99-e26236de101d · outbound

This paper cites Learning to summarize from human feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Learning to summarize from human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.885733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.885733Z digest=sha256:7ed1c87bfe1c42e225755f27c8001bc97b53f6d1e3629e2820025a2510400755

Observation fe93116b-d111-474a-a295-271f3d6f21d5 · outbound

This paper cites Interpretable preferences via multi- objective reward modeling and mixture-of-experts.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Interpretable preferences via multi- objective reward modeling and mixture-of-experts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.993349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.993349Z digest=sha256:3744115d91cf45246c9a868f5ef87a0b0d79f82cd6b7bf8ba67ee2bfc8dea1b4

Observation 5f3038dd-c30c-4500-b988-cc7392c35913 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Transforming and combining rewards for aligning large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.466949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:19.143851Z digest=sha256:1fddc482164088ee61d18cfc090bf05a0dc1b7ed62408fb4d509bd6d199fd10c

Observation 2bfa2f02-2fad-454c-b3c1-0db4f0351f45 · outbound

This paper cites E., and Stoica, I.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning E., and Stoica, I

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.274208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:19.259824Z digest=sha256:740d551cc0e409a7c2c771f19e42c5d6270f84c917bf577f6dc21708e2d97049

Observation 3e40a4df-025b-4e51-bfef-1be957b8cf7f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.329410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.329410Z digest=sha256:dfb97ff55c187dc83b37fca7e168528a1a89e10c4dce45d7ff113fe6765b0dbc

Observation 29deb25a-cbbc-4ee4-ba28-98eb74ad27e1 · outbound

This paper cites confidence.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning confidence

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:19.995163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T23:55:19.365804Z digest=sha256:e73a12336f6701f1522d5d732dfc91a599f89396812997de10206960ee2bc6df

Pith citing papers

Observation eb9dc29b-af21-4a33-837f-297378dc94bc · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 262

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.294691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:b6c953c4614765f3b3a8644182b5eb60b6768bd590efecdb5e8b3f57015759bd

Observation ada27720-584f-40ab-a0a2-a0c9f39baa49 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.225595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.225595Z digest=sha256:ebcf190618033521d8f3721633fae25f3975e41d13bf54b0d546056246212557

Observation 0ca00ec7-4d2a-4e55-9fcd-e12388adae42 · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.451741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:ebca12fc9f170ab2ced1e831027dc11b7b96f94bca847e97268848f95c7c5139

Observation bfb0290e-40fb-4bfc-a48a-58a088745556 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:05.330212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:2033dfba4101e906a02f3ff3dbfe4d51256aa29b2f5845c867d32c8c721ed6ce

Observation abdb9b9b-5c93-48c9-94f5-4a5ab4fc5248 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.112770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:c5aee08ffc999ced1c7a4e8cb1400e1ef63cc3a7bc6687fd91c1b3061ad7a95c

Observation cee7903f-9cbe-4bc3-b9c4-dbedb0d01671 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.439793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:76ed78cd0a2b12361c57e267cba1583dd2d03726df36aca1df2497895a7c93fc

Observation e2dc734c-b2b6-479d-9a4c-e9bd28f1bad1 · inbound

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge cites this paper.

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.618105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:21:37.937934Z digest=sha256:2685de735455e09de57bea3a1f4e32032bb8fc058667690d451e5cac4f87eed7

Observation 2107dac5-c1d6-48f2-8801-6c2dd0e50e20 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.707284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:cb55b1def2399a142f98838e7e7ba135a0c221bcd52a716bc0723556329aad7e