Pith. sign in

Paper Citation Record · LEDGER

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression

As of 21 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2501.12698.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12698 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:57:52.060181Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 521a53c1-ea8b-40a7-8167-eef05a458336 · outbound

This paper cites Prompted LLMs as chatbot modules for long open-domain conversation.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Prompted LLMs as chatbot modules for long open-domain conversation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.338230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:51.970700Z digest=sha256:f6cbb161cdbfcd8b7471183a0d3723f3336c69af48e0cde39f85dd8a607be9c0

Observation ea5beff2-4615-4f24-8e9a-8f1aadd1bfa1 · outbound

This paper cites Introducing chatgpt.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Introducing chatgpt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.328017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:51.974860Z digest=sha256:acc41043757bd942dffacc093cc32fe741eda3844b028e4a59a07b986ea405a2

Observation 3a19af56-10e3-4508-874a-f6197ada1066 · outbound

This paper cites Gpt-4 technical report, 2023.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Gpt-4 technical report, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.318353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:51.978324Z digest=sha256:c9e0dfe51724729b9568b2263e13d4c84504bfb3ed309c5d955e67717be569eb

Observation 66ba5ebd-f5e0-46ed-8998-13098b2eb584 · outbound

This paper cites An overview of bard: an early experiment with gener- ative ai, 2023.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression An overview of bard: an early experiment with gener- ative ai, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.308978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:51.981974Z digest=sha256:9f0b8fc641ea43b14a487f6150733283efd683cea47116941dd5274cd35abc1b

Observation ed77d2ac-2aa8-4392-b6f3-5590116d7565 · outbound

This paper cites Training language models to follow instructions with human feedback.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:51.985853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:51.985853Z digest=sha256:e877461ee1b77ad7d6127b12746aa64e625873cc626d69596a4cbebcfe04ec9d

Observation 3e10dced-73bf-4fdd-9bd6-f0c2d33117fa · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Fine-Tuning Language Models from Human Preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:51.990020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:51.990020Z digest=sha256:28e62eea68e2b223d360500606ccf91d9a8ed539e841d059115f600c76430d34

Observation 832aacf0-b176-46d4-b560-dc9e332f5244 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:51.993988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:51.993988Z digest=sha256:2bea7f1bc7b73d07f441ccb58751dd2c16ae2c753a8bb284b335fb0237bab948

Observation 60a576c6-d012-4a67-affd-de1c60fab68b · outbound

This paper cites Deep reinforcement learning from human preferences.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Deep reinforcement learning from human preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:51.998503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:51.998503Z digest=sha256:20641cab71d2bd083ae626a57d53fb6f68508ff29c74145cf5ee8d25dadd79e9

Observation 1ce9df95-f6bc-4ad1-b62b-afb4c512dc17 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.003062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.003062Z digest=sha256:f2852e8c802bfe62a1a0af52c6e44246dfd675adec75ff4dba2640ae3b83632b

Observation b070e946-c315-49d8-9c69-342c55e82370 · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.007090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.007090Z digest=sha256:657acd3b405e55c17eb4c995768b5e60240baaa4356b56940f18e8af3c58a7fe

Observation db4173e4-44c4-4c45-82be-0a61b557fa87 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Constitutional AI: Harmlessness from AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.010908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.010908Z digest=sha256:039f565364c401a7bea2f6177b095cd568c2625a5c26588deb0bc2fb30a94c09

Observation 969e6cc9-452c-4abc-8729-c90d81096288 · outbound

This paper cites Reward design with language models.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Reward design with language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.298769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.014998Z digest=sha256:a017c5083aaded717bc7ad813880ada4c4c48cf90eed65efca05511482b49c97

Observation a41c6a63-1ba1-4545-8a68-1de16ad1ad32 · outbound

This paper cites Beyond human data: Scaling self-training for problem- solving with language models.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Beyond human data: Scaling self-training for problem- solving with language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.289489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.018762Z digest=sha256:dd9929cfc836c371aa4727ecca00a1fe3c2aa833f0ef1c2eb4d9665882552ac5

Observation 258e32cf-bd87-469b-bd20-876ca8eff796 · outbound

This paper cites Language instructed reinforcement learning for human-AI coordination.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Language instructed reinforcement learning for human-AI coordination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.279162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.022299Z digest=sha256:6974e53bef8c79e84e10a5e972e3de9926121d1724afedc2772bf5a270c952a9

Observation 82349e81-bd88-4602-93ff-296fe541781e · outbound

This paper cites Unsupervised evaluation of inter- active dialog with DialoGPT.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Unsupervised evaluation of inter- active dialog with DialoGPT

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.268531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.025682Z digest=sha256:fe2bcc3ea65230f5f7726132cb2f05d5e527527bef1e6ae8767947e1b1e78f82

Observation a2943b5d-b4c7-4568-8f55-54dccde7f58a · outbound

This paper cites MEEP: Is this engaging? prompting large language models for dialogue eval- uation in multilingual settings.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression MEEP: Is this engaging? prompting large language models for dialogue eval- uation in multilingual settings

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.257899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.029045Z digest=sha256:d8f8c463e443451e1f9282a1ab10b0b8c7c05931255c70bff418e53175ff9153

Observation e1ed076c-d424-43c1-9f00-719841471d8e · outbound

This paper cites InCharacter: Evaluating personality fidelity in role-playing agents through psychological interviews.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression InCharacter: Evaluating personality fidelity in role-playing agents through psychological interviews

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.240108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.035469Z digest=sha256:6c0e1a69ea62379813146d14af231f548c6741037e5aa2eaf087b8b62041c02e

Observation d0e7d349-74da-4d6f-a641-d8a13fb10946 · outbound

This paper cites LLM-eval: Unified multi- dimensional automatic evaluation for open-domain conversations with large language models.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression LLM-eval: Unified multi- dimensional automatic evaluation for open-domain conversations with large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.229521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.041516Z digest=sha256:f88a0755843ff56176d187e04e4da6e05ff8739c5bb0b9d6120e1e19a45c13a8

Observation 2c3434ab-19bc-41e0-bb8e-54b030592973 · outbound

This paper cites G-eval: NLG evaluation using gpt-4 with better human alignment.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression G-eval: NLG evaluation using gpt-4 with better human alignment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.217994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.044585Z digest=sha256:3e5df6b04401936d5b539cf372e867187fcf39b9263cbb49e20c42dec5b8f4e5

Observation 6b8d4bfd-ad4a-4442-ae2f-011cdc7b4110 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.047712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.047712Z digest=sha256:eaae1bdfeebd7bcc595c61b36d6452c65f9ae198e1dcc9e332e8e6420b230508

Observation 234df465-0dc3-4f76-bec8-fae5966eb80a · outbound

This paper cites Empirical analysis of training strategies of transformer-based japanese chit-chat systems.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Empirical analysis of training strategies of transformer-based japanese chit-chat systems

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.206924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.051140Z digest=sha256:aa077589cb4ca2c2dec750721e66dc21c0220a89c13204007044fbc618fa2aa9

Observation 2c2ebb21-10e3-4e9c-94a5-b727b6d76725 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.054156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.054156Z digest=sha256:80269be1a10f9af28693c482d2c48d7a444143dd606e3d5dc7a1034cfb1ff094

Observation 025fdf04-d36a-4cdd-b901-6bec0bf1177f · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.057031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.057031Z digest=sha256:043004d1e8c86464b0572975569c3e3c291e920ff21856fc4290104a18e22128

Observation c1383729-fbb1-4a0f-be58-4436cfb526fe · outbound

This paper cites Towards empathetic open-domain conversation models: A new bench- mark and dataset.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Towards empathetic open-domain conversation models: A new bench- mark and dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:57:52.195351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T16:57:52.060181Z digest=sha256:bddeb17742f280e40d172e0846e84ce4f03dae5e96a1f3d45c6dc8b69bb9f29a

Observation 10529605-ac9f-4ceb-a939-db5382ff6b40 · outbound

This paper cites an unresolved cited work.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.038496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.038496Z digest=sha256:e94603476081d32330d1c42299d4353042caee9662bf0ac6612ab3dd3b111707

Pith citing papers

No inbound Pith citation observations are available.