Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:57:12.869989Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 14 inbound Pith citation observations for arXiv:2502.04306.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:57:12.869989Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:02.869499Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T00:12:50.066065Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d3730524-bea0-40f6-9a2c-abdb6403fa3f · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f5f5bf-e1a0-46e3-a21e-6ac3c9782fb8 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization PaLM 2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4037b6a2-85c3-4b69-b849-17dc8a9ab4b2 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abad24c7-c5c3-4cd5-a387-e10f708402d2 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization A general theoretical paradigm to understand learning from human preferences
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6715730b-9f74-4afe-beba-3889b44ef644 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Rank analysis of incomplete block designs: I
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783e1512-f46d-4dda-8cd1-732d3404c361 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Preference Learning Algorithms Do Not Learn Preference Rankings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db93fe2f-df8b-4c0a-8d2f-fef6567543bb · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization AutoAgents: A Framework for Automatic Agent Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe9f0a0-813a-492a-8ed9-a5e07f1aeae5 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39132a60-e940-4c92-b3b1-5acf31439456 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed8e6d1-280b-488f-a08c-4529977cc6b1 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a688a3-97ee-4aec-b0a7-5bcd43a35937 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 523419a9-71b1-49a0-bd12-65b3a6e44f69 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Data Interpreter: An LLM Agent For Data Science
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17995c6-1247-46e8-8fd5-82887e161ca6 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66f9f549-15d9-4508-9cce-51f1e93e5fb4 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Automated design of agentic systems
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9036684-3777-4203-a531-692cecb8b42b · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34217c21-d590-4e49-9398-d3b734f99674 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eacf067-512e-41a4-946a-f8b685c3d21f · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Dspy: Compil- ing declarative language model calls into state-of-the-art pipelines
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a57525fb-7c3a-44ad-8b1d-0b1ebe9974ad · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aec9b39-8049-4888-bd3c-0ba7e85892e1 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization AutoFlow: Automated Workflow Generation for Large Language Model Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9182281b-7a15-4fd4-9b4f-7ffc17c828cd · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization DeepSeek-V3 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dcd6c3-db4d-424a-997a-ca66c168ca8c · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization A dynamic llm-powered agent network for task-oriented agent collaboration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68acb43a-21cb-41fb-8287-a80311037be4 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Self-refine: Iterative refinement with self-feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation afa58b2c-fcc2-4aad-b839-c1d5bb82114a · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c524b6e-e454-4f60-8800-0d5e6d45c54e · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 995c473c-f3e3-40b5-b607-113fb98b3c4b · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Training language models to follow instructions with human feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e09a88e1-aa5e-4373-adf5-19920e30f149 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Disentangling length from quality in direct preference optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96c6840c-58d5-4b4e-9a3d-c39a35a5bcaa · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Direct preference optimization: Your language model is secretly a reward model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730793c5-2d09-4834-bd38-8a53edb308ad · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7637e98-1356-47e2-bc61-66235a1b08d6 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Archon: An Architecture Search Framework for Inference-Time Techniques
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590e455e-1797-4d30-b5a0-3605336b837d · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1552c035-da6d-403d-9411-129bc9a05d11 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Reflexion: Language agents with verbal reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf0d7583-db8f-4d2a-a930-ba87247585b7 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Adaptive In-conversation Team Building for Language Model Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 957d2528-5de7-43b6-9827-a7763559cb4f · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa4f3c3-9a2d-4b3b-a7b0-b030e6a684a6 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Self-consistency improves chain of thought reasoning in language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a66ec3b-ca81-4111-8059-3af9c63ef40c · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi- persona self-collaboration
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0d8999c4-e272-4ada-9252-1e886f0e2e92 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Chain-of-thought prompting elicits reasoning in large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75839ae2-a980-484e-ad82-47651ca6bd13 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Is dpo superior to ppo for llm alignment? a comprehensive study
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 29de14fb-a760-4134-8026-859dbf0b7960 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Lemur: Harmonizing natural language and code for language agents
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af447a64-30b5-454f-a568-7a1014933a12 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Qwen2.5 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872f6fb6-0f25-49dd-8a5c-51a2cfa4610e · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Large Language Models as Optimizers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34433255-8367-41b9-b74a-996117b5ff42 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Buffer of thoughts: Thought-augmented reasoning with large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8523c9a-d6f9-495e-a4c8-9ba3a264929d · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-Correction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eebd04e1-989f-4d39-bc8d-b7abc32cb453 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 056913e2-58e7-4a41-b99d-e8a92a24c65e · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization TextGrad: Automatic "Differentiation" via Text
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28aae596-721c-49e4-9ebf-e193d166fa80 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84e06ba-b204-4691-869a-70b5adf62783 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization AFlow: Automating Agentic Workflow Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa3e9ba-73b0-476d-9f36-219b6f159f2b · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae5b654-d6f7-4043-b10e-578f07115c06 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Symbolic Learning Enables Self-Evolving Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2e3f73-e73a-4640-99e3-4f25bb359481 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization "" This is a wor kf low graph
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7763ed-99b6-40b2-b47c-9a5e5f0c87b0 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Format MUST follow : custom ( i n s t r u c t i o n : str ) -> str You can modify the i n s t r u c t i o n prompt
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f7132786-fb17-40e0-877a-7e974988f781 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e5e3a19-ae31-4779-a77e-fa615ea6dd22 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Format MUST follow : a n s w e r _ g e n e r a t e () -> str For example : so lu tio n = await self
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7e8f3ce8-6a70-498b-9671-0c4c8b1cde36 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37c5180b-f107-4b6e-b2c7-fd36eaa2b531 · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b0b2d1a0-ab32-4f7e-a8de-7e1c27894b5f · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8576080f-d72b-4221-b476-829f078754ea · outbound
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization "" This is a wor kf low graph
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0a41fbf5-200c-4f0b-a399-205d9a46a1b9 · inbound
FlowReasoner: Reinforcing Query-Level Meta-Agents ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c72814f9-1626-49eb-81e5-f5da80437d1b · inbound
HALO: Hierarchical Autonomous Logic-Oriented Orchestration for Multi-Agent LLM Systems ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e15cffa-39ec-400a-9577-83390127f903 · inbound
MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a721dc97-9037-4479-b644-6c9584fbabac · inbound
GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f3c0730e-c85b-4733-8c99-a39eff0612f1 · inbound
Rethinking Query Optimization for Multi-Agent Systems [Vision] ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff54a86d-7777-42ff-84a2-9d05f1a66d12 · inbound
Autogenesis: A Self-Evolving Agent Protocol ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b2c07b09-5da3-4bfd-8109-fc522d1443b7 · inbound
Autogenesis: A Self-Evolving Agent Protocol ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 879697a1-c5a3-4c0c-ab0c-6d9f2a50c1e1 · inbound
EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a8b83235-1d34-42e4-a096-2ae6ff7980ac · inbound
Towards Direct Evaluation of Harness Optimizers via Priority Ranking ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e287924a-aa6a-4bed-ba34-0e1a35ca3380 · inbound
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 75c452f2-a934-4964-995b-d6c0c695d327 · inbound
A Workflow-Aware Serving Layer for Agentic Applications ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af735b10-0f92-44ae-9c45-29658aba67bf · inbound
CONTRA: Red-Teaming Configurations of Personalizable Agents ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32dd730-36fd-4726-b31b-54bf42d8bf75 · inbound
Reward-Free Evolving Agents via Pairwise Validator ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689c9664-2d66-48fa-9e2b-80feda11857e · inbound
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.