Pith. sign in

Paper Citation Record · LEDGER

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

As of 17 August 2026, this Paper Citation Record lists 100 of 148 outbound references and 12 inbound Pith citation observations for arXiv:2412.15838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15838 v2

Coverage vector

measured 100 of 148 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:09:13.456124Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:48.396266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T01:32:22.464378Z

Reference resolution

100 of 148 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e81517b-58b7-46cb-bd5b-b1732ee3f832 · outbound

This paper cites GPT-4 Technical Report.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.852944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.852944Z digest=sha256:05fc97aa06844907c6360918ceb4d7cc5ef986b8369be09f6c485023f3e1e6a7

Observation d52a87ef-a321-4f93-845b-1826919ccf25 · outbound

This paper cites MusicLM: Generating Music From Text.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.859946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.859946Z digest=sha256:f4e65e1f34824e57856fa92b9b9221e4953075c0a8321ab6482a7beb06d35a24

Observation bac8d475-881c-409c-a107-3ae2a2ef5b20 · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.868191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.868191Z digest=sha256:9a49a9221129ba9748c04e5445836fcf761fbd6447dc3f0818a389e587d73c99

Observation fd8555b2-c7c2-4670-ac18-d9c31f2016e3 · outbound

This paper cites Audio visual scene- aware dialog.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audio visual scene- aware dialog

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.874672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.874672Z digest=sha256:9fd700e052572930544ac8702edf94c0c47b9075b0cd4234be288cb03990b3a3

Observation 1503e659-c43e-4fd1-901d-6540d2303c1a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.881406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.881406Z digest=sha256:3f4aa36f5201296d7b35dee7e0bb94cf3687a7b4aae7ff1097cf0c1e99ea703c

Observation e2663749-bb91-47aa-9c7e-0f672f63ae3b · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.888377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.888377Z digest=sha256:b61baf0bcd946b1696fdedb825babf7f8f30374eee9b6308d56f9f9490bc517b

Observation 4c6d5649-96cb-442c-9911-4100989e731e · outbound

This paper cites Claude 3.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Claude 3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.895831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.895831Z digest=sha256:766848eb7fe2be5a45f93e06a89e23b18e071d41afa679d5a168e905cdd42e44

Observation de26c618-7e77-47b9-b8c4-c38592d6a3e8 · outbound

This paper cites Pika art.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Pika art

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.901531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.901531Z digest=sha256:708f7d6837622c47b16fe8408243939b7dbc4d532a1c340a18954558cc022771

Observation 09007dbb-32d6-4376-bb9e-42bb9c44066c · outbound

This paper cites Qwen Technical Report.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.908026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.908026Z digest=sha256:3cbbb8f8ddc69835d576aea3fc89655df2772d00435d5ee5c6ac042c18b72562

Observation 91a851fe-83d9-46b3-90fc-ecc76891fbfb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Constitutional AI: Harmlessness from AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.913726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.913726Z digest=sha256:b97ce6e143e6d13245297f30e5738173e4ceba599915712fdc69956e30c3e13b

Observation e9a9a1cc-13cc-42ea-a12d-61609ab083bf · outbound

This paper cites Training diffusion models with reinforce- ment learning.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training diffusion models with reinforce- ment learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.919821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.919821Z digest=sha256:57cdd5ac64392e6a972f7a14554e7a0e7251d7cbf91611c6527d5e20ffe9b8cc

Observation b31dba60-d8e6-4584-9d42-24d676fe3243 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Activitynet: A large-scale video benchmark for human activity understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.925690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.925690Z digest=sha256:ecd41d3fcfd3ec42f2f88ffc7c5e932180b384ab36aa6717dbb1d922a5108ce5

Observation c1a21545-2c5e-419e-929b-ef44887ff30c · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.931762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.931762Z digest=sha256:3b37f21b01ac1fd7f63e1476c51bae92be4a7b2ef9e2c27d7d0a330d9335b374

Observation 38661ef8-7c50-4580-8710-fa8e250bcd0e · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vggsound: A large-scale audio-visual dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.937011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.937011Z digest=sha256:fdb3990cf5994f21f09549264eb30bdd8d4bef9f05edf5c79d9d35b8f8795ef5

Observation 60f55d3b-1bb1-406d-8821-88539e848548 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.942401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.942401Z digest=sha256:f9c855bd1743a23cd5bc03f9271a6460c1e466f1082c0ba5869c255a6a687b6b

Observation 88c47207-6c7b-49ab-bc00-ebfba7a9d09e · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.947500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.947500Z digest=sha256:d178b554ab4c70ee316e213a3ef4f7085826ee183d028d7db86460cafad3385a

Observation 247aed36-34ab-473b-9817-39513e3658e6 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.953832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.953832Z digest=sha256:799f566922d950c2f08a10261a72541099289d2ac8ec186508c14a3b784ff274

Observation 45e7e75a-7fdc-4bdb-a770-977b127b493a · outbound

This paper cites Vast: A vision- audio-subtitle-text omni-modality foundation model and dataset.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vast: A vision- audio-subtitle-text omni-modality foundation model and dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.959129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.959129Z digest=sha256:c6790792f9194b237d2241d8130349c38c92b912a977d9fbee82eb55b4569ef3

Observation ed598795-6dd7-43e4-b9f9-770376e7ef9c · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.964448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.964448Z digest=sha256:8c494dd46ed6c7d0c0d2af13804d4dcf82e4a1a1c8f05b767c049d29fce9eb63

Observation d2ea859f-d959-4b56-935c-af0e1f21295d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.969123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.969123Z digest=sha256:9f59f0004c8481ac38affb676110218b087a95ba42a6cace94d5cc7912c45508

Observation f8336829-5615-42f6-9e6f-b014cd58f082 · outbound

This paper cites Mitigating Hallucination in Visual Language Models with Visual Supervision.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mitigating Hallucination in Visual Language Models with Visual Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.974082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.974082Z digest=sha256:7a188ca68f359e8177cfc824b0a648f7e4e999037d88101c40433db396d8094f

Observation d6d10d20-c643-4d77-9942-832649580bad · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.980482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.980482Z digest=sha256:c184a99eae384a75435537ad76ae23a98d1d22f7be100ff10c6f346de751a9a6

Observation 07f4a4f6-c6c2-44d8-a12f-ec79be421720 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.986477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.986477Z digest=sha256:b14ac0c18d67e5a3c05710ef359af6d94e9d6783652553dd7674166cab1a7eeb

Observation 1a00611e-49c7-4d9a-9719-ca4cbf45eab8 · outbound

This paper cites Qwen2-Audio Technical Report.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen2-Audio Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.998854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.998854Z digest=sha256:e4f62a8c98423c2f32ebc27f986d4670197bd7e5fc43b1d669eb0236551ae4c6

Observation fa7b9795-b534-44f9-b4b7-bf9001eb97e6 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.005404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.005404Z digest=sha256:79f920491174a9f87460ac9d9e611f0b0e4b62505460688b9a27f3732367cb02

Observation 1e6fc323-5bdd-4293-8318-25237c8ff182 · outbound

This paper cites SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.011046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.011046Z digest=sha256:28d93c991ea7e1bfe5c4b7b3242c45dcf46b6a025a919377ad00bb18a61a424a

Observation 0d122bb7-bdbb-4cb4-8b9c-1ce52a23bfa3 · outbound

This paper cites Cogview2: Faster and better text-to-image generation via hierarchical transformers.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Cogview2: Faster and better text-to-image generation via hierarchical transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.017269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.017269Z digest=sha256:7e1e60dcae0f747cc72b0c141f3965cedf3deee0123ba0685daac543f424332c

Observation 5472f8e1-8f7e-45b1-8ce4-9313c5683e90 · outbound

This paper cites En- hancing chat language models by scaling high-quality in- structional conversations.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback En- hancing chat language models by scaling high-quality in- structional conversations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.023864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.023864Z digest=sha256:f1f480dfad62cb1bcce108f08b3cdaf5db15e61783e0680f733786272ea939b7

Observation f3205791-f2ca-4b91-bbd0-d17997e0b8df · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale, 2021.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback An im- age is worth 16x16 words: Transformers for image recog- nition at scale, 2021

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.030121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.030121Z digest=sha256:38091886d4b18d0a74617878fae063a57230cd973e28c4756f97edc3bb4903e7

Observation 656d4e90-b8a3-4bac-8946-e8d20df43bf0 · outbound

This paper cites The Llama 3 Herd of Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.037132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.037132Z digest=sha256:1bc112e874e1d6799be4d9c89e5a434cd638057d7295c963c3079a5518416a95

Observation 2a5529be-a93c-4d83-835e-5ddc9c3608f0 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback KTO: Model Alignment as Prospect Theoretic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.045914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.045914Z digest=sha256:343da0ef2584ae0037b1ad7763f94106e3c76914a0fefb8fad7876f3723b4013

Observation 690c0005-3455-4214-9ef2-dddab68c9748 · outbound

This paper cites Stable Audio Open.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Stable Audio Open

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.052993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.052993Z digest=sha256:780cad3ffd356ef1fbdcebbfba38e6f21ce06a4513b352c2d7fc2e97e8136249

Observation 4ea0a9ce-51d8-465f-bc6c-460ce4fd60b6 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.058336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.058336Z digest=sha256:932cc7870094c2f80869e7523d233f0d2c945f2caf9532781b9295f3572eb2c1

Observation b536adab-cfed-43a9-bfe4-191f80c54efd · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.065334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.065334Z digest=sha256:0300f848ba4347716b3dc231f5ea7d5b2e199748b913c0430ced5a2edb4a2d04

Observation 85f84202-71d9-425c-bf37-bd33c12ac764 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.071001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.071001Z digest=sha256:68d50a1bc12080becc5dcd25791f3e934e9f54f0d1a8d82c1aca10d09400fb41

Observation 63d74bcb-37a2-4a4c-ac6c-b893ed47c36f · outbound

This paper cites Make-a-scene: Scene- based text-to-image generation with human priors.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Make-a-scene: Scene- based text-to-image generation with human priors

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.076897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.076897Z digest=sha256:e724e97899adc7213a6eeac5d35000d6ecc9eb69a69c0567ab07c10e2b940f9d

Observation 148dd874-f385-44f2-b83d-6e3e7103bb79 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Imagebind: One embedding space to bind them all

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.083148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.083148Z digest=sha256:121c11e7e4c2e7bdbbf4664a98445bece2c2149d0c9b3f8b8f2bfc9f0a066da5

Observation 15418cca-6bf9-4945-ac03-c2774b6029c9 · outbound

This paper cites Listen, Think, and Understand.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Listen, Think, and Understand

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.089824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.089824Z digest=sha256:7365540051d45577be01055a1b23499320d6fa7b48f993eb5efe71e2a7e49dab

Observation 6da612fd-fbc6-455b-907b-2e37c8a06fab · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Ego4d: Around the world in 3,000 hours of egocentric video

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.098017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.098017Z digest=sha256:a5f657a3fbd1793b16b33abfc4ea83ab5fa73d8c9876493122618eebdc0e2fcb

Observation 0c8b9add-d6fd-45b3-8f30-5cdf73f7f249 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Denoising dif- fusion probabilistic models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.104130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.104130Z digest=sha256:7aad844199edac4461883ae04b345525498e40629b1cfb1ea0bd0cbdb701623d

Observation d2eb6f4f-f26e-4607-acfa-6eb2ed2f802d · outbound

This paper cites Reference-free monolithic preference optimization with odds ratio.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Reference-free monolithic preference optimization with odds ratio

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.110707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.110707Z digest=sha256:f6eda2e40778c26a6250a4649ac7e7767492548c7102ff4837d3852661174797

Observation a7c67a3f-5bac-44b7-95f1-92b36799e431 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.117135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.117135Z digest=sha256:c380b46b3f0408948f68733fc7d4de91099c944529bb7dbe3ea2397d0958c571

Observation dbef11dc-5629-4e4a-93d2-5a95896d7cd7 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness eval- uation with question answering.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tifa: Accurate and interpretable text-to-image faithfulness eval- uation with question answering

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.124856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.124856Z digest=sha256:f6035997a1129ab69fe8aa2162b3de937c6a8ebbd02fba726ae3c357de7545f8

Observation 539d44e8-d27c-4b28-bdc9-dc4624d101e9 · outbound

This paper cites Movienet: A holistic dataset for movie under- standing.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Movienet: A holistic dataset for movie under- standing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.132283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.132283Z digest=sha256:f25100459d260186f14194e0d3c136fd728e1c3da38b61093637f18c51d117ba

Observation 58290bbe-8e3e-418e-a569-0aba224d9d9a · outbound

This paper cites The Platonic Representation Hypothesis.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback The Platonic Representation Hypothesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.138523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.138523Z digest=sha256:573e8bd6a0fc9a6c56ae04823b7f435de5029406b1571e41b44f2434170a3071

Observation f2ca3893-4af5-4b83-98ae-806bcfb036ec · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback AI Alignment: A Comprehensive Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.145058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.145058Z digest=sha256:4d1f689708b70970321ea97a83a5b39cb521e22c6e26b5c9e127577041a9ca5e

Observation b22ca459-5270-4b9e-9092-fe28562c7324 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.150282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.150282Z digest=sha256:a2784e216e88dbb25addac718e9c02e5da4fb10561bfd0a228ecf88edcbc9d8c

Observation deb3d062-4a3e-4bb1-97cf-86a80fe9a10f · outbound

This paper cites FollowBench: A multi-level fine-grained constraints following benchmark for large language mod- els.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FollowBench: A multi-level fine-grained constraints following benchmark for large language mod- els

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.155900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.155900Z digest=sha256:7bef5fab1e912c20abcc5634f831f97eb9c7f6894ced57afca60b4dbe0b78f2f

Observation 54d42660-f8c6-4c10-84e8-6790d5ddcc56 · outbound

This paper cites FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.161441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.161441Z digest=sha256:ea731feaed199747d495427fa4dd1c6b7bd8e9ffbba870c1faf4d0afd02348f6

Observation 6ba9b290-341c-40d4-865e-ab39c8d08231 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audiocaps: Generating captions for audios in the wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.168320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.168320Z digest=sha256:62f934e7ff32b69d44df3c3e3a8133f8bcee7aa60770f0678bcc24e0b098f08c

Observation ab6a9768-b1f8-412d-b194-6ddd321c003b · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.174808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.174808Z digest=sha256:0b0c5f27d4291364fa3d4a4975c8ec879d86f2d6f3806ac4faac8b5701584c0e

Observation ac391ca1-d780-4cbd-bb08-0d867c7e2f7c · outbound

This paper cites black-forest-labs/flux (github reposi- tory).

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback black-forest-labs/flux (github reposi- tory)

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.180171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.180171Z digest=sha256:4c8e0b3fdbfa213492b778198abd723cb230974dbae847c171702c5aebda0302

Observation fbe258e0-7c4a-4e17-a62c-f9692fc2191a · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Learning to answer questions in dynamic audio-visual scenarios

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.185870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.185870Z digest=sha256:d4c2e69a40b7a3849699d820baf19f0c570ec88c78463ecc9b5fb0ca2072166a

Observation dfb43062-79c5-4271-abdf-65ab53a644c2 · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.191588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.191588Z digest=sha256:3ed098761ff86933ddc475a856cf2e98c17453523e479688cbe76b5045e59f45

Observation f814b7b1-38c7-4a4e-bb03-f837ea2f3347 · outbound

This paper cites Vlfeedback: A large-scale ai feedback dataset for large vision-language models alignment.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vlfeedback: A large-scale ai feedback dataset for large vision-language models alignment

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.196963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.196963Z digest=sha256:0156d1ddb4fc1a1ecf98c060f068e42cdc38bbc7c057829b01c05ed816441620

Observation c589e657-ea0b-4c5d-894e-81d0c94b7e3d · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Evaluating Object Hallucination in Large Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.202147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.202147Z digest=sha256:242f5dce909bf67f6ad353b1d3d2defe78defccc703b8535554179036afcda61

Observation c37f3e56-2fc1-474b-bb12-d2cb118584c0 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.207320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.207320Z digest=sha256:f476b385b0a50339e02b38d782bafd7e321496cf1f70d1d4b1cb6760c9ef68fa

Observation 0a379f52-60ef-4813-ad2c-dc54a77c03a4 · outbound

This paper cites Rich human feedback for text-to-image generation.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Rich human feedback for text-to-image generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.212418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.212418Z digest=sha256:90e188b2168ca1bcd43d5a4bffb3951cb7f8dc75f2bada280b5bec90c0070d08

Observation 8522f6eb-80c3-41a6-9fc1-56ff80d9efde · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.217467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.217467Z digest=sha256:c2b82c1be0d5d450be9577848d6923f482601a30d1bc51186a8698f0cb7e1922

Observation d0825298-a859-4483-9ea1-fb78a0f7f935 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.223340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.223340Z digest=sha256:8b90805296def49ae28a35491e0bd36bc5520357e260d6572af3c54d8719f982

Observation 0c0f45cd-d5aa-4da9-9497-a27f57a32f15 · outbound

This paper cites Improved baselines with visual instruction tuning.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Improved baselines with visual instruction tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.229027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.229027Z digest=sha256:6a2339ab3e108f36e097852b157752c05108c8e19cffb6634eb975dbb4f463ff

Observation bb3f9848-a34b-4b67-a4a2-bc73d4aec5b4 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.234393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.234393Z digest=sha256:17ed3ec34efb4a5102deeed500f49c8a492b4ca8d92f6a5a9d4b813266d998b5

Observation 6faa12d4-3d54-4025-8229-7f2bd9c1f394 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.239542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.239542Z digest=sha256:df08154dca2e93a5ea94c289c657ef25c46096b7ee47aa70d07c55b19bdb5492

Observation adf548e5-3e79-4a64-af9d-f71177351bbe · outbound

This paper cites Visual instruction tuning.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Visual instruction tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.245176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.245176Z digest=sha256:8fda9cd13cf599a69efe1c5d6feba5c34deccda2b9e2ebcc3a1fa5de23b2286b

Observation e99e0ec4-3dc0-40cd-ad29-17910e6568d4 · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.250765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.250765Z digest=sha256:bb035b489417eb9000f7f5b6fa32bac70d3de3e5361f4e727587601a1753d9b1

Observation 02eebe75-08c4-4422-804a-bf205189901c · outbound

This paper cites Holistic Evaluation for Interleaved Text-and-Image Generation.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Holistic Evaluation for Interleaved Text-and-Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.256793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.256793Z digest=sha256:1a1533d699f510d5fbcf5722f022114e6210f469b18a21623ad41c7c72ec9457

Observation a77701c6-a3f3-4c84-b42f-7a25a942e53b · outbound

This paper cites FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.263881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.263881Z digest=sha256:3b9f2d01748be7650f0f0dffc28dd9e26ec9243e962fd0bd5f19a876292e7683

Observation 15ac0d80-c375-4c60-ab13-190d2e617a0e · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models, 2024.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mm-safetybench: A benchmark for safety evaluation of multimodal large language models, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.269400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.269400Z digest=sha256:6c28627f76442ae17ed92ae7262917c97c5dc3c867cfe87cdb6610100f303e72

Observation c6eee69c-3885-4fef-b765-55b7dd97713e · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Texts.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Safety of Multimodal Large Language Models on Images and Texts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.275081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.275081Z digest=sha256:8c680340abf006c189d647a2e43510cb24fd7fe9ec0c956fdac020d44ad4b266

Observation d3283ad6-f64b-4170-97df-cd6f4a8ae507 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MMBench: Is Your Multi-modal Model an All-around Player?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.281735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.281735Z digest=sha256:c0f84b3409bf9bf60e4ff0150af30bb7c1ee8b1ab6d74893a012802883bbd0ec

Observation 6b5f8743-0330-4cb7-82d6-23c254d038ce · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.286960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.286960Z digest=sha256:1df75ebd5a829ea172012be3c195fdee1b0b74c3f1ed2be9b4fcb6fc16361b4b

Observation 38a574c2-2717-4d77-8499-d1f2f411986b · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.292499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.292499Z digest=sha256:8bc339fbe1ca8709575932caf2b151a003c42c227f2d78531e93cd7feea2b607

Observation 7aa8b609-4f99-4c6e-8301-032bda97a982 · outbound

This paper cites Deepart: Learning joint representations of visual arts.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Deepart: Learning joint representations of visual arts

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.298764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.298764Z digest=sha256:c93f75c7e978f692d5f67864c9a788e9d6f749f0f30992435ad51d040bdc5b51

Observation 1f3314e2-f5e2-4386-a6e1-bcdce444ca2b · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language mul- timodal research.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language mul- timodal research

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.304095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.304095Z digest=sha256:b46ee5d62e53edf75faaad6cf77802a9af1337ca25c893c4037560af71c321fc

Observation a48d06eb-8497-4b7b-899f-0363ca074d8b · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.309116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.309116Z digest=sha256:a7860099ce85134ee707615711e073592686a7cd1f67dd6183be24ab2d97ebcd

Observation 1292cc2a-c380-4140-925c-ced7abc06225 · outbound

This paper cites an unresolved cited work.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.316332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.316332Z digest=sha256:1d963dff3f9d582d21ac2aaeee8ae1a526cc9203f826a3f923cdc4463655bdfb

Observation bb41fb0f-944d-4070-b003-bbe5d7a13250 · outbound

This paper cites Glide: Towards photore- alistic image generation and editing with text-guided dif- fusion models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Glide: Towards photore- alistic image generation and editing with text-guided dif- fusion models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.322159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.322159Z digest=sha256:4f95b1bb1f9e3de78105f8ea79ce704353a445ca25a7b31844134f44ff46b2a3

Observation 65f5c485-f92e-4fb4-a78c-7f68d539b65a · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.328262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.328262Z digest=sha256:675012bf56e7eb3abb4c58a05e31a43a516ae6c34db73ba351b0ed1586de291e

Observation 792f6d48-2799-4f2e-a42f-979557028780 · outbound

This paper cites an unresolved cited work.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.333596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.333596Z digest=sha256:4c54b93761394a89f0aea7c911971ddbcfc64125f2a525f497504a2f45275a47

Observation 6a393a2e-3def-474b-a41b-2ec1db2f8a59 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training lan- guage models to follow instructions with human feedback

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.338938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.338938Z digest=sha256:aa86dd65baca697de22a96cdae65992dc848b85b4a4ccb0511b324d7af3a5cad

Observation 9dfff67d-4011-47cc-a59d-13249dc8822f · outbound

This paper cites Instruction Tuning with GPT-4.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Instruction Tuning with GPT-4

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.344688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.344688Z digest=sha256:dc45edeea08d37f3f3bde281cd14f32334e6df254c5d983b7c9816529c541c99

Observation 77bcb3a0-cf8b-4a33-9f20-2a3bb13a0ed2 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.350863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.350863Z digest=sha256:ec795698cb2e8d47271b64fffb3f01cec5f44c7d15ce41978175f8dc3d6b5ca3

Observation 8d53fff6-dafb-4234-aaef-8de70082acaf · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.356520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.356520Z digest=sha256:e0d851f536b578416b9eb658509a1f7aec64327af720d8cad3a4d6181b0578d8

Observation b8d7cdaf-7928-4f65-9456-fea6b68a93da · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Learn- ing transferable visual models from natural language super- vision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.361872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.361872Z digest=sha256:6fac514d40b61151f59a8e45401e07076a15ce53c7b8bd3f72b6d40f87f494d2

Observation d6c87cb7-6535-44b1-b4d7-8671b2ca5335 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Robust speech recognition via large-scale weak supervision

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.367541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.367541Z digest=sha256:43f82fae8499734b5609ace88afc121efbc3eeec5a3f12e27ce20226749b334e

Observation 2f43c59d-282b-4b0c-9a2f-abdaf43a0c5b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Direct preference optimization: Your language model is secretly a reward model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.373748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.373748Z digest=sha256:64013ce72205181925020d348d4a23cf0ec4cc0a05fa8cc5aab204e1285ba156

Observation 4f5f0552-d6e3-426c-966f-f76319de9f54 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.379593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.379593Z digest=sha256:383b80c980631d6c518d3e31884cba5b0cd02f5de4b59dbf9e6fbbf773f2d29f

Observation 0b7dda5a-2f61-4a5c-88b1-bec75b6caa8b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback High-resolution image synthesis with latent diffusion models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.385630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.385630Z digest=sha256:c7506cd70b17712bae89532e119143b90cc0da7222311f57c7ee514901cb5fc6

Observation a29faebb-d600-459a-8674-d798484c7c9a · outbound

This paper cites Training Language Models with Language Feedback at Scale.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training Language Models with Language Feedback at Scale

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.391447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.391447Z digest=sha256:77cde7a95b673d4710db2411e7edf785c51d714f469792aa5d23f8e024a2c3e9

Observation 3a263e53-5154-44a7-b398-1ec6944f159d · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.398296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.398296Z digest=sha256:519900c70c23c34116dec592f3bb4235267fa08b76ba12fc1e8191aa3da3abb6

Observation 2b4105e0-7127-4fed-bef7-18909a64e155 · outbound

This paper cites A simple baseline for audio-visual scene-aware dialog.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback A simple baseline for audio-visual scene-aware dialog

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.403926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.403926Z digest=sha256:73fe6e58e8238e065d7af3e3c49f5582f00ff72512f6eca536a5c4af1cbba139

Observation 737cd126-fc78-41fd-8c23-7e54a4c9fff8 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback A-okvqa: A benchmark for visual question answering using world knowledge

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.409686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.409686Z digest=sha256:901b2f8dbf2e39831d8cd7e281824a7496e022ff7c5c1cdb0fe01a56946c7b3a

Observation bed0152b-4cbc-4f67-b121-ce3dd752567f · outbound

This paper cites Generative multimodal models are in-context learners.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Generative multimodal models are in-context learners

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.415895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.415895Z digest=sha256:04082a8e08c39cf14f31d1b3051595b4a046c290dfad14bc48f3bcac0b93c1b2

Observation cb798616-9c1b-4179-a82c-1222d587289e · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.422023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.422023Z digest=sha256:15e53e688645d3409e2f4f9dcccaa10445ba13f717fecf10c3e3de93ce6db041

Observation 5a9428b8-20ce-4787-92b9-9c5499fde97b · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.427621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.427621Z digest=sha256:776b2d2977e04d6ae2d6d5156ebf1728785d9ae7b7646b594d3af4193fac4670

Observation 99a7ac28-8fef-40eb-b771-76cf080dd083 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Stanford alpaca: An instruction-following llama model, 2023

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.433729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.433729Z digest=sha256:22e135e60aef3bc9feaf6fc2d491cff5d8b17c3effabb1e9aa1571ca24e5a5fd

Observation 041491c7-0ee5-46b9-b6ac-f421147097c6 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.439082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.439082Z digest=sha256:91151500ab447b8da18b061a0cde191dae845a9ea4779fd8985a4c8475f2ff10

Observation 2e34cd5f-f90a-475d-be72-7d7970964ff3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.444820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.444820Z digest=sha256:67233cdd2ad6121d7fa3858fa304bcf962f6e95481599899b80cb2bed9ceb91c

Observation 68ddf252-05cf-42e1-b1b8-1c695200ab9e · outbound

This paper cites Multimodal interaction: A review.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Multimodal interaction: A review

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.450234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.450234Z digest=sha256:95f7da0ce93384a66a65c867b7ef8d6bd295dadb5087ead082844631e6180f38

Observation 3894a934-24c2-44c2-8439-27ff9bc7f8ae · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.456124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.456124Z digest=sha256:6ee44d0b7b287a97a951faf0dd406c4ac24cc59ee4616b1496614dc515a32c9a

Pith citing papers

Observation 08d72886-ec01-4b9a-a0e3-aa02e463e154 · inbound

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction cites this paper.

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:48.396266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:17:48.396266Z digest=sha256:5d1981f2c0ad40c3872d1824c13b5fb19ba62639edb01e865bbbb770771e2e61

Observation 68979fad-f720-4082-b6ee-53bd0998d496 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 284

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.497327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:84151eb10eb567dc9f8d35c833fff5072c6991f1300b3c2857e96b9db31bd97a

Observation 5960097a-5d40-4e62-b534-0e45e1f6e7d5 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.467784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:e655cb17ebbe2689147e98850dcf850f3168cf02959fa0cce430de79e78e1461

Observation e4fbbf22-fdf0-409d-b2e1-5dc8c0824026 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.554751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.554751Z digest=sha256:34fe6fc0b85d037cc568cab3986e0f9d97e6da8a4948b6abfa37c7487bdf9928

Observation 41965ca6-c991-4de9-88d5-55ec184c603d · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:27.009844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:27.009844Z digest=sha256:e36a9553c8b03ae0a343f6feb27250451433b9493a368b057ccdbf468259a595

Observation 235f0a51-cc3f-4255-85f3-c40e292c7bf1 · inbound

From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary cites this paper.

From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:22.423350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:22.423350Z digest=sha256:29f26cc04cc0f569faa4eb600dde93b4d2c15a7895acb7e0a54e4adf3f925b15

Observation 3313bbb2-a990-4a28-ab9c-bcf0eb042741 · inbound

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong cites this paper.

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:16.405895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:16.405895Z digest=sha256:4b0f3e4d5c0b1a1183866be444b62b28bf3ba518e12568929b6f0289dbe12d11

Observation 98bc5119-0e18-4399-87e6-325f9abdd670 · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:42.053686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:42.053686Z digest=sha256:2acc5307c92ae90b13f86d699eaf157de1bb3411ea50bedef5f3f91c848bb28a

Observation 3a5b4f03-2e3d-46e9-b251-80b90ffcab23 · inbound

Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics cites this paper.

Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:55.128834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:16:23.776505Z digest=sha256:240574f1758d247a0f291f78d41229d43cf1a72a5c14d66c69db17cde0311d61

Observation fc88937f-7680-458d-8545-f77f7945806f · inbound

Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics cites this paper.

Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:20:00.192149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T10:16:00.039638Z digest=sha256:1fe55bd9bfdb3a31481a7964cea36dc84ea8d4dffc5674a9618d2f6759e77581

Observation 1ab7cce1-dd4d-47eb-b7be-dc48274b0e9f · inbound

Step-Level Preference Learning for Generative Agents in Social Simulations cites this paper.

Step-Level Preference Learning for Generative Agents in Social Simulations Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T02:00:25.588376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:00:25.588376Z digest=sha256:0f228335ead7b65d2d3a5936e08f0c712ea7462813c542e60842dff57ffc3ecb

Observation ebb96359-7d81-4f7c-a9ed-a2e1029b10b3 · inbound

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers cites this paper.

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:14.295305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:14.295305Z digest=sha256:500335c2eb8f1ea78ca20761cd7650225fa32700ddd9bed70b2dfd1c3221346d