Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:09:13.456124Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 148 outbound references and 12 inbound Pith citation observations for arXiv:2412.15838.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:09:13.456124Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:48.396266Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T01:32:22.464378Z
100 of 148 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1e81517b-58b7-46cb-bd5b-b1732ee3f832 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52a87ef-a321-4f93-845b-1826919ccf25 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MusicLM: Generating Music From Text
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac8d475-881c-409c-a107-3ae2a2ef5b20 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8555b2-c7c2-4670-ac18-d9c31f2016e3 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audio visual scene- aware dialog
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1503e659-c43e-4fd1-901d-6540d2303c1a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Flamingo: a visual language model for few-shot learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2663749-bb91-47aa-9c7e-0f672f63ae3b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c6d5649-96cb-442c-9911-4100989e731e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Claude 3
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de26c618-7e77-47b9-b8c4-c38592d6a3e8 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Pika art
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09007dbb-32d6-4376-bb9e-42bb9c44066c · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a851fe-83d9-46b3-90fc-ecc76891fbfb · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Constitutional AI: Harmlessness from AI Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a9a1cc-13cc-42ea-a12d-61609ab083bf · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training diffusion models with reinforce- ment learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31dba60-d8e6-4584-9d42-24d676fe3243 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Activitynet: A large-scale video benchmark for human activity understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a21545-2c5e-419e-929b-ef44887ff30c · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38661ef8-7c50-4580-8710-fa8e250bcd0e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vggsound: A large-scale audio-visual dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60f55d3b-1bb1-406d-8821-88539e848548 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c47207-6c7b-49ab-bc00-ebfba7a9d09e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247aed36-34ab-473b-9817-39513e3658e6 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e7e75a-7fdc-4bdb-a770-977b127b493a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vast: A vision- audio-subtitle-text omni-modality foundation model and dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed598795-6dd7-43e4-b9f9-770376e7ef9c · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ea859f-d959-4b56-935c-af0e1f21295d · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8336829-5615-42f6-9e6f-b014cd58f082 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mitigating Hallucination in Visual Language Models with Visual Supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d10d20-c643-4d77-9942-832649580bad · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f4a4f6-c6c2-44d8-a12f-ec79be421720 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a00611e-49c7-4d9a-9719-ca4cbf45eab8 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Qwen2-Audio Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7b9795-b534-44f9-b4b7-bf9001eb97e6 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6fc323-5bdd-4293-8318-25237c8ff182 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d122bb7-bdbb-4cb4-8b9c-1ce52a23bfa3 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Cogview2: Faster and better text-to-image generation via hierarchical transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5472f8e1-8f7e-45b1-8ce4-9313c5683e90 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback En- hancing chat language models by scaling high-quality in- structional conversations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3205791-f2ca-4b91-bbd0-d17997e0b8df · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback An im- age is worth 16x16 words: Transformers for image recog- nition at scale, 2021
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656d4e90-b8a3-4bac-8946-e8d20df43bf0 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback The Llama 3 Herd of Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5529be-a93c-4d83-835e-5ddc9c3608f0 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback KTO: Model Alignment as Prospect Theoretic Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690c0005-3455-4214-9ef2-dddab68c9748 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Stable Audio Open
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea0a9ce-51d8-465f-bc6c-460ce4fd60b6 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b536adab-cfed-43a9-bfe4-191f80c54efd · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f84202-71d9-425c-bf37-bd33c12ac764 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d74bcb-37a2-4a4c-ac6c-b893ed47c36f · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Make-a-scene: Scene- based text-to-image generation with human priors
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 148dd874-f385-44f2-b83d-6e3e7103bb79 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Imagebind: One embedding space to bind them all
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15418cca-6bf9-4945-ac03-c2774b6029c9 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Listen, Think, and Understand
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da612fd-fbc6-455b-907b-2e37c8a06fab · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Ego4d: Around the world in 3,000 hours of egocentric video
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8b9add-d6fd-45b3-8f30-5cdf73f7f249 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Denoising dif- fusion probabilistic models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eb6f4f-f26e-4607-acfa-6eb2ed2f802d · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Reference-free monolithic preference optimization with odds ratio
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c67a3f-5bac-44b7-95f1-92b36799e431 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Cogvideo: Large-scale pretraining for text-to- video generation via transformers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbef11dc-5629-4e4a-93d2-5a95896d7cd7 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tifa: Accurate and interpretable text-to-image faithfulness eval- uation with question answering
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539d44e8-d27c-4b28-bdc9-dc4624d101e9 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Movienet: A holistic dataset for movie under- standing
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58290bbe-8e3e-418e-a569-0aba224d9d9a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback The Platonic Representation Hypothesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ca3893-4af5-4b83-98ae-806bcfb036ec · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback AI Alignment: A Comprehensive Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22ca459-5270-4b9e-9092-fe28562c7324 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deb3d062-4a3e-4bb1-97cf-86a80fe9a10f · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FollowBench: A multi-level fine-grained constraints following benchmark for large language mod- els
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d42660-f8c6-4c10-84e8-6790d5ddcc56 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba9b290-341c-40d4-865e-ab39c8d08231 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audiocaps: Generating captions for audios in the wild
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6a9768-b1f8-412d-b194-6ddd321c003b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Pick-a-pic: An open dataset of user preferences for text-to-image generation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac391ca1-d780-4cbd-bb08-0d867c7e2f7c · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback black-forest-labs/flux (github reposi- tory)
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe258e0-7c4a-4e17-a62c-f9692fc2191a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Learning to answer questions in dynamic audio-visual scenarios
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb43062-79c5-4271-abdf-65ab53a644c2 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f814b7b1-38c7-4a4e-bb03-f837ea2f3347 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Vlfeedback: A large-scale ai feedback dataset for large vision-language models alignment
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c589e657-ea0b-4c5d-894e-81d0c94b7e3d · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Evaluating Object Hallucination in Large Vision-Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c37f3e56-2fc1-474b-bb12-d2cb118584c0 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a379f52-60ef-4813-ad2c-dc54a77c03a4 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Rich human feedback for text-to-image generation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8522f6eb-80c3-41a6-9fc1-56ff80d9efde · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0825298-a859-4483-9ea1-fb78a0f7f935 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0f45cd-d5aa-4da9-9497-a27f57a32f15 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Improved baselines with visual instruction tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb3f9848-a34b-4b67-a4a2-bc73d4aec5b4 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6faa12d4-3d54-4025-8229-7f2bd9c1f394 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf548e5-3e79-4a64-af9d-f71177351bbe · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Visual instruction tuning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99e0ec4-3dc0-40cd-ad29-17910e6568d4 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02eebe75-08c4-4422-804a-bf205189901c · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Holistic Evaluation for Interleaved Text-and-Image Generation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77701c6-a3f3-4c84-b42f-7a25a942e53b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ac0d80-c375-4c60-ab13-190d2e617a0e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Mm-safetybench: A benchmark for safety evaluation of multimodal large language models, 2024
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6eee69c-3885-4fef-b765-55b7dd97713e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Safety of Multimodal Large Language Models on Images and Texts
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3283ad6-f64b-4170-97df-cd6f4a8ae507 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MMBench: Is Your Multi-modal Model an All-around Player?
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b5f8743-0330-4cb7-82d6-23c254d038ce · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a574c2-2717-4d77-8499-d1f2f411986b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aa8b609-4f99-4c6e-8301-032bda97a982 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Deepart: Learning joint representations of visual arts
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3314e2-f5e2-4386-a6e1-bcdce444ca2b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language mul- timodal research
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48d06eb-8497-4b7b-899f-0363ca074d8b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1292cc2a-c380-4140-925c-ced7abc06225 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb41fb0f-944d-4070-b003-bbe5d7a13250 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Glide: Towards photore- alistic image generation and editing with text-guided dif- fusion models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f5c485-f92e-4fb4-a78c-7f68d539b65a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792f6d48-2799-4f2e-a42f-979557028780 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a393a2e-3def-474b-a41b-2ec1db2f8a59 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training lan- guage models to follow instructions with human feedback
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dfff67d-4011-47cc-a59d-13249dc8822f · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Instruction Tuning with GPT-4
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bcb3a0-cf8b-4a33-9f20-2a3bb13a0ed2 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Sdxl: Improving latent diffusion models for high-resolution image synthesis
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d53fff6-dafb-4234-aaef-8de70082acaf · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d7cdaf-7928-4f65-9456-fea6b68a93da · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Learn- ing transferable visual models from natural language super- vision
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c87cb7-6535-44b1-b4d7-8671b2ca5335 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Robust speech recognition via large-scale weak supervision
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f43c59d-282b-4b0c-9a2f-abdaf43a0c5b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Direct preference optimization: Your language model is secretly a reward model
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5f0552-d6e3-426c-966f-f76319de9f54 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7dda5a-2f61-4a5c-88b1-bec75b6caa8b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback High-resolution image synthesis with latent diffusion models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29faebb-d600-459a-8674-d798484c7c9a · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Training Language Models with Language Feedback at Scale
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a263e53-5154-44a7-b398-1ec6944f159d · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Laion-5b: An open large-scale dataset for train- ing next generation image-text models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4105e0-7127-4fed-bef7-18909a64e155 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback A simple baseline for audio-visual scene-aware dialog
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737cd126-fc78-41fd-8c23-7e54a4c9fff8 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback A-okvqa: A benchmark for visual question answering using world knowledge
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed0152b-4cbc-4f67-b121-ce3dd752567f · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Generative multimodal models are in-context learners
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb798616-9c1b-4179-a82c-1222d587289e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9428b8-20ce-4787-92b9-9c5499fde97b · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a7ac28-8fef-40eb-b771-76cf080dd083 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Stanford alpaca: An instruction-following llama model, 2023
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041491c7-0ee5-46b9-b6ac-f421147097c6 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e34cd5f-f90a-475d-be72-7d7970964ff3 · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ddf252-05cf-42e1-b1b8-1c695200ab9e · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Multimodal interaction: A review
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3894a934-24c2-44c2-8439-27ff9bc7f8ae · outbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d72886-ec01-4b9a-a0e3-aa02e463e154 · inbound
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68979fad-f720-4082-b6ee-53bd0998d496 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 284
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5960097a-5d40-4e62-b534-0e45e1f6e7d5 · inbound
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4fbbf22-fdf0-409d-b2e1-5dc8c0824026 · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41965ca6-c991-4de9-88d5-55ec184c603d · inbound
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 235f0a51-cc3f-4255-85f3-c40e292c7bf1 · inbound
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3313bbb2-a990-4a28-ab9c-bcf0eb042741 · inbound
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bc5119-0e18-4399-87e6-325f9abdd670 · inbound
A Survey on Training-free Alignment of Large Language Models Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a5b4f03-2e3d-46e9-b251-80b90ffcab23 · inbound
Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc88937f-7680-458d-8545-f77f7945806f · inbound
Identifying Topological Invariants of Non-Hermitian Systems via Domain-Adaptive Multimodal Model for Mathematics Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1ab7cce1-dd4d-47eb-b7be-dc48274b0e9f · inbound
Step-Level Preference Learning for Generative Agents in Social Simulations Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb96359-7d81-4f7c-a9ed-a2e1029b10b3 · inbound
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.