Pith. sign in

Paper Citation Record · LEDGER

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

As of 4 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 100 inbound Pith citation observations for arXiv:2503.01743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01743 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T22:22:27.455361Z

measured 158 of 158 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 147 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:04:10.561277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:37:34.431677Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact45
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6b54e1f-1742-4df7-8671-a420a9af934a · outbound

This paper cites Phi-4 Technical Report.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Phi-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.827502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:e51a1b7feb1ba9e05c34dea46c130f203df2ea74828b84cbf6556523a96387f1

Observation cc4a8c2d-201b-49c4-b918-a9f9822e660b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.598919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:6a7b9fa0a01e071f74519334fb507f553bda8cf4eac19e26b3c14fd279853da4

Observation a367daaf-6138-4906-b78f-8600010187ec · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.618918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:b11f269cd1ae61a52288d63f371e211982306d2a26dea3c85e2182b27d4ffb3f

Observation ac13f7c7-ac1a-458a-b22a-cb0403bec366 · outbound

This paper cites Program Synthesis with Large Language Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Program Synthesis with Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.629998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:c55ea64e2e22753118779e3a0b11e8c0116bb5aee32fa35a0245d05cf1f62a73

Observation 1e0fe995-2dc4-4c5f-9b30-d352767e45fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.640594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7107dd5bd2fa6ecc0ec3e63418a8592f5be75d41c5e31740ba84975475b06383

Observation 5a1425e3-4730-43e6-9e81-1244100eafd3 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:27.658413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:050ae29ae075d3d000113a3418bc9c689783de698a1d37e6b1faaa270e0d1b8c

Observation 853c25c5-6def-460c-9651-7b1dcaa73068 · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:55:18.098549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:eeced2c612d09c7c9dfe65ffa833ab9e4282416a01d0343f7f3ced2f59bfd2d1

Observation 86aae764-7ee3-4a0c-8152-b5a2ce7b2f43 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.684224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:289ad2585e787fdf3497013e1726e0c426acdd5b6e64a6cf8537f73c1ec140e6

Observation 1faa8fbc-8fc9-45c1-bfb4-71cbdf4a149f · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.195230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7ad0afd654342485dcc203fa7d9439021d7fc22ec89b302011a7a9b4a2a964bf

Observation 0a0a9faf-59aa-49ac-9657-e0cf7d262a49 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.211121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:0a5805a3412ba004c857fce1c7b60ffaf3b28cb91e9e9db3e3bfaaee5d4174ed

Observation 3a059973-2a2b-4636-9b9f-9ab0c3887ebe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.842608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:6b0c959afac65dcea2c01e02ce35e1d546ae84c558cfecb88af8ec96d12f07e8

Observation c7a11588-f703-4351-9c61-23e4cc1b9ad3 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.653229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:d0680a2d28a21581ac8a7c1bab40e1931b9994100139b3a52b978bef6d73fe4e

Observation 0be5604f-6156-400f-90c1-811155d64806 · outbound

This paper cites Qwen2-Audio Technical Report.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Qwen2-Audio Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.887359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:dc0249a05336fc7a4d49e8472425d31fb4680bfc4809171c80b4bbd7c4f3df40

Observation 951b5426-f4ed-4f0a-983e-cc0efc7b2ecd · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:e4aef6cd21c93c71787d12cce6cb481b30db6fe8334bb5ade37a74227b7a6b0e

Observation a8adfca7-6afb-47be-9979-59c4881f0f9c · outbound

This paper cites The Llama 3 Herd of Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs The Llama 3 Herd of Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.929950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:54d500597eb49a135e8b4b22a2d5d7bdef091d8b3391cb729a87cf25e8f90935

Observation ac6a79fd-1b7d-415d-ac27-11b631a7e9c9 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs NVLM: Open Frontier-Class Multimodal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.942674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:1614d059045adb36ef9e9385e073f654b63fa3e810c95e8557679699ed57afc4

Observation cfc165d0-00e0-4861-a4cc-ae115cef8770 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:28.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:218ddf13e818d6506e74b0797f875cd73f5896d43f0e4f38aa0938ba27023982

Observation 2e39c14c-8d32-4b9b-9453-ec93b0e661ce · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.974405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:a8b7a8db2a97830dbaddc10dd999a6cc5d996c2fb5ee30b03449d80206a7fe14

Observation cb91d072-e5e7-429c-b9d4-f256aecc44af · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:0181c2ce2b24809a546835b08b603a8c3a1efaafc8804b3b45c02506a4fb1f03

Observation 073f4210-4bfd-4aab-ad56-feccfb6dfc50 · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Audiochatllama: Towards general-purpose speech abilities for llms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.229639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:d01f604020dcaf399baad3de3dfc00348ef3da915ffef84bf1a19e995bfe4fff

Observation 48cb63da-7dd0-4086-87bd-a5f0a67c8f9d · outbound

This paper cites Joint audio and speech understanding.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Joint audio and speech understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.242536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:6fa641796f15f8cae34e60d1d2af26bb8819f5b79865a5a73bfb66acdc8f0d17

Observation 538d58b4-7391-4752-af4f-70ac42e45fa0 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Conformer: Convolution-augmented transformer for speech recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.248124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:2e9fe7f82ced9d6a5baed39f8eb0b6f2ba624ea3b69f04b52631c8d1e17cc579

Observation 603a0f70-7dc0-45ef-acb9-4082d8366ef4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.002945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:1de232bcba38c0e1ad1eaadbe68dad65369aad03a740dd10d3430b886c7c982e

Observation 357f1c9f-f3fa-4729-a656-9f00bb096e5e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Measuring Massive Multitask Language Understanding

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.020311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T01:08:06.256034+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:899259bf07d5779077b12a32c6c2bc357bf523191db4daf2852bf9a2476a0c00

Observation 0a8e7430-2547-4ca4-a50f-a1ec18d00da1 · outbound

This paper cites GPT-4o System Card.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs GPT-4o System Card

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.028354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:4a6e020b82ec52f4aca6ce9519cce87e9604b9d6e8877b7ab58e6e1fcf20c924

Observation 70a8d384-67ab-46f0-b5ec-c1d024c3c516 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint arXiv:2411.04997.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint arXiv:2411.04997

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.041542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:40cb9710d50d81426172e53321d3e001c623859c2778c144fbb09a025a23cdb4

Observation 835fd8aa-9747-47af-8739-3f54d9f683b5 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.052476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:5929e0b3dcf2ec591133323e83c97f95a355d9714bfbd9c8a430ef8070c0acab

Observation 462faadc-9f2d-4962-9074-c05666608f0e · outbound

This paper cites [LBX+24] Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs [LBX+24] Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.297769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7b982f6d3790fb5b0dfb4b1b3d62efe9fd3c7492d45ee749b9da06830b7b3cf6

Observation 5c790fe7-78b1-4371-99d8-da30fb8126cc · outbound

This paper cites Let's Verify Step by Step.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Let's Verify Step by Step

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.062136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:e866152431b57f7de4cf6b83cb853c3693750f099aa39b744bd6fde85b65a9e6

Observation aab323ed-6852-4e11-8995-334ee85bf04c · outbound

This paper cites Red Teaming Visual Language Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Red Teaming Visual Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.068521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:8f63e4180306cd6fc80ce1f60a82607ca7bca12817c8a7f3f3c0e9d2d198ff37

Observation fd5de2bb-c7be-4f87-bf34-5a3b69153488 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:04:18.346890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:b1ff364c5bc680f8884ee8de9dd2d4d3c8b3c035b0b046874f8ef9b9afb8117e

Observation e19f28f0-ed57-49c5-90b5-9df326138136 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.089575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:84691c495ceaa58a1c18d12e88bb717a84f3b706c3545123c07dc59cdae8b7ee

Observation 51e28001-376d-4aa2-9742-47d02087c54e · outbound

This paper cites American invitational mathematics examination–aime.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs American invitational mathematics examination–aime

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.291545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:dfcca943bce1e4add98bb780a2e0c780b9e81ab752723356acebc270a4a13a0c

Observation 5d366aa7-3bfc-495d-ba6d-dad78571317d · outbound

This paper cites Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.098692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:560ec52494099a74ce26861269fd65aec936e71175810e67b5c1a7d9a97d9e3a

Observation 682f4fef-587f-4785-ae4c-2914a410cca8 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.265591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7f2810d9bf6d11ceba7fa302aa5d4d293c0d61c0f8334c82cba7aa4a06372d4c

Observation 1c10b859-8f3a-4ad8-9ab6-d80841c52553 · outbound

This paper cites s1: Simple test-time scaling.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs s1: Simple test-time scaling

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:22:28.105516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7a83fabb4538e75f56b747bbbcba139f8b94826e3a18e5becace4995bccfded4

Observation d8b6ce4a-1bdd-4baf-aefd-d430af5807b6 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Robust speech recognition via large-scale weak supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.281517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:cd9ff10e071b51d880eb710543c1cffa5a7be6a822d9c0d0eeb3219950b5b17c

Observation 04e502c6-18df-4333-9084-149b01fb5adc · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:20:15.117976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:70cc2d66dc88ae35181d84a6b69f499ab157335cda43e82a912a66b34b8fcb38

Observation 77834b61-efb2-47c0-b102-4e45ffcf2f31 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:22:26.192007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:520f25dd8d3c6b40dac529936c52ae66a9855d7f83f7d31fddbe06849f1edc9e

Observation 2ad8cbf1-a96d-4bee-a255-1e061f6aeede · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.139480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:97df71f288dfd0aef02e105b1cbe4fca5fa1579684f90e825aac01ed4819d56d

Observation df1f6ee2-b138-4566-a6bc-dfb3f59c20d5 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:04:55.589197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:1c88e3b3ac07f686eb75b00a9487cf99834ea1110dce11aa364d28d1ad0319e5

Observation 3823b17c-5a9d-4351-b21e-f6d63a1eda20 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.160647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:cda9bb0409956f9d4975780f2124b6ba3ffa054660da8a3302c00847269d0544

Observation b8d9ed14-0905-4955-96cf-3b85e68f5226 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.171212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:95a81dc46d99aaa09e17cbe750b28ef5d5f8235ce59a662a68152d8026aa2c9c

Observation 9ac5b64d-3da6-4e3e-8bc4-74ece11226b0 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Gemma 2: Improving Open Language Models at a Practical Size

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.180304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:d0aa2d38ad0334859df1b742055770d17b1bb46b10ff47dd26acb97570441d33

Observation e5b681c1-99f6-46aa-a744-03fd1880f7fa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.189326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:3dc97e44cb5a49b40e3d72ce7e1dd3e08a4cc38c7d9e48c164d385204bda2c71

Observation 6c1efb9b-3909-43c1-b2d1-7091cdabdc2a · outbound

This paper cites SQuALITY: Building a Long-Document Summarization Dataset the Hard Way.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs SQuALITY: Building a Long-Document Summarization Dataset the Hard Way

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.693553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7668c015605189c2707aa3871aaf1ba2ed7c66c53913713bb026d2f1c6cdca42

Observation 8990929a-fe4b-408f-af02-62258c2dcb4b · outbound

This paper cites Covost 2 and massively multilin- gual speech translation.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Covost 2 and massively multilin- gual speech translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.271000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:eb1322b97b30347ce73c1f7fd588660d6dff54b733e461a7026b9c521fe36eab

Observation 6809925f-4b03-49bd-89e9-1edaa1d13904 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.710097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:011e1aff2ed4f6f159a16a6e14114f4cd01e416b9d4eaa32495f08de508a6241

Observation 0b29ff4f-6dda-483e-bc91-7765c994cb15 · outbound

This paper cites LIMO: Less is More for Reasoning.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs LIMO: Less is More for Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.899840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:2c03e2e0fa9194b327d52ac35438bb639cfb79195e649e586907bdd39e8d3d20

Observation 2ce20adf-103f-4274-952b-153118eaa12a · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs AIR-bench: Benchmarking large audio-language models via generative comprehension

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T22:22:28.258493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:cb9672b1dee5fd4a2e63a3883402405ca3f5b5df63838407af7cfd0701ae6b48

Observation 8182dada-8e4e-4b85-b829-818fd0c1e904 · outbound

This paper cites Qwen2.5 Technical Report.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Qwen2.5 Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:27.732347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:5a6911fe50d51d2e7e8fd77f9b7e2e965dd1769e58e936e186bd562887637f4b

Observation 51e1c557-58a4-4bf6-889e-06feafbedde0 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.633687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:6f0cdd3b6927f2536d77da85d94b359cb092a172c70246dde584c8b84b1b985c

Observation aeb0577f-abd5-4ab6-bc11-1462b817b8cc · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.759237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:c80b74ea2a50899700af7ffe52340ff91bce71168ce5adfc5b286c12dc9e9c7f

Observation 8483c4d9-565f-4c1c-a4cd-6bf3729d0d31 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.772873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:914b41e5dd1c62d06dddace6edc9e6d3a44c2a9d89296e7c536892885c2f2e7f

Observation 75b57934-9208-4775-af26-56708e043fd2 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:27.783266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:7dd3cac61f905e48610ac26e252d16b7472de0e80f77bbb95f35b5f30040de12

Observation e4f0b3ff-3453-4591-a366-942211b6c35c · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:09b79c17da2264965f130fbd2a94c3698aa073cd678a111d9392d4bdd10d193f

Observation 86149ad6-9a2a-4921-a378-1339e2d11331 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:114e7c498570dc5326c04977ad55448db7542f235cc080a3c680eb177c2100ae

Observation 2e30a073-f24a-4245-a6ac-28c6e8cab7f1 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:30:10.582609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:dd0bf87bbdfa3dd37916797a31928a0d286b44d8174eeb032ed62aaefedc64cd

Pith citing papers

Observation fc50935c-2548-474b-8451-a7780fbbe5d4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 153

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.932026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:038642e1a2f9a472ee5d8b9ec9a4e558647ee73e3b1ab6c7ee83459211109fa5

Observation 57ae7c8d-8655-45fb-b70d-4d79c383c401 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:45:08.095879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:5209bbc008411af67485e9b22dddacbcf1d3be9c4f697b1250d19bd40eb152eb

Observation 8f2789c6-f92e-41f1-aa8e-49bd94c96e37 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:42:45.329287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:3e3b3bbc44b93b511a8f4ec9cdc9ac69794de2bcc421e1e6c072449a9f88b5de

Observation fb8f2733-32a7-48cb-b36a-42a92b07b0c8 · inbound

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models cites this paper.

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:41:59.995046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T02:38:27.510872Z digest=sha256:dcb85d858ebb894ea0ce4eed303ae7da0a74237798a5d2bebc19abc0029b669d

Observation b41b52fd-ff1e-461b-9ae1-a166d668d306 · inbound

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation cites this paper.

Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:28:41.934494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T21:28:41.904725Z digest=sha256:26f4ac7c7040fea56e8ea9d5546934b71fbfeba663abb374e04bf061c84998d1

Observation 2bade1fd-f17d-41c8-8fcd-d058ffc92a10 · inbound

ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation cites this paper.

ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:16:54.039837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T23:14:58.555068Z digest=sha256:19884ebb5600319ab033922049e6edbfcee2e621b960be119e8db1261200a287

Observation 7e801151-e8b9-4e50-b28a-4746425bba15 · inbound

Evaluating the Impact of Verbal Multiword Expressions on Machine Translation cites this paper.

Evaluating the Impact of Verbal Multiword Expressions on Machine Translation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:51:50.257012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c66e6838d65f00b7431c59e8696917c4e02a0c121db14dc9b0f855e682058e13

Observation 70747e53-9514-4468-9a13-549f43fcbc9e · inbound

Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies cites this paper.

Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:16:51.174160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T21:14:16.601070Z digest=sha256:a091d9bf9d3961016247dd1881864266ddd8b28ffb9a568af0a86346f66049da

Observation 349de331-6c6b-4bfa-9182-3db61011571b · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:02:37.479387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:ca0399ff3fc897b08f25dfc3b7a7ab262aec8a3f3f73910267c4d58ba1327760

Observation f15fc938-5188-44a6-9259-c8a8a484dae1 · inbound

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems cites this paper.

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:12:37.223016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T13:12:34.387094Z digest=sha256:c5af4b69f797bb71ba1781a062de748cefbe04203e74b413f5f55944b82144e0

Observation c85b32d0-1da5-40b8-8a98-3b3ab4235798 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:11:17.911796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:88c2826294fec83400f0ece2e9cd028dd8ae56f6ee89516ec7a5d8d32702ad4c

Observation 2245ba5c-9be9-4832-983d-fcde7624f0d8 · inbound

Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior? cites this paper.

Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior? Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T22:04:10.561277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:04:10.561277Z digest=sha256:e050ec8f1d80b8d1d94c16ea51a05fea6de2c660b443fcaf22ce4b8028926897

Observation 40843380-1b75-42f8-ac3d-a1620e9a27e9 · inbound

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds cites this paper.

TRANSPORTER: Transferring Visual Semantics from VLM Manifolds Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:01:34.678805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T05:59:14.478279Z digest=sha256:86d7b4977b03037002aa741b7a51f7c691a7ba61a0a279f04ac8d6cb7b5f4499

Observation c1cf2137-0847-4b67-9c2c-0c74ebe91a48 · inbound

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models cites this paper.

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:41:48.653117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:41:48.653117Z digest=sha256:0cfa2b0e006bf4a4d201dc1b81a6db0cdd0e0ff486142596f76d0178f999fa94

Observation 89fc915a-8bc4-4449-96ce-5056f5e7d23f · inbound

Different types of syntactic agreement recruit the same units within large language models cites this paper.

Different types of syntactic agreement recruit the same units within large language models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:43:53.235603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-17T02:43:51.273343Z digest=sha256:6abe803c3732efd43d4933b793ac30b0fd8eae56db958f50eda6b0e822ed15df

Observation 1aa76f50-759a-4e00-9a4a-da4a95f0bf81 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:51:17.818489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:712c1a35107c881f06d6afa8ee4916fe0dfb9e0f192c3bd21d77d50d64e9e712

Observation 24ff696b-f949-4fcc-a3e9-5f6a45c96737 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:01:58.870711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:01:58.870711Z digest=sha256:d2fbdd54e1df105aeb9cb644f297f40080e53c05f56efe5400784e481bc7de8e

Observation 1146e15d-c8ba-4895-9d60-a9253b2fd06c · inbound

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus cites this paper.

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:31:01.952272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T14:30:46.664861Z digest=sha256:13391e40338abaac9cc248d71305541c55565d09ca45366c1f171f5072725f3a

Observation 06a6f66f-5af8-4c96-affb-752da934bb85 · inbound

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering cites this paper.

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:07:58.668659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T14:04:17.935630Z digest=sha256:bb2f8ce88608fdc91b21ece550fc7b79392ee9c22408b7f1ea40e1513eaca127

Observation 89b64e93-047f-42b4-9867-1c9ea27f28e7 · inbound

PAL*M: Property Attestation for Large Generative Models cites this paper.

PAL*M: Property Attestation for Large Generative Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:37:48.703148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T11:36:55.160572Z digest=sha256:1cd413414d81bdc07161519be5a0f74d8a08b53ffef4c1d41345eabc36d23f65

Observation d11f1a18-6b2f-40d1-8637-a8799e30cd2e · inbound

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation cites this paper.

*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:50:43.132868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:50:43.132868Z digest=sha256:6c1700899cc0c8144647a2f70bd3ed3665145ebe473c7c40462f0b92054c0fbc

Observation 0d04c0e6-31a9-4951-9d8e-717538f328cf · inbound

GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant cites this paper.

GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:30:14.474460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T18:27:16.527361Z digest=sha256:4dee653d2ab99913a9944014d8fb3c103ab7a3710006646050f67e222ed3eee8

Observation 3222f388-5a43-4662-9b73-556720b5c28e · inbound

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding cites this paper.

Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T13:58:09.323985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:58:09.323985Z digest=sha256:13ca1dc459ede8da08ea42da91f56d813e908e2008a65f49b73407899f3a19b0

Observation 0cb3c854-847d-4ea5-ae15-6bedb7ad260b · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:45:28.322239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T10:41:22.138517Z digest=sha256:50f7cf3abb87553f1fe1649242b7a5a8edd2773fa71f7a5266e33eab491b87d2

Observation 3aff7a7e-66cf-4c83-a733-f6787ccb92ca · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T20:46:17.286576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:46:17.286576Z digest=sha256:5dda721570a390daec3e3e87438b1d47d50826c647be7318e85d911a6245484b

Observation e288d52e-4ea1-4669-98eb-4a89e9b48209 · inbound

Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models cites this paper.

Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:48:24.987287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T00:45:43.705160Z digest=sha256:ed78564ba4763cdaeabdc82a25981fb7a6f31f83bd32f60009750c3376999d97

Observation 36ab39a9-6ffc-4be1-a908-eb17e0495041 · inbound

Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification cites this paper.

Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:33:51.519897Z digest=sha256:1dead49bcf64d30a998417f5fc2fac710a345b0dc48ef4fcfa38c90e3c0aad5c

Observation 57aa726b-99ee-441b-b6c7-ebeae55e506c · inbound

DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues cites this paper.

DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:08:18.876839Z digest=sha256:c37323b3d54ff70e5738b9ced1724eb96ba297b7cdeedcd8a94ab356e0455056

Observation 4b361755-3521-4b04-90b6-f311654cb79a · inbound

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing cites this paper.

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:56:12.097628Z digest=sha256:fa846f18edb3a6cb95a0359d7db950772dbbe7ddda218734b2808afebcded985

Observation 5e717bab-89e8-4e41-9a2c-19ca42fb60ca · inbound

Demographic and Linguistic Bias Evaluation in Omnimodal Language Models cites this paper.

Demographic and Linguistic Bias Evaluation in Omnimodal Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:05:01.819578Z digest=sha256:dde4a8d7dcd70756257aafa95362d9e8b48343e0b64e2c2e1846f7aad407b69d

Observation c87e02ee-03cf-49ae-9398-94046dc3f5d3 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:c202968146a4f7e64559beaa85b39ea518e3789d294f29e15ef117591cb0d4c5

Observation 2872c852-abe0-4d71-808a-0a2b2931e3ac · inbound

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms cites this paper.

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:21:48.653698Z digest=sha256:ca54b72d9d78a47c1b2c33918a727db1526ebf44873011366a8e6177c14c9b1e

Observation 5c181fb5-36d4-4875-a844-2cb19b3469e8 · inbound

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms cites this paper.

CheeseBench: Evaluating Large Language Models on Rodent Behavioral Neuroscience Paradigms Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T00:59:19.369998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T00:54:55.296944Z digest=sha256:fa47a1fb52fbc736caf22d6406f53a19f10d1c44a4f010f9f0ac5b357e26e9b2

Observation 95db6952-3fa9-4d1a-bbe3-9fd43d583bf3 · inbound

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction cites this paper.

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:35:20.515352Z digest=sha256:79781807f8710c15768a32032e45a0b2d377524a392fbead8a77d994b008bf58

Observation a1d92992-a423-4fc9-b6c0-899a0b3f3905 · inbound

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs cites this paper.

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:09:51.365573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T08:06:45.477344Z digest=sha256:e38daac8ad13b5a6fd721f1c5efb81a09e4c4b1cad377c8441d9e5b771700368

Observation 4f47e51d-fc89-4997-9516-e857365c2277 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:b434e0ec70ec6f3bc86af6a304c7fb494eaef909cbfd1796e920193a3fe6fdfb

Observation 284115d3-ad83-4938-b6fc-a30a5dd9bd2f · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:89aaa16bc7e70d9485654a6facd04b912206a4e08a2e8656399e21ab1d51a486

Observation 5c3933bd-3a26-4a96-87c3-1f7770dfa598 · inbound

GroupDPO: Memory efficient Group-wise Direct Preference Optimization cites this paper.

GroupDPO: Memory efficient Group-wise Direct Preference Optimization Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T09:43:18.432084Z digest=sha256:edc8b1a8a65aa6f039891374f1b263a1abe37fbae65cecc998db78604af0246c

Observation 69d20138-12fc-4bc3-896f-7d12f047b861 · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T09:20:22.200577Z digest=sha256:88ad891a2863b022fea01db99e729adb12c1750f4bbe68866c1d1c0608542703

Observation c971776e-6541-4b9c-a51e-166e3f6bb18f · inbound

MUSCAT: MUltilingual, SCientific ConversATion Benchmark cites this paper.

MUSCAT: MUltilingual, SCientific ConversATion Benchmark Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T00:49:19.556887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T00:45:01.115573Z digest=sha256:3630c1fd3b3781724eead31cc43930ebd44949fdd5c40611b96cd46941fb33e4

Observation 043d299a-f7f1-4bdd-bbd3-fe76ecae512c · inbound

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech cites this paper.

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T05:46:50.923340Z digest=sha256:41e2a7f98a7117e192a46b27b063987f1356c8269e2d6d2b84a60a5197a37e35

Observation 67445640-9fcd-43a2-a00b-1e415eb7a2aa · inbound

UniMesh: Unifying 3D Mesh Understanding and Generation cites this paper.

UniMesh: Unifying 3D Mesh Understanding and Generation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T06:55:42.679323Z digest=sha256:705ffafe390d39df591249f138640521814c6a9eb991185ec1fdda656551f5d0

Observation 73f61951-d7e7-4640-9e0c-e49ef425e73e · inbound

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models cites this paper.

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T01:34:54.375266Z digest=sha256:59903a50bc7a4b20de1bb121315827cd5839580b19bf3bf895ad0afecd8ff46d

Observation 16bd3282-88d4-47ec-8ff1-283c90a56f76 · inbound

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling cites this paper.

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T01:14:16.831333Z digest=sha256:a93fe78eb6c8e6054794054dc475da13dbc54077e29d7466a121b804168994eb

Observation 034b8f9e-1d1c-4f36-98cf-c08814fee02b · inbound

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA cites this paper.

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T22:11:27.005071Z digest=sha256:61c63fb5746dd56459d3fc499eb774a488844b594a58ddb74f211dae70cb3d97

Observation 0b32c577-78b7-47c1-bb5d-5e9f967ab1af · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:11c33788a882ef738cbd87586c7e2c491fe5644a6e44d6ff33e8c6c354e7dc1d

Observation e9899b37-44e8-40af-bc00-b577a5b281de · inbound

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models cites this paper.

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T05:20:30.823304Z digest=sha256:21fe9a47f4457d52d28f0f67bee3416b2eb2df2d837a97d663410e9fbe4deb37

Observation 71fa4ab1-7c5b-43dd-ba03-0aea4193ab5e · inbound

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation cites this paper.

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:36.277354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T17:39:38.234052Z digest=sha256:60da6ee0a11eeff1cb54f8e0e5e0c42a8098f02be52a7023414f6a29c454bc90

Observation b8f59b96-2c41-4705-9995-ca30cc122839 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:41:26.438714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:597a14511af0b8124a4fc3a462f915901b5363b377df78a0fcf3d514fe5e5356

Observation 62893fbb-299f-4133-8436-3b242fc3ee0d · inbound

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology cites this paper.

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:16:26.388012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T11:46:28.742487Z digest=sha256:2dc8f408b5a9b2d74b4fd5d6309fb24f24e7abd75c46c22f4720310fc3323975

Observation 70cbbf45-b33a-48ff-ab37-7c7b7e72927b · inbound

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs cites this paper.

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:17:27.351727Z digest=sha256:cc980a1f29f985ae19a61cf64e55f8eef2f43c71975055fbb506702c00afeca8

Observation 1b7778fc-fc47-4cba-b6ba-f68cfb58fc82 · inbound

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR cites this paper.

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:46:27.897018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T09:12:14.094145Z digest=sha256:bacd00a940c53903eee3d5ff32fcd5890e9a839e6fc5042497ac002d87b0dc94

Observation 56adbc4e-cbc1-46e7-942b-24a93f79c35d · inbound

SENECA: Small-Sample Discrete Entropy Estimation via Self-Consistent Missing Mass cites this paper.

SENECA: Small-Sample Discrete Entropy Estimation via Self-Consistent Missing Mass Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T18:57:31.542345Z digest=sha256:119c39af4bb547da9c799913c82468b94e9ec76609c8bd05d75fc83bfb15f3bb

Observation 42d1680b-dfeb-4d33-8036-6f100a545789 · inbound

Multimodal Data Curation Through Ranked Retrieval cites this paper.

Multimodal Data Curation Through Ranked Retrieval Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T17:59:49.890800Z digest=sha256:6396567d96c38186cdb41e567cd03819701df4a5534529c79ea57a6bd4179460

Observation cc1fed03-1f8c-4942-94bb-f2dd98a4e270 · inbound

RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI cites this paper.

RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T09:12:00.835585Z digest=sha256:d49a9f99c415e64688aeef2b9f8f86c85b5860bb0af41c2cb80f6624f4320732

Observation e418c78d-0996-4969-be86-3171c5a958a3 · inbound

RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI cites this paper.

RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:35:07.752240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T23:27:32.526304Z digest=sha256:a5bbf9705fcf72c5800e876429058922abcf7309c88e78d59c75283fda5b3283

Observation 8f7f689e-fdfc-4527-b622-50b66d9dfb90 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:3d8e2d70215d1ce9df55786d3fd6e7ae34dbe8b6da3646847c2124529c00fca0

Observation b1fcc56e-8a75-4b8c-86d3-6acf933db2e5 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:12.100964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:12.100964Z digest=sha256:0f7cda9934aef1742ce0f21b255201adba94c19a195b4e7fdf8d52286f74a372

Observation f1665edc-74f6-42f9-be50-114b16f0cb87 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:b4da9cd0938d5d3c3f5b1a3f27a45ab18006a8003ac365384b5ceadfab2da61e

Observation 47553457-1369-49ca-bb38-923fa5a77944 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:d40777b103f90501ea84f42961e85d013bc29e8fb3c0a2f10591fab04d00c256

Observation 9affb265-1f43-4a13-9fa8-3e5d2889b799 · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:22:28.302846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T03:01:13.168529Z digest=sha256:205dc66ce9cba46fc96aeec59afe24ec99e22bee5ac73afa8cf0918ba9fdb3e6

Observation 18d09668-4518-4b80-b812-32da377dc361 · inbound

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation cites this paper.

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:24.938742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T10:30:11.309910Z digest=sha256:c09f8d28cd004273f5bf93a4343b66e12d6cc6e4255c732bff92ccbfb1e0e91b

Observation 08c5d7b5-7783-402c-9bdf-aa30407e24c6 · inbound

Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills cites this paper.

Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:56:27.312250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:49:49.484361Z digest=sha256:ca4b98ce11723a8dd618d57847463b82d88eab9a59f0435feccca319eb93e98c

Observation 9277b1a2-b629-4636-8892-4324cd353a98 · inbound

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization cites this paper.

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:24.427716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:01:23.060562Z digest=sha256:dd65edfadea3622cbcd4f4fe9007e95bd98b80fa32f240fd95a542b6d426ac43

Observation ddec2916-69e0-4a60-b4c6-831bf7949106 · inbound

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models cites this paper.

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:35.186477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:00:50.940530Z digest=sha256:a596c542c508d0709a879bdf31056a04d08ec9b945b39cfd230867c4dd92aebf

Observation 25368d89-2efe-4311-bef9-dd183ae8ab52 · inbound

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA cites this paper.

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:56:25.303513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:48:09.071812Z digest=sha256:f9faff828417ba781cb3eebac1672e455eeeb3818e7b0010e051b2dd1141bed0

Observation 5051896e-5e5a-4762-913b-cebfd0e56904 · inbound

Differences in Text Generated by Diffusion and Autoregressive Language Models cites this paper.

Differences in Text Generated by Diffusion and Autoregressive Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:59:27.192307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:59:06.804446Z digest=sha256:e131934a39ad79daec965f2c62c0d71468865ef2d7e8fb152b470aaa5697d105

Observation 176a8cba-f680-4354-bdec-bec0e0afdd25 · inbound

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores cites this paper.

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.289781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T21:31:14.913156Z digest=sha256:ba42c2c39dc305f971adf32a8b50a51267814719191b282cb2d11087add58d85

Observation bdeb7108-6c71-4401-9751-87770047b854 · inbound

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models cites this paper.

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.230702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T21:00:25.664841Z digest=sha256:0a44d1c833efe9f5584eb1b853d5f3ccf17511dc2c3ad8fc5d5127b1b86c5db0

Observation 3ea3408b-e9f4-46b0-baf5-00a09aafeaaa · inbound

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents cites this paper.

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 118

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:49:53.807525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-21T08:45:56.550821Z digest=sha256:fea9199f89039a966485ba3775fe8a26c5886bcba74a1094b0ea9952acebce88

Observation f69373fe-0e40-4211-b73e-43ab0adebcae · inbound

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs cites this paper.

MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:09:44.605032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T03:04:27.417831Z digest=sha256:ca6e2403229694508a631eaf8f5bd899912d986f019e06d3991ef900111ced67

Observation 5f0db0b1-fd7b-4393-b486-ea38e3944eaa · inbound

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making cites this paper.

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.408415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T14:48:55.203993Z digest=sha256:1a7384719ad6cdf6f3de8b783da758e31abcb3e48b4ebd472bcceb70e110973d

Observation 21f2d9f6-9b08-4086-8101-fafa1fbd444e · inbound

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings cites this paper.

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:28:19.106437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T13:25:09.024054Z digest=sha256:d3b1fc1892c0780e666a4e9da37513de69f0eb51e09103c7abc1f4541a2ed6ee

Observation 1c1bd359-02b1-4eec-8e9a-b2ad6134961f · inbound

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions cites this paper.

PAREDA: A Multi-Accent Speech Dataset of Natural Language Processing Research Discussions Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:33:14.431845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T11:30:33.672467Z digest=sha256:2641aaae046d5be8e4e77ebc7914545114496ea9e2c7aa42ca113b0fd7615752

Observation a94cb6ff-fe0c-4df3-bc32-128b74fc7656 · inbound

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction cites this paper.

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:13:15.941762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T12:13:14.295659Z digest=sha256:72aed508108752530e53034f4a1574d3b6854e0824d7654e362aaa79fe39d687

Observation d8ab7666-9b65-4167-8ba5-c9fefde59fb4 · inbound

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding cites this paper.

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:08:13.489000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T11:06:59.215750Z digest=sha256:280ff3769d30a72640a8b9bebc7f3c6441ca970d863026771d21daae4c76da5d

Observation 5edbd3f2-3b67-4f3f-911c-c560858b7bd7 · inbound

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning cites this paper.

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.007229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T05:25:51.655208Z digest=sha256:13c106fe0db856bd7b8c43873a3966e9cbe7baa11652dd6eb689872ade489d36

Observation ea3e0ebd-899e-4157-9c75-eb8cf43c7b8a · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:13:58.426891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:0737074c8e05f828e1ee5eaf96ee6c1a1974068a0d5fdff297ee5cf6fba77736

Observation 428e3860-292e-4a38-9643-6a22984567e0 · inbound

X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation cites this paper.

X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:04:47.067371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T10:04:35.334305Z digest=sha256:9621291765de9ee68899a5101782e513eaaaeaf49e3cc2b113c46b1f9b73f689

Observation e316f673-279b-4fcc-a9e9-677f72a9f445 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:55:25.040607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:a85e7dfd9912ccb4121313d6fbb25d4f33b341d4134ce724737654566e72f3fd

Observation 5ae4628e-1bbc-44a3-bd87-07b61e3dff07 · inbound

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs cites this paper.

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:25:04.331128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T21:22:15.413485Z digest=sha256:af2829688db92eab4a0ac087aeff4ee02243c01355fded1c8f645d63faf4ae7f

Observation b2c422b4-a748-4a71-a388-5dbfedde6161 · inbound

Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling cites this paper.

Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:38.723471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T11:06:21.527926Z digest=sha256:7a61f5e7aa61b0800a065ead8b9eb75ab7744e25a43bedbbee51ea068d24658b

Observation 77d12bc1-f31d-4c81-8966-89a4d999cc7f · inbound

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models cites this paper.

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:00.665863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:53:52.992152Z digest=sha256:e645d13fa04903c111a6db68d64cf3321c70c7e31e59cc919f574c4df98c7895

Observation 043e5342-89fe-4824-b2f5-cef53576b92e · inbound

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding cites this paper.

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:23:59.913079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:22:53.215257Z digest=sha256:e1a6cade57c67c8a3d23248e7b231bfc87284efc14100a5175a28173c38a2efb

Observation db54f728-3bf2-4e03-bdbf-8a8e2b8d32de · inbound

The Future of Facts: Tracing the Factual Generation-Verification Gap cites this paper.

The Future of Facts: Tracing the Factual Generation-Verification Gap Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:33:50.402930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T18:31:04.169632Z digest=sha256:62e593c2723064fea712d183d0b812fdf30c30095413ddb14a4bfe2c917e5c14

Observation 06aeb066-d758-4876-b277-6bf484e2eab2 · inbound

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs cites this paper.

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:36:09.049165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:16:27.574100Z digest=sha256:789cacea8e7ed7c598c77b2be4913c2a208da7afd72282310d55c1bde1c4c806

Observation 424d0c7d-918b-4566-ab4d-1c75c5420324 · inbound

RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting cites this paper.

RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T00:02:49.592598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-29T00:02:19.803078Z digest=sha256:3524814a901b8998bce43bc10dffa498caa5da517bfc2a819856f1de121b42d3

Observation 30499a84-c286-415d-8e55-7667d0aac91f · inbound

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors cites this paper.

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.173932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T19:23:19.422735Z digest=sha256:d282611c9476646648d0fc657eba068aa4e384d5ad90489b35ac83f0ccc345b1

Observation 8f14a4ba-c57d-4d2a-94d3-ed1843609af7 · inbound

Differentially Private Datastore Generation for Retrieval-Augmented Inference cites this paper.

Differentially Private Datastore Generation for Retrieval-Augmented Inference Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:36:15.468352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T16:39:27.164341Z digest=sha256:fb62f618c5716c8650ec771d1077a66357459520d89f3523e6e0945d8ae9635c

Observation dd674201-08e3-43c0-a27e-cf7529506a43 · inbound

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring cites this paper.

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:27.755036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T10:44:00.642366Z digest=sha256:34362228a5469baca621f02db30daed4456e4e60b92e244397f5d67221fe157a

Observation 435a9621-9ee6-4fc8-9c4a-eb7f8b31320f · inbound

Do Transformers Need Three Projections? Systematic Study of QKV Variants cites this paper.

Do Transformers Need Three Projections? Systematic Study of QKV Variants Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:17.325422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T15:14:49.475561Z digest=sha256:0991141be41f599a7b0522cf8ede66ea23dcbe7b113050319529b4f9b33b079f

Observation 1f7d7371-2bed-413b-a1f7-63e608569244 · inbound

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention cites this paper.

Entity Binding Failures in Speech LLM Reasoning: Diagnosis and Chain-of-Thought Intervention Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:46:46.246146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T06:41:25.592270Z digest=sha256:f4a26419f17219f36788c2b87cc21242f59601574aa0b1f9aa7724105b262b0d

Observation 4f586ca6-4abe-43d3-8227-8289100e9a76 · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.330569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:43f41ead8fd4f0c838eff731fe688b7187da291176367154ffee0505618385b1

Observation 96a7e54d-9606-499a-974b-6bdb0496d58c · inbound

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction cites this paper.

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.971075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T02:57:54.629402Z digest=sha256:229befad8b535e45a290a8c68b01f7424cc38557856590b3902e765bc83bb4ce

Observation 0c21627d-9c6c-4692-a620-dc1ff02ddb33 · inbound

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios cites this paper.

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:36:59.048458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T01:11:56.037674Z digest=sha256:6b68cd3f512db918ae2baf75c17c16e9755ced7d2327194fc3da582a60423eea

Observation ed67616d-4ecc-4e8d-a061-ec2edc2ba58f · inbound

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding cites this paper.

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T17:37:15.018370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T21:54:57.736518Z digest=sha256:5919f005ab7e9a96c9a1983dde7e389372fd6a356a9d9171719dcfdc4320a069

Observation 843a2925-253d-4f09-a27b-2f2af3eddc3b · inbound

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER) cites this paper.

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER) Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:44.979232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:51:15.098039Z digest=sha256:a0b2085c6984bff6b849e76dfefc4df0c927e4972afe9a2486f2c09ddb69abb4

Observation af222d8e-548a-4726-bb9f-ee3dffeb813a · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.441366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:9ef67050e221e74c34127445b6fb5b5240297b902791f8cc7cea37df81f3b6bf

Observation 80f64914-1c6f-483e-b062-253aec610cd6 · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:07:26.490409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:3f716392f216033c0e37ffc38a536725d6868766a6ee7ddbc72dc9c61d272c02

Observation 649b619d-444d-4692-8a82-376eec771c58 · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:47:35.737632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T14:40:40.673091Z digest=sha256:d5877800085abdd621292ff940481d9f872f4e8d0810478bd4c0b8b887a92ad1