Pith. sign in

Paper Citation Record · LEDGER

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2411.16761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16761 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:51:00.548967Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T11:12:41.130806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:13:02.889016Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98657fc5-a208-4ada-b229-c5b56bdc3602 · outbound

This paper cites GPT-4 Technical Report.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.328980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.328980Z digest=sha256:86cd41dd068a4a2e845cf7cc4efa9c20adcd158841284e6e5c28efb2716ed8fa

Observation 92df7e73-02b3-4a55-b01c-0f6accf25c7b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.333999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.333999Z digest=sha256:74163cafb81726e18d380c35f7d9b0ac68d5eee43680b573b453a55b6820a5c6

Observation 8d042221-d275-4515-b938-ba29d7cd1c47 · outbound

This paper cites Claude 3.5 sonnet model card addendum.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Claude 3.5 sonnet model card addendum

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.383100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.338497Z digest=sha256:5224f40c0624d828926c08c8092809e36e4b8c4fdf1b4be75f7763d9ffa3d6a1

Observation 9d7703ce-7359-454d-84ce-77bc2a1a4653 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.342915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.342915Z digest=sha256:c5bdc16392f9b6df232bf52ce04477e5a3f9ac52ddbb81d8fa19bfff212a479b

Observation d5c3bd4e-2345-4f79-8f22-a7f799a9f909 · outbound

This paper cites Egocentric vehicle dense video captioning.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Egocentric vehicle dense video captioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.347288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.347288Z digest=sha256:f68350def5d8a125536f397e37dab109742cb81f3662d5ce97b860eb8423e12f

Observation 3652958d-4a66-4fb2-abf7-20a1e0836eaf · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.351623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.351623Z digest=sha256:0aa5e57be6787d1645285f863cb0c67652b24fc535861523281d387bd4d89f68

Observation 7d66eed1-557e-4b08-afad-00300d730aaa · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.361261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.356234Z digest=sha256:bb8175e56822c190df5782aab58d6eb35b665685dbfe8ef700fc315b05708709

Observation e8a86ecc-dfb7-4ca6-81b6-dcb33b12e495 · outbound

This paper cites A benchmark for reason- ing with spatial prepositions.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning A benchmark for reason- ing with spatial prepositions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.347413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.360458Z digest=sha256:ccb3133ca3a6b8ea41bb63b48703b0b60c3534948559298632cd67cc9d4b29eb

Observation dad0ab57-fb61-42bb-8e0a-b7e920424163 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.364493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.364493Z digest=sha256:fe689cb5ac11cc418f6e6fedf1eea7b5e811a6bc4bee8e46d3dcb11c36f8a03c

Observation 7aa87d5a-031b-4b8e-8a3b-eccbd0dd0010 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.324596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.368658Z digest=sha256:eb531ab970574db76d71f77ec69e21aa5c04187ef066559390b4d70f689df677

Observation 65f975cd-6008-4083-bb41-586171b570ae · outbound

This paper cites Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.308194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.372958Z digest=sha256:0158616b367b2530ca30715b3e54e0462819eeae3314775e4fb7c4541c1d8081

Observation 6260e3ea-0abc-4d22-9a38-23913b96f1df · outbound

This paper cites Pedestrian detection: An evaluation of the state of the art.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Pedestrian detection: An evaluation of the state of the art

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.289802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.376989Z digest=sha256:c18bf69944a372b96274f64c3df13d0c16266fca560fda05e8abe37c0f2b6896

Observation 0debd8bf-b013-495e-8e81-4a24490c601b · outbound

This paper cites Pedestrian movement direction recognition using convolutional neural networks.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Pedestrian movement direction recognition using convolutional neural networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.273896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.380797Z digest=sha256:c981172bd0457fab45f878d921344931b413ebc2a0bd38e1736f13c611f4be36

Observation 468b7edc-3849-497a-90b3-76bd82053dda · outbound

This paper cites Palm-e: An embodied multimodal language model.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Palm-e: An embodied multimodal language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.259070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.384377Z digest=sha256:53f3b7924313410305f79653f4cc68061c5fff09f4d6d343c21aee94f37d07ab

Observation 166481fb-dcd3-4cb3-9471-eb55c41bb38f · outbound

This paper cites Integrated pedes- trian classification and orientation estimation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Integrated pedes- trian classification and orientation estimation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.245433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.388481Z digest=sha256:c0e4829e7d6b522d955d522a094b894617371ac3483542ff27ebc9b3cf46809c

Observation f6285ba5-5f0a-4c36-a443-9e336c2b9168 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.392494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.392494Z digest=sha256:c4832560936e97e0723812064f9dd28e62f7837c01441d02136d62142c90c2f4

Observation f4eba972-771a-4015-a955-4f89dc31be2f · outbound

This paper cites Image based estimation of pedestrian orientation for improving path pre- diction.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Image based estimation of pedestrian orientation for improving path pre- diction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.231701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.397035Z digest=sha256:975b37fbf9efb319be3559b76f7318c53f98b698193f93f59b29d0224fc6593f

Observation 85b136eb-d572-46ef-86b2-6850ecb7323d · outbound

This paper cites Phys- ically grounded vision-language models for robotic manipu- lation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Phys- ically grounded vision-language models for robotic manipu- lation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.217051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.401239Z digest=sha256:2e48b816444b4e6554d6d05e055d985d23507399e71b55a9a00703795b64b9d1

Observation c2a4c464-c645-4b87-be3f-806d71ab2873 · outbound

This paper cites Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.405551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.405551Z digest=sha256:07f8e9fcaec44e32945290ebf632f5eacc98413c799b499d30f06ad0414d261c

Observation 5b3f2ce3-ce83-4c87-9ba7-2635f92b801b · outbound

This paper cites Human brain dynamics accompanying use of egocentric and allocentric reference frames during navigation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Human brain dynamics accompanying use of egocentric and allocentric reference frames during navigation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.199835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.410766Z digest=sha256:771c9431cb886f1a13e3d0e8858249d7a158b848a30735a621a703116e627e1d

Observation aa94c7c4-0278-4840-baee-96c27dd51a17 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.415796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.415796Z digest=sha256:e2b83261bde8cd214fa73756df2db3aa37b610f54be3d9f438dcf2264efdfaff

Observation 3bc83d3c-d248-4095-bb19-9cab35fd42e9 · outbound

This paper cites Loc-zson: Language-driven object-centric zero-shot object retrieval and navigation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Loc-zson: Language-driven object-centric zero-shot object retrieval and navigation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.177741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.420453Z digest=sha256:ffcddbc2b4cac30e9ce1841940c958c422a63d3c97a410db8f543ae94bb22d60

Observation 49d902cb-5626-41ec-b50b-a6afef179bcf · outbound

This paper cites Lora: Low- rank adaptation of large language models.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Lora: Low- rank adaptation of large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.159976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.424618Z digest=sha256:f3f1c3da3744f7c2d2f7cf4d30d6dc1568e474a132eee0b038df701338f60812

Observation 3825dca1-753e-4a4f-83a7-3065d42a2508 · outbound

This paper cites Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.145535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.429428Z digest=sha256:776ad6105fb97993cf21df4d5deeb7495e87bccced6a56657b4894ca91b3e091

Observation 30254ba8-f78c-4466-a255-87442d5bdf37 · outbound

This paper cites What’s “up” with vision-language models? investigating their strug- gle with spatial reasoning.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning What’s “up” with vision-language models? investigating their strug- gle with spatial reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.126517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.433937Z digest=sha256:4a5be2442a5cd5af40aa0a7aae82814365aeaab8a7355104368d7938ac2d7334

Observation c9f4f3f9-f795-42f9-b4a4-bb8c3b1aff51 · outbound

This paper cites Frames of reference and molyneux’s question: Crosslinguistic evidence.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Frames of reference and molyneux’s question: Crosslinguistic evidence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.106047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.438133Z digest=sha256:2df1bae19f1b884c67ee3e41afcc0426f259794ba4cd77917bdbed301c22f766

Observation f01eadcb-cf25-414b-91d5-35854a66354c · outbound

This paper cites Deeper, broader and artier domain generaliza- tion.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Deeper, broader and artier domain generaliza- tion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.090421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.441932Z digest=sha256:80e1a2b95218bdbce67cdfef55785c094f3c803474f522297fbf2ed7ddd0b0df

Observation f9f762bb-497d-4eca-b956-d567e5a32bc4 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.446313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.446313Z digest=sha256:2536511aa47db09d22fe9de1170a94aa4aea7ed7e466b34b90f22698a79b799f

Observation 2aa98f29-2aae-4a5d-9c04-89eed2b1baf2 · outbound

This paper cites Microsoft coco: Common objects in context.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Microsoft coco: Common objects in context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.450789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.450789Z digest=sha256:7ca7e8cacd0fd00213b32889fa6d3a407a7edc853d0bef2d700bcb37c2c0b18c

Observation 0edff087-e152-4151-ab17-17ed8fb502d7 · outbound

This paper cites Visual instruction tuning.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.067161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.454515Z digest=sha256:675d0a5eebe4e93d7b8448ac8e99a5e99a34c378b6d8e16ae8d1a4d2c6215061

Observation 5f68169d-043f-4477-a9c6-c61676db6a2c · outbound

This paper cites Decoupled weight de- cay regularization.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Decoupled weight de- cay regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.458219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.458219Z digest=sha256:1c13089226c7d44c02cc0ade34b99ae54b00f481cf161f5213244c8454dc07e9

Observation a5d6e9c2-2afb-419b-978a-cd830d9868ed · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.462256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.462256Z digest=sha256:79c687a6c0bdcc96ab39fcb7c86fe4d5d2f61106cb016a7bf6948f7c2d2f15db

Observation bba0b25c-2550-41d1-9cbb-51a87562aa3f · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Embodiedgpt: Vision-language pre-training via embodied chain of thought

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:01.034723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.465825Z digest=sha256:94ce16911b5659cc4b5121bd6d8bfc9b4b76b3129c44a95609562cdbe2fd7270

Observation 9406d84a-1528-4af5-a02f-6e5c1e1fe225 · outbound

This paper cites Pose es- timation for category specific multiview object localization.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Pose es- timation for category specific multiview object localization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.907779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.469613Z digest=sha256:ea9e4cc0066f2da1752c18e5157a582efdae10196fcc5c19bc6b534576584064

Observation 12005bce-e790-4a9b-af33-c1806d6fdbca · outbound

This paper cites Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.473380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.473380Z digest=sha256:0ca402990e70c964bce9b10129e6f70aee7bb3486459de4afedcf3d257bf8991

Observation eb825cd3-e964-4672-ab7d-5c3a1596af12 · outbound

This paper cites Moment matching for multi-source domain adaptation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Moment matching for multi-source domain adaptation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.477740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.477740Z digest=sha256:f191f684a1b34272203ced2262d49d4dd25d9c87d30ec29dedebaf66bbf12f7f

Observation 05368f0b-6c6e-4a22-8417-e7f4f4db5ee3 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.481880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.481880Z digest=sha256:cbc09a03e74df057c212553c75ef39e5912fd216ccf94f385a80553f457dde3f

Observation 08156588-335c-4d2e-939f-55b195cbde3f · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Dreamfusion: Text-to-3d using 2d diffusion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.486107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.486107Z digest=sha256:6d5afe3ed9f585743d634e309c645f19fe6e116582a05affe41da03d13edb8d6

Observation 6b02b073-5dcf-4928-82fb-bf9d71c1e174 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.490426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.490426Z digest=sha256:7be54d5ad295eb30c6698c693228539079dba05bc741e830e1a44290937a9473

Observation 7fed912b-d058-4c8c-94d0-64fc8a1f78d9 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Lmdrive: Closed-loop end-to-end driving with large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.866393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.494259Z digest=sha256:2b0e8a2348dc16a57807b0325d3570fd491f64b4990f4dc9b35ed61004634028

Observation ea9594f9-1af7-4836-882b-80ef7feae745 · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context (2024).

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context (2024)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.850417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.498377Z digest=sha256:9ef321d71554abe89a51ac3464c3f23cfc6718443e0e2a585bf843f1f9e4b11f

Observation 8a4acee8-9565-4872-a764-39eaea3d3b19 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.834661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.502355Z digest=sha256:6dcb6a3a2e8bc2531f79f28a27cc59594ddb9f4f9b74069b6cb3a1c03cb41c0f

Observation 14be0946-ce2c-41d8-b6c2-f9fd0333e926 · outbound

This paper cites A fronto-parietal system for computing the egocentric spatial frame of refer- ence in humans.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning A fronto-parietal system for computing the egocentric spatial frame of refer- ence in humans

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.816525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.508084Z digest=sha256:2eb766831ce06e8fe2236b9f4b955b27e7101127ad28f0420fd741f6069864a4

Observation aa019bca-add5-494f-a46d-ca74d7eb0e9d · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Visionllm: Large language model is also an open- ended decoder for vision-centric tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.514094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.514094Z digest=sha256:78e0a131382fd204eb21ad911acbd59355fb27966536617a1223996d55fd4ac3

Observation b662f74c-ac89-4e91-8e0e-f8c78be60371 · outbound

This paper cites Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.518417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.518417Z digest=sha256:4ea87c98e3765fd8d82c84fc9ebdf0dc2afde8959b3db3e0ae4cc0656aa2ac37

Observation 16c38780-a7d7-45ea-9403-66026202efe5 · outbound

This paper cites Editable scene simulation for autonomous driving via collaborative llm-agents.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Editable scene simulation for autonomous driving via collaborative llm-agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.523691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.523691Z digest=sha256:e21514b76bc7c21fefc929c3897ebd6426e8f69093f8328a988235b381211cc3

Observation 751cf349-11d6-48df-bad6-8f3640c01f88 · outbound

This paper cites Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.769966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.528072Z digest=sha256:e67a05e67db702e761896d7c17d64e1ca9256b06b94a5519a645eb3e78a9411d

Observation e6056df6-dfc3-4ef8-93b7-d56e6a06ac5a · outbound

This paper cites Retrieval-augmented egocentric video captioning.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Retrieval-augmented egocentric video captioning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:51:00.532156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:51:00.532156Z digest=sha256:1611625b80fbb9483988e21ef0fffc24b0a4cb1dabc8571647824ed9a54b353e

Observation c85efa5a-0a7d-4395-b640-954d00a9471e · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.745525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.536393Z digest=sha256:d62042af37a459cf47229cc1ab0280b59c7cdab6ee824ce9ba23482488eb9364

Observation dfe4577d-c8ed-4f46-9b5a-d89befcbdc5b · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.731522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.540490Z digest=sha256:0b38f6c855ab9bc113f9c30c2b8568a6d00aa953c62577e3598030a2465572d4

Observation c9e08196-fae1-4fdd-b705-42daaaa2cb96 · outbound

This paper cites The neural basis of the egocentric and allocentric spatial frame of refer- ence.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning The neural basis of the egocentric and allocentric spatial frame of refer- ence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.716897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.544475Z digest=sha256:f3514bc3cc20b4e7064824a39d04f2bd764951eefbfe1dbcb29c2ae836ce6732

Observation 4f05f16d-9d49-4f6e-9fa0-21376ae970bb · outbound

This paper cites yes” or “no.

Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning yes” or “no

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:51:00.700836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T13:51:00.548967Z digest=sha256:3a3f5cc7076d7e2e3a84a1b91674b335cb766599e5fe82239379a7d695bdd785

Pith citing papers

Observation f9d7873b-8f6f-47f2-8402-09bbd1246c7f · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.890631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:ca670de5d5fa4b271726640537e0a63933f06f070662a4a5177a2173f63ea3b0

Observation 390e76f5-04cc-4123-aafa-e0252ab906b1 · inbound

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning cites this paper.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.787696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:f8ea6c3cc2c6fad486f2f5350f61d252b73f4e53b36b2a0e1d5e9f7d6891a853