Pith. sign in

Paper Citation Record · LEDGER

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2412.01550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01550 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:22:04.066726Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:49:08.968786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T22:49:09.109028Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab8e3d1b-a633-4657-af7b-cea2f5bb4f09 · outbound

This paper cites GPT-4 Technical Report.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.859307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.859307Z digest=sha256:50182af60aafc74960987a3fb9d026a32c3f6e3bdbcc8b2fffd177c9916c83f0

Observation a937ec8b-c65a-4d26-add7-504d835e4de8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.863710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.863710Z digest=sha256:dd3b93c104b56f87da05ab160a0fac68eb666ccaa558f21c26f4dd36a9423228

Observation fb8853da-f817-41b1-9c71-204d10ba00fc · outbound

This paper cites 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.868030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.868030Z digest=sha256:4d6c924b9ff2b83f12731ddf3c69df4596c8fc18ed3c300c120dd8550ce6ae86

Observation ba620139-edf7-4a7e-ab42-c056edbf3db2 · outbound

This paper cites 3D-TAFS: A Training-free Framework for 3D Affordance Segmentation.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model 3D-TAFS: A Training-free Framework for 3D Affordance Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.872664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.872664Z digest=sha256:c5977201c62960a58a89202121ba87cb73032a61a9b5ed07ca3acfc82c3d7772

Observation faa3f16d-433c-414b-b3fc-4b4c039def66 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.810474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.877203Z digest=sha256:d1ac80428aab2f60d17835281786e27cb85923e795c8ca8a694b8a6d16a2c344

Observation 9659b92e-b005-43ea-a0cb-c595f645da6f · outbound

This paper cites Scene- fun3d: fine-grained functionality and affordance understand- ing in 3d scenes.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Scene- fun3d: fine-grained functionality and affordance understand- ing in 3d scenes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.800290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.880848Z digest=sha256:4fb6eafc7812119a0579f80e99b311befa2e3169c0481b046f50d725c8adbbb8

Observation 4ab097ee-b741-4d9e-aafd-71af3453c2cd · outbound

This paper cites 3d affordancenet: A benchmark for visual object affordance understanding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model 3d affordancenet: A benchmark for visual object affordance understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.790236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.884814Z digest=sha256:5f130fee896366e9247e683fe6103798d345f82244b24bc50111a50d7b27465b

Observation ac39b2a5-e6fb-4928-bf02-2139eb936263 · outbound

This paper cites Learning 2d invariant affordance knowledge for 3d affordance ground- ing.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Learning 2d invariant affordance knowledge for 3d affordance ground- ing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.888579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.888579Z digest=sha256:3abf79376f87d8b3f51e24a343778578e556275133ceddc9893bec5663965e28

Observation cf87460f-5780-472e-ac89-754453a834ab · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model 3d-llm: Injecting the 3d world into large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.891868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.891868Z digest=sha256:ddfc49e744acad3f7d2538770caaa8774f0e01d1c0084c2c647f6f3a4e7bfab5

Observation 9797d96f-474d-4909-97fc-7b433f886771 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.895445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.895445Z digest=sha256:28ed2936db42cd1eda5e5d066f65cfb2ae254e0814a2e39fc876fa8525a987f5

Observation 444fe85c-5a16-4e63-8f04-37f80e7cd82b · outbound

This paper cites Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot ma- nipulation.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot ma- nipulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.774147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.899343Z digest=sha256:7e4b6cbe4bff923f1061cbb63d6668cebefad8ec99749c9c163baa3ddd91127a

Observation 69943d3b-05b3-4f17-ba23-39d1e283c041 · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.763757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.903257Z digest=sha256:7426148ed555e6e862d6a9f5c00ea0b62087faa15dba4896d306b49191c9cdcb

Observation eeefe02c-7f12-4951-83f5-b56c7c5e341c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Lisa: Reasoning segmentation via large language model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.907546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.907546Z digest=sha256:ffc59edb26aff5444feb6c683a7528f6a9ebe1ef4a0517b93d63985cf6f14673

Observation adbd5ecc-11f6-4f97-90b6-493d3d0a3e6d · outbound

This paper cites One-shot open affordance learning with foundation models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model One-shot open affordance learning with foundation models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.746768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.911800Z digest=sha256:3ce48739e22f31feac0cbd0ad203bc707f07241c0613970d4e9b2ae38dd2fdeb

Observation eeff70d7-cba7-4186-9b64-adc2f9042cdd · outbound

This paper cites Referring transformer: A one-step approach to multi-task visual grounding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Referring transformer: A one-step approach to multi-task visual grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.915594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.915594Z digest=sha256:b201fc08214fe58624a271f830d566dab7d89792eb55cb968cb7e0a2d28fc984

Observation 8bf1e2e4-155e-41fa-8f70-c34c6940b220 · outbound

This paper cites Laso: Language-guided affordance seg- mentation on 3d object.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Laso: Language-guided affordance seg- mentation on 3d object

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.730704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.919805Z digest=sha256:b95b3c12942831ab127da5cf80619688214ee78dfbdc57ec3847bbe8d6a7d568

Observation 6bf88876-59e8-4037-9d2d-f0dde4f82d93 · outbound

This paper cites Gres: Gen- eralized referring expression segmentation.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Gres: Gen- eralized referring expression segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.719827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.923863Z digest=sha256:66900cff0e9a2614e37fb7080d3016617ffeaaa1a472692fda55562c5fa079f5

Observation 49669313-8f62-4f44-adcf-60bf0c46eba9 · outbound

This paper cites Visual instruction tuning.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Visual instruction tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.927144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.927144Z digest=sha256:420206e2b22ad06337ab8c22688cb158deb85d08922165f10387585f4bd5a415

Observation 646cfba1-47ec-4448-a8de-f6e128e61dcf · outbound

This paper cites Openshape: Scaling up 3d shape representation towards open-world understanding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Openshape: Scaling up 3d shape representation towards open-world understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.702372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.930917Z digest=sha256:20089e2280bc6811638b45d734684c447968bde663f56f6dd4b31dedb0031eaa

Observation bae82463-d3b1-490d-90aa-0b6d388c2c04 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.934026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.934026Z digest=sha256:1155c92ba2c93b3dbbc61ae37bb64bb6b2fda5805993b60ded5a0593bead645b

Observation 8a180739-6da4-4d7b-ba22-f36db095805f · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.938006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.938006Z digest=sha256:2681b8920ec25b399d20f21829b29b808372049f60c6f42e828880d48bb5c22f

Observation 4b540fe5-7ffd-4e99-998e-1a25ac8b0bdf · outbound

This paper cites Auc: a misleading measure of the performance of predictive distribution models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Auc: a misleading measure of the performance of predictive distribution models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.691382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.941867Z digest=sha256:f240507db90770d58cf8a3ae7cf882c298172710543c54581628c7ea50d178e9

Observation 4c78e5e9-fee6-4399-93c8-2f7dce28c78f · outbound

This paper cites Decoupled Weight Decay Regularization.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Decoupled Weight Decay Regularization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.945876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.945876Z digest=sha256:667ab78079022205833e29a372afce784888f7416e0bc4deb5292dba780a3de8

Observation bf8cfec3-38cd-4bf5-b2c9-5db1fb9ca42d · outbound

This paper cites GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.949625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.949625Z digest=sha256:a01ed934b173c1f4216181d3ffe5e83afb6cd17011deee1cd5990a1ea26715f1

Observation deef5323-929c-4b82-b5b3-8bd734e117ce · outbound

This paper cites 3d-sps: Single-stage 3d visual grounding via referred point progressive selection.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model 3d-sps: Single-stage 3d visual grounding via referred point progressive selection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.679691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.954442Z digest=sha256:e86a81132481bb3660c71a184adfe9f7a4b1f3c8064b42d6d0e52ce1c4319198

Observation 8ffaf169-f5f0-4703-abc9-638c71c4321b · outbound

This paper cites Partnet: A large- scale benchmark for fine-grained and hierarchical part-level 3d object understanding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Partnet: A large- scale benchmark for fine-grained and hierarchical part-level 3d object understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.668687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.957852Z digest=sha256:ee9ebe9c1f99fd1d30835664e82c5c428e0beb60ebfb4fbccdaadb8dfdfb6912

Observation 5a12ca22-527f-420c-9ecb-d9d7196c618a · outbound

This paper cites O2o-afford: Annotation-free large-scale object- object affordance learning.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model O2o-afford: Annotation-free large-scale object- object affordance learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.656116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.962387Z digest=sha256:39de2b9fd4067dabd060bce7cbc1e1c203ba970a637f45d8ed24cf7582caca1c

Observation d5376ed9-0cf7-4a83-a290-622e4646770e · outbound

This paper cites RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.966607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.966607Z digest=sha256:1477a3512ea7e03be97b784eb3f98bffb06c516bf3384b4cca4ed0f1882c9840

Observation 50671ee5-14f9-40c1-a36a-229c1c433451 · outbound

This paper cites Open-vocabulary affordance detection in 3d point clouds.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Open-vocabulary affordance detection in 3d point clouds

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.645548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.971256Z digest=sha256:05a700d0ee8858d5baf5eefdbafa5fb523b5b285b2a1c06fa4e3fa99d03210c1

Observation cfad631b-d865-42f1-84f1-e1ae302718c2 · outbound

This paper cites Where2explore: Few-shot affordance learning for unseen novel categories of articulated objects.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Where2explore: Few-shot affordance learning for unseen novel categories of articulated objects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.636149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.974685Z digest=sha256:3a7b5a43a93bf8d543310573de6b877efb195c5cc5af063fc237e457875c4568

Observation 68a1afd0-09ae-4652-8a1f-3f49512cda7d · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.978024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.978024Z digest=sha256:875691b5fa182062d1a9d076db2d1b28fdda68a4223713702eb037a161a784a9

Observation 34129aa7-d87b-478f-b684-de823d1b2d94 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.619748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.981553Z digest=sha256:fb458cde04725a8e8ccd927a8994734e7a087342ffdb552a2ba15d0c26b54e2b

Observation 68c35e8d-34d8-43b9-a628-0deea8da1a67 · outbound

This paper cites Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.607948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.985752Z digest=sha256:9c9cea527cf65159818ab753d6ec94c3ad66485c194a7e3ac451b7a5f075333a

Observation 04be1169-c03b-451b-b780-fdaffc1f3d72 · outbound

This paper cites ShapeLLM: Universal 3D Object Understanding for Embodied Interaction.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model ShapeLLM: Universal 3D Object Understanding for Embodied Interaction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.989163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.989163Z digest=sha256:7ac61494322e7b1638d94e93344762a714d749dc4c4b85fdb31b973db16ea4b8

Observation ff3c5759-9ed8-4cc2-a132-7188a2bd74c3 · outbound

This paper cites Affordancellm: Grounding affordance from vision language models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Affordancellm: Grounding affordance from vision language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.597166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:03.995473Z digest=sha256:8f2936e85db2ce156878b186bda5fdcf35a6b8e5ad478d440c2397a1b2270cc7

Observation 94cb39af-65f2-439f-9e8f-587d903f6120 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.999478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.999478Z digest=sha256:51736842a54fde744eae0f0f373c0215bc9fbc50fc59b17fbb12bffd0a6d3ab5

Observation 74d02153-9498-411f-a560-79a114b7f42d · outbound

This paper cites Optimizing intersection- over-union in deep neural networks for image segmentation.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Optimizing intersection- over-union in deep neural networks for image segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.579377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.003391Z digest=sha256:8f2af7f599a5aa036df06191386636939814d576181e8754353430a2be2747d3

Observation 47ba5652-44dd-42a8-a062-adf68c30b165 · outbound

This paper cites GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.006460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.006460Z digest=sha256:4854417207919aef1e712362d3098610d9fbec1c5f5af5af6a2d9ecf0bacdeaf

Observation 21476325-1ecd-4d30-85a7-f6deaa41a147 · outbound

This paper cites Color indexing.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Color indexing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.568404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.009844Z digest=sha256:d9473d491fd9f05c6634777a654d7f16dc1519c8e2f48bf77442147fdd67d845

Observation a00a30f4-ee81-498e-8a6d-585f8fc04bbb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.013237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.013237Z digest=sha256:7b1b1b4fd5b09947798f470c8a1a6a1965f04950cc373898d59dbd433a53d012

Observation 790adcfb-82c4-43d9-9b4e-9ea0058dcc99 · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Visionllm: Large language model is also an open- ended decoder for vision-centric tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.016706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.016706Z digest=sha256:438a74ba2112f5621e2ba44ef89576916f5b784ae2d3c8d475ba96fe019fee13

Observation d8947f22-a89c-46ae-a9a5-02726637a8f5 · outbound

This paper cites Dynamic graph cnn for learning on point clouds.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Dynamic graph cnn for learning on point clouds

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.550904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.019849Z digest=sha256:42f23b5f5b10d7b7e2df03769541961937a85bf7a09d43713bb6f68477fd0ff0

Observation 35681b06-9f66-45a8-b805-4fb98dd501fe · outbound

This paper cites Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.539977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.023477Z digest=sha256:af25b4c569a82dde447c0ee11c3647d2bae9f80f31cc6b1d8eb407d8ecb34ed7

Observation c05242bf-45db-48ad-9743-6e7cc79538d2 · outbound

This paper cites Learning environment-aware affor- dance for 3d articulated object manipulation under occlu- sions.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Learning environment-aware affor- dance for 3d articulated object manipulation under occlu- sions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.529682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.027108Z digest=sha256:75601354e799b481e4f8a0d1ee127013cd869555f6c3f741a68fed864aca1374

Observation 67b0b7ab-fe0e-47c4-8985-11ff6c8f4cb8 · outbound

This paper cites AffordDP: Generalizable Diffusion Policy with Transferable Affordance.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model AffordDP: Generalizable Diffusion Policy with Transferable Affordance

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.030516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.030516Z digest=sha256:2d398849a4291b65b86a651b35d33de506369c90a30bddce000cf6ec3b37d469

Observation a00cb060-c011-4fa1-826a-1702af9ba9bc · outbound

This paper cites PartAfford: Part-level Affordance Discovery from 3D Objects.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model PartAfford: Part-level Affordance Discovery from 3D Objects

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.034679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.034679Z digest=sha256:915abcbfe9b8970503224849393330257ae867358b4f9f1574b919026e8a2427

Observation f78edc63-97e3-40f5-8742-2c85b156b9a5 · outbound

This paper cites Weakly-supervised affordance grounding guided by part-level semantic priors.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Weakly-supervised affordance grounding guided by part-level semantic priors

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.518779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.038666Z digest=sha256:03108156bfdc27eadeb1717845a866fee4ccc57ecef0fd61c5342d21861f3888

Observation 41a64d97-2b2f-4ee0-a5ce-84dfd9281f06 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.042646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.042646Z digest=sha256:405b4bd2fc8f5eec99955e698a388f3225bb0a01882212619d93f839e5a77949

Observation 23d12af1-3457-46e9-84a6-21082957c3d2 · outbound

This paper cites Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.508196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.046371Z digest=sha256:b8bda1f7ea7d7a42571049301b663dd7b0f30eb31832d91c4ac4dae19c859017

Observation 44b14df7-2f30-471e-bc22-79e192377a99 · outbound

This paper cites Grounding 3d object affordance from 2d interactions in images.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Grounding 3d object affordance from 2d interactions in images

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.496613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.049340Z digest=sha256:590a937bc48e6cdd6df3986e8768099812073dbdc5d2d4e9356e230bee7d23be

Observation 46b23a12-9091-4d5a-b482-5f9e51f2648a · outbound

This paper cites Lemon: Learning 3d human-object interac- tion relation from 2d images.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Lemon: Learning 3d human-object interac- tion relation from 2d images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.486321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.052564Z digest=sha256:2384bc0701c2160ef84f43517710a8157620592851d8b31fc463fce145321877

Observation f69d034d-0738-4355-993f-2de9a35a7dec · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.055997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.055997Z digest=sha256:0882e56de040c603fde7353bd22cf9c9d0dea1837f175e3cb3f635712ec58b16

Observation 63edb73d-6da7-4125-9340-83cc90a7a3c0 · outbound

This paper cites Uni3d: A unified baseline for multi-dataset 3d object detection.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model Uni3d: A unified baseline for multi-dataset 3d object detection

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:22:04.475924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:22:04.059689Z digest=sha256:6e74af344a4e8fdd6fcbe65a9640744278dedf02cbfaa2d89f6f26e3967dcbc3

Observation a846904f-5e78-4b50-b971-f9ef41629ef4 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.063177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.063177Z digest=sha256:11b84caa287bb04b835458e718b005c49f5ecf10c4318acc96363c2e21426ad8

Observation cb685245-b92f-4eb9-a991-16602a54c21e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:04.066726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:04.066726Z digest=sha256:6e50563bcaa54c9b25662740d0cd27c3865ec34e574fd78ff0a53c4f863bf967

Pith citing papers

Observation 85b2a812-ccc1-438d-bc66-f05c37773df8 · inbound

AffordDP: Generalizable Diffusion Policy with Transferable Affordance cites this paper.

AffordDP: Generalizable Diffusion Policy with Transferable Affordance SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-11T22:49:09.116483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:49:08.968786Z digest=sha256:c69b9ccf707773b13168edf6762a2029cd2e7fcba48c2183f50bb1dcd4590f5b