REVIEW 3 major objections 3 minor 76 references
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read LM-AD routes the affordance word through BERT with cross-attention over point-cloud features, reporting 41.98 mIoU full-shape and 35.26 partial-view on 3D AffordanceNet.
desk verdict A sensible Flamingo-style transplant to 3D affordance detection with large reported gains, but the ablation evidence does not yet isolate what causes the improvement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Affordance Query Module (AQM) is a stack of blocks built on BERT, a pretrained bidirectional Transformer language model. Inside each of BERT's 12 layers, a cross-attention layer is inserted between self-attention and the feed-forward network, with the text tokens serving as queries and point-cloud features (PointNet++ output after a 1-D convolution and batch norm) serving as keys and values. Each block produces a text representation that has looked at the point cloud, and the final block's output is decoded by another cross-attention layer, where the point features query the aligned text features, followed by an MLP. This repeatedly fuses the two modalities instead of scoring them with a single cosine similarity.
What would settle it
Train the identical LM-AD pipeline with the AQM's 12 inserted cross-attention layers replaced by 12 plain cross-attention layers of matched width, parameter count, learning-rate schedule, and random seeds, and compare mIoU distributions over at least five seeds; if a plain baseline reaches or exceeds 41.98 mIoU on the full-shape split, the paper's attribution of the gain to AQM's LM-specific fusion is refuted.
Extended reading notes
Core claim
The paper's central claim is that a point-cloud affordance detector should not stop at a single cosine-similarity score between a text embedding and point embeddings. Instead, the text should be processed by a pretrained language model whose layers include cross-attention over the point cloud, so that each text token can interrogate local geometry and the final text features carry fine-grained alignment. LM-AD (Language Model-guided Affordance Detection) implements this with 12 BERT layers, each augmented with cross-attention over PointNet++ features, and reports mIoU of 41.98 on full-shape and 35.26 on partial-view 3D AffordanceNet, with accuracy 68.60/62.65 and mean accuracy 68.89/59.09. The paper also reports that swapping AQM for plain cross-attention drops full-shape mIoU to 19.38, and it presents this gap as evidence that the LM-based fusion is what drives the improvement.
Load-bearing premise
The main result depends on the assumption that the comparison to 'simple cross-attention' controls for everything except the AQM's design, so the measured 22.60-point mIoU gap is caused by fine-grained LM alignment rather than capacity, training schedule, or chance.
Editorial extensions
If this is right
- Full-shape mIoU on 3D AffordanceNet goes from 22.33 with OpenAD-KD to 41.98 with LM-AD, and partial-view mIoU from 20.48 to 35.26.
- Replacing AQM with plain cross-attention drops full-shape mIoU to 19.38, which the paper reads as evidence that the LM-based fusion module is responsible for most of the improvement.
- Accuracy reaches 68.60 full-shape and 62.65 partial-view, compared with the below-50% accuracy the paper cites for prior point-cloud methods, so the gain is not limited to the IoU metric.
- Both full-shape and partial-view tasks improve, and the paper highlights partial-view inputs as the realistic setting for a robot observing an object from limited viewpoints.
- Because the text input is assumed to be a single word following prior work, the architecture itself is compatible with longer text and could be evaluated on multi-word affordance descriptions without structural changes.
Reading between the lines
- If AQM's alignment pattern is the cause of the gain, the same 'text queries, geometry keys/values' insertion should transfer to other open-vocabulary 3D grounding tasks, such as language-driven part segmentation.
- The paper restricts text to a single affordance word, so the module's capacity for compositional or multi-word instructions remains untested; that is a natural next experiment.
- A matched-parameter, multi-seed replication would determine whether the 22.60-point ablation gap comes from fine-grained alignment or from added model capacity, since the current comparison does not control for those factors.
- The reuse of a frozen language model as a fusion engine suggests that expensive cross-modal training may not be needed for point-level language grounding, which could lower the cost of robotics perception systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LM-AD, a method for 3D point-cloud affordance detection built from a PointNet++ point encoder and a pretrained BERT text encoder. The core module, AQM, inserts a cross-attention layer into each BERT layer so that text hidden states attend to point-cloud features, followed by a final cross-attention from point-cloud features to the aligned text representation (Eqs. 1-3). Experiments on 3D AffordanceNet report large improvements over OpenAD and OpenAD-KD in mIoU, Acc, and mAcc for both full-shape and partial-view tasks (Table 1), and a two-row ablation (Table 2) credits AQM with a 22.60 mIoU improvement over a "simple cross-attention" baseline. The paper's central claim is that capturing fine-grained multimodal alignment is what drives these gains.
Significance. If the Table 1 results are reproducible, LM-AD would constitute a substantial empirical advance on a standard benchmark, roughly doubling the mIoU of the previous best method on the full-shape task. The architecture is simple and plausible, and Equations 1-3 describe a coherent forward pass. The evaluation uses an external benchmark with train/validation/test splits rather than a self-constructed setting, so the main comparison is not circular. However, the evidence for the mechanistic claim is currently thin: the ablation baseline is underspecified, no seed variance or confidence intervals are reported, and no code or hyperparameters are provided. These deficiencies make it difficult to determine whether the reported gains are due to the proposed alignment mechanism, to the pretrained language model, or to uncontrolled experimental factors.
major comments (3)
- [Section 5.3, Table 2] The ablation does not isolate the mechanism claimed in the title. Method (i) is described only as "simple cross-attention layers" with "the number of layers unchanged"; the text never states whether this baseline uses BERT, whether BERT is frozen or fine-tuned, how many parameters it has, or how it is trained. If method (i) drops BERT, the 22.60 mIoU gap conflates the pretrained LM's capacity and linguistic priors with AQM's interleaved cross-attention. If method (i) retains BERT, the comparison still varies the fusion architecture jointly with unspecified training choices. A control that keeps BERT fixed and varies only the alignment structure (e.g., one cross-attention at the output versus interleaved cross-attention in each layer) is needed to support the central claim.
- [Section 5.1-5.2, Table 1] The comparison with baselines is not fully controlled as reported. The paper does not state whether the OpenAD, OpenAD-KD, and ZSLPC numbers in Table 1 were re-run under the same 70/10/20 split, point sampling, and training schedule, or whether they are quoted from the original papers. If the numbers are quoted, differences in data preprocessing or evaluation protocol could affect the comparison. The authors should report the exact protocol for every method in Table 1, or re-evaluate all baselines in their own framework.
- [Section 5.1] No experimental configuration is reported: learning rate, batch size, optimizer, number of epochs, number of input points N, maximum token length L, point-cloud normalization, and BERT fine-tuning strategy are all omitted. No multiple-seed results or confidence intervals are given for any table entry. For an empirical paper whose contribution is a performance claim, these details are necessary for verification and for assessing whether the large margins are stable.
minor comments (3)
- [Section 4.2, Eq. (2)] Please define the input to the next AQM block explicitly (e.g., x^(i+1) = g^(i)) and state where residual connections and layer norms are applied; as written, the recursive definition of the block is incomplete.
- [Section 2.3, Ref. [18]] Reference [18] is cited as the source for BLIP-2, but the BLIP-2 paper appears to be reference [71] (Li et al.); please correct the citation and verify whether the bibliographic entry "InstructBLIP 2" is accurate.
- [Section 3] The notation x_txt ∈ {1,0}^{V×L} suggests a one-hot sequence; it would be clearer to say that the input is a token sequence of length L over a vocabulary of size V, since the actual input to BERT is token ids rather than one-hot vectors.
Circularity Check
No circularity: LM-AD is an empirical benchmark comparison; the ablation confound is an experimental-control issue, not a definitional reduction.
full rationale
The paper contains no claimed derivation of a prediction from a fitted input. LM-AD is evaluated on the external 3D AffordanceNet benchmark with standard 70/10/20 splits, and the main results (Table 1) are measured mIoU, Acc, and mAcc against published baselines. The ablation in Section 5.3 compares AQM (method ii) with "simple cross-attention layers" (method i) while keeping the number of layers unchanged; even if this baseline is underspecified and may conflate the pretrained BERT backbone with the alignment mechanism, that is a limitation of the experimental control, not a circularity. The reported 22.60 mIoU gap is an empirical difference, not a quantity that equals its own input by construction. BERT is invoked as an external pretrained model, and BLIP-2 and Flamingo are cited as prior inspiration rather than as load-bearing uniqueness theorems or self-citations from the present authors. No fitted parameter is renamed as a prediction, no result is forced by a self-citation chain, and the method is benchmarked against an external dataset, so the derivation chain is self-contained. The circularity score is therefore 0.
Assumptions & free parameters
free parameters (3)
- Training hyperparameters (unreported)
- AQM cross-attention configuration (unreported)
- Maximum token length and point count (unreported)
assumptions (4)
- domain assumption A pretrained LM with inserted cross-attention can transfer its linguistic priors to condition on 3D point cloud features.
- domain assumption The 70/10/20 split by shape semantic category is a valid zero-shot protocol comparable to prior baselines.
- domain assumption Standard optimization with cross-entropy loss will train the inserted modules without breaking BERT's useful representations.
- domain assumption PointNet++ is a sufficient point cloud encoder for affordance detection.
Cite this review
Pith. "Pith review of Capturing Fine-Grained Alignments Improves 3D Affordance Detection." pith.science (2026). https://pith.science/paper/WPGRLQF7
@misc{pith2026250619312,
author = {Pith},
title = {Pith review of: Capturing Fine-Grained Alignments Improves 3D Affordance Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/WPGRLQF7}},
note = {Machine review of arXiv:2506.19312}
}
read the original abstract
In this work, we address the challenge of affordance detection in 3D point clouds, a task that requires effectively capturing fine-grained alignments between point clouds and text. Existing methods often struggle to model such alignments, resulting in limited performance on standard benchmarks. A key limitation of these approaches is their reliance on simple cosine similarity between point cloud and text embeddings, which lacks the expressiveness needed for fine-grained reasoning. To address this limitation, we propose LM-AD, a novel method for affordance detection in 3D point clouds. Moreover, we introduce the Affordance Query Module (AQM), which efficiently captures fine-grained alignment between point clouds and text by leveraging a pretrained language model. We demonstrated that our method outperformed existing approaches in terms of accuracy and mean Intersection over Union on the 3D AffordanceNet dataset.
Figures
Reference graph
Works this paper leans on
-
[18]
InstructBLIP 2: Extending Vision- Language Models with Fine-Grained Instruction Tun- ing,
H. Chen and T. Xu, “InstructBLIP 2: Extending Vision- Language Models with Fine-Grained Instruction Tun- ing,”IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 2023
work page 2023
-
[22]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark,et al., “Training compute-optimal large language models,”arXiv preprint arXiv:2203.15556, 2022
arXiv 2022
-
[1]
J. Jiang, G. Cao, T.-T. Do, and S. Luo, “A4T: Hi- erarchical Affordance Detection for Transparent Ob- jects Depth Reconstruction and Manipulation,”IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 9826– 9833, 2022
work page 2022
-
[2]
Deep Affordance-Grounded Sensorimotor Ob- ject Recognition,
S. Thermos, G. T. Papadopoulos, P. Daras, and G. Potami- anos, “Deep Affordance-Grounded Sensorimotor Ob- ject Recognition,” inProceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition (CVPR), July 2017
work page 2017
-
[3]
Af- fordance Transfer Learning for Human-Object Inter- action Detection,
Z. Hou, B. Yu, Y. Qiao, X. Peng, and D. Tao, “Af- fordance Transfer Learning for Human-Object Inter- action Detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR), pp. 495–504, June 2021
work page 2021
-
[4]
Pre- dicting 3D Human Dynamics From Video,
J. Zhang, P. Felsen, A. Kanazawa, and J. Malik, “Pre- dicting 3D Human Dynamics From Video,” in2019 IEEE/CVF International Conference on Computer Vi- sion (ICCV), pp. 7113–7122, 2019
work page 2019
-
[5]
Predicting hu- man activities using stochastic grammar,
S. Qi, S. Huang, P. Wei, and S.-C. Zhu, “Predicting hu- man activities using stochastic grammar,” inProceed- ings of the IEEE International Conference on Com- puter Vision, pp. 1164–1172, 2017
work page 2017
-
[6]
Af- fordance grounding from demonstration video to tar- get image,
J. Chen, D. Gao, K. Q. Lin, and M. Z. Shou, “Af- fordance grounding from demonstration video to tar- get image,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pp. 6799–6808, 2023
work page 2023
Show all 76 references
-
[7]
Affor- dance Research in Developmental Robotics: A Sur- vey,
H. Min, C. Yi, R. Luo, J. Zhu, and S. Bi, “Affor- dance Research in Developmental Robotics: A Sur- vey,”IEEE Transactions on Cognitive and Develop- mental Systems, vol. 8, no. 4, pp. 237–255, 2016
2016
-
[8]
A survey of visual affordance recognition based on deep learning,
D. Chen, D. Kong, J. Li, S. Wang, and B. Yin, “A survey of visual affordance recognition based on deep learning,”IEEE Transactions on Big Data, vol. 9, no. 6, pp. 1458–1476, 2023
2023
-
[9]
Visual affor- dance and function understanding: A survey,
M. Hassanin, S. Khan, and M. Tahtali, “Visual affor- dance and function understanding: A survey,”ACM Computing Surveys (CSUR), vol. 54, no. 3, pp. 1–35, 2021
2021
-
[10]
3d af- fordancenet: A benchmark for visual object affordance understanding,
S. Deng, X. Xu, C. Wu, K. Chen, and K. Jia, “3d af- fordancenet: A benchmark for visual object affordance understanding,” inproceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pp. 1778–1787, 2021
2021
-
[11]
Open-vocabulary affordance detection in 3d point clouds,
T. Nguyen, M. N. Vu, A. Vuong, D. Nguyen, T. Vo, N. Le, and A. Nguyen, “Open-vocabulary affordance detection in 3d point clouds,” in2023 IEEE/RSJ In- ternational Conference on Intelligent Robots and Sys- tems (IROS), pp. 5692–5698, IEEE, 2023
2023
-
[12]
3D ShapeNets: A Deep Representation for Volumetric Shapes,
Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3D ShapeNets: A Deep Representation for Volumetric Shapes,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2015
2015
-
[13]
Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World Data,
M. A. Uy, Q.-H. Pham, B.-S. Hua, T. Nguyen, and S.- K. Yeung, “Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World Data,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019
2019
-
[14]
Open-vocabulary affordance detection using knowledge distillation and text-point correlation,
T. Van Vo, M. N. Vu, B. Huang, T. Nguyen, N. Le, T. Vo, and A. Nguyen, “Open-vocabulary affordance detection using knowledge distillation and text-point correlation,” in2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 13968–13975, IEEE, 2024
2024
-
[15]
Transductive zero-shot learning for 3d point cloud classification,
A. Cheraghian, S. Rahman, D. Campbell, and L. Pe- tersson, “Transductive zero-shot learning for 3d point cloud classification,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 923–933, 2020
2020
-
[16]
Generative zero-shot learning for semantic seg- mentation of 3d point clouds,
B. Michele, A. Boulch, G. Puy, M. Bucher, and R. Mar- let, “Generative zero-shot learning for semantic seg- mentation of 3d point clouds,” in2021 International Conference on 3D Vision (3DV), pp. 992–1002, IEEE, 2021
2021
-
[17]
Zero- shot learning of 3d point cloud objects,
A. Cheraghian, S. Rahman, and L. Petersson, “Zero- shot learning of 3d point cloud objects,” in2019 16th International Conference on Machine Vision Applica- tions (MV A), pp. 1–6, IEEE, 2019
2019
-
[19]
Flamingo: A Visual Language Model for Few-Shot Learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barrei...
2022
-
[20]
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need.”
-
[21]
BERT: Pre-training of Deep Bidirectional Transform- ers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transform- ers for Language Understanding,” 2019
2019
-
[23]
Detecting object affordances with Convo- lutional Neural Networks,
A. Nguyen, D. Kanoulas, D. G. Caldwell, and N. G. Tsagarakis, “Detecting object affordances with Convo- lutional Neural Networks,” in2016 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pp. 2765–2770, 2016
2016
-
[24]
Affordancenet: An end-to-end deep learning approach for object af- fordance detection,
T.-T. Do, A. Nguyen, and I. Reid, “Affordancenet: An end-to-end deep learning approach for object af- fordance detection,” in2018 IEEE international con- ference on robotics and automation (ICRA), pp. 5882– 5889, IEEE, 2018
2018
-
[25]
Object-based affordances detection with Convolutional Neural Networks and dense Conditional Random Fields,
A. Nguyen, D. Kanoulas, D. G. Caldwell, and N. G. Tsagarakis, “Object-based affordances detection with Convolutional Neural Networks and dense Conditional Random Fields,” in2017 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS), pp. 5908– 5915, 2017
2017
-
[26]
A multi-scale cnn for af- fordance segmentation in rgb images,
A. Roy and S. Todorovic, “A multi-scale cnn for af- fordance segmentation in rgb images,” inComputer Vision–ECCV 2016: 14th European Conference, Am- sterdam, The Netherlands, October 11–14, 2016, Pro- ceedings, Part IV 14, pp. 186–201, Springer, 2016
2016
-
[27]
A deep learning approach to object affordance segmentation,
S. Thermos, P. Daras, and G. Potamianos, “A deep learning approach to object affordance segmentation,” inICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2358–2362, IEEE, 2020
2020
-
[28]
Cerberus transformer: Joint semantic, affordance and attribute parsing,
X. Chen, T. Liu, H. Zhao, G. Zhou, and Y.-Q. Zhang, “Cerberus transformer: Joint semantic, affordance and attribute parsing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pp. 19649–19658, 2022
2022
-
[29]
Learning affordance grounding from exocentric im- ages,
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao, “Learning affordance grounding from exocentric im- ages,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2252–2261, 2022
2022
-
[30]
Affordancellm: Grounding affordance from vision language models,
S. Qian, W. Chen, M. Bai, X. Zhou, Z. Tu, and L. E. Li, “Affordancellm: Grounding affordance from vision language models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pp. 7587–7597, 2024
2024
-
[31]
Visual instruc- tion tuning,
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruc- tion tuning,”Advances in neural information process- ing systems, vol. 36, pp. 34892–34916, 2023
2023
-
[32]
Semantic labeling of 3d point clouds with object affordance for robot ma- nipulation,
D. I. Kim and G. S. Sukhatme, “Semantic labeling of 3d point clouds with object affordance for robot ma- nipulation,” in2014 IEEE International Conference on Robotics and Automation (ICRA), pp. 5578–5584, IEEE, 2014
2014
-
[33]
Affordance detection for task-specific grasping using deep learning,
M. Kokic, J. A. Stork, J. A. Haustein, and D. Kragic, “Affordance detection for task-specific grasping using deep learning,” in2017 IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids), pp. 91– 98, IEEE, 2017
2017
-
[34]
Pointnet: Deep learning on point sets for 3d classification and segmentation,
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” inProceedings of the IEEE conference on computer vision and pattern recognition, pp. 652– 660, 2017
2017
-
[35]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space,
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,”Advances in neural information process- ing systems, vol. 30, 2017
2017
-
[36]
Dynamic graph cnn for learning on point clouds,
Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bron- stein, and J. M. Solomon, “Dynamic graph cnn for learning on point clouds,”ACM Transactions on Graph- ics (tog), vol. 38, no. 5, pp. 1–12, 2019
2019
-
[37]
Point Transformer,
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun, “Point Transformer,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 16259–16268, October 2021
2021
-
[38]
A robustly opti- mized BERT pre-training approach with post-training,
Z. Liu, W. Lin, Y. Shi, and J. Zhao, “A robustly opti- mized BERT pre-training approach with post-training,” inChina national conference on Chinese computational linguistics, pp. 471–484, Springer, 2021
2021
-
[39]
Deberta: Decoding- enhanced bert with disentangled attention,
P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding- enhanced bert with disentangled attention,”arXiv preprint arXiv:2006.03654, 2020
2006 arXiv
-
[40]
Learning Transferable Visual Models From Natural Language Supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” inProceedings of the 38th International Conference on Machine Le...
2021
-
[41]
Multimodal Alignment and Fu- sion: A Survey,
S. Li and H. Tang, “Multimodal Alignment and Fu- sion: A Survey,”arXiv preprint arXiv:2411.17040, 2024
2024
-
[42]
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,” 2021
2021
-
[43]
Multimodal repre- sentation learning for tourism recommendation with two-tower architecture,
Y. Cui, S. Liang, and Y. Zhang, “Multimodal repre- sentation learning for tourism recommendation with two-tower architecture,”PLoS One, 2024
2024
-
[44]
I can listen but cannot read: An evaluation of two-tower multi- modal systems for instrument recognition,
Y. Vasilakis, R. Bittner, and J. Pauwels, “I can listen but cannot read: An evaluation of two-tower multi- modal systems for instrument recognition,” 2024
2024
-
[45]
Bridgetower: Building bridges between encoders in vision-language representation learning,
X. Xu, C. Wu, S. Rosenman, V. Lal, and W. Che, “Bridgetower: Building bridges between encoders in vision-language representation learning,” inProceed- ings of the AAAI Conference on Artificial Intelligence, 2023
2023
-
[46]
Be- yond Two-Tower Matching: Learning Sparse Retriev- able Cross-Interactions for Recommendation,
L. Su, F. Yan, J. Zhu, X. Xiao, and H. Duan, “Be- yond Two-Tower Matching: Learning Sparse Retriev- able Cross-Interactions for Recommendation,” inPro- ceedings of the 46th International ACM SIGIR Con- ference on Research and Development in Information Retrieval, 2023
2023
-
[47]
Touchformer: A Transformer-based two-tower architecture for tactile temporal signal classification,
C. Liu, H. Liu, H. Chen, and W. Du, “Touchformer: A Transformer-based two-tower architecture for tactile temporal signal classification,”IEEE Transactions on Multimedia, 2023
2023
-
[48]
Mix-tower: Light visual question answering framework based on exclusive self-attention mechanism,
D. Chen, J. Chen, L. Yang, and F. Shang, “Mix-tower: Light visual question answering framework based on exclusive self-attention mechanism,”Neurocomputing, 2024
2024
-
[49]
Towards artificial general intelligence via a multimodal foundation model,
N. Fei, Z. Lu, Y. Gao, G. Yang, Y. Huo, J. Wen, and H. Lu, “Towards artificial general intelligence via a multimodal foundation model,”Nature Communica- tions, 2022
2022
-
[50]
Multimodal Reranking for Knowledge- Intensive Visual Question Answering,
H. Wen, H. Zhuang, H. Zamani, A. Hauptmann, and M. Bendersky, “Multimodal Reranking for Knowledge- Intensive Visual Question Answering,” 2024
2024
-
[51]
Dif- ferentiable cross-modal hashing via multimodal trans- formers,
J. Tu, X. Liu, R. Lin, Z.and Hong, and M. Wang, “Dif- ferentiable cross-modal hashing via multimodal trans- formers,” inProceedings of the 30th ACM Interna- tional Conference on Multimedia, 2022
2022
-
[52]
Towards User Friendly Medication Mapping Using Entity-Boosted Two-Tower Neural Network,
S. Yuan, P. Bhatia, B. Celikkaya, H. Liu, and K. Choi, “Towards User Friendly Medication Mapping Using Entity-Boosted Two-Tower Neural Network,” inIn- ternational Workshop on Deep Learning for Human Activity Recognition, 2021
2021
-
[53]
Fusing information from multifidelity computer models of physical sys- tems,
D. L. Allaire and K. E. Willcox, “Fusing information from multifidelity computer models of physical sys- tems,”2012 15th International Conference on Infor- mation Fusion, pp. 2458–2465, 2012
2012
-
[54]
Seg- Net: A Deep Convolutional Encoder-Decoder Archi- tecture for Image Segmentation,
V. Badrinarayanan, A. Kendall, and R. Cipolla, “Seg- Net: A Deep Convolutional Encoder-Decoder Archi- tecture for Image Segmentation,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, pp. 2481–2495, 2015
2015
-
[55]
Sensor fusion of camera and LiDAR raw data for vehicle detection,
G. Danapal, G. A. Santos, J. P. C. L. da Costa, B. J. G. Praciano, and G. P. M. Pinheiro, “Sensor fusion of camera and LiDAR raw data for vehicle detection,” in2020 Workshop on Communication Networks and Power Systems (WCNPS), pp. 1–6, 2020
2020
-
[56]
A Model-Level Fusion-Based Multi-Modal Object Detection and Recognition Method,
C. Guo and L. Zhang, “A Model-Level Fusion-Based Multi-Modal Object Detection and Recognition Method,” in2023 7th Asian Conference on Artificial Intelligence Technology (ACAIT), pp. 34–38, 2023
2023
-
[57]
Learning to combine local models for facial Action Unit detec- tion,
S. Jaiswal, B. Mart ´ ınez, and M. F. Valstar, “Learning to combine local models for facial Action Unit detec- tion,”2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recogni- tion (FG), vol. 06, pp. 1–6, 2015
2015
-
[58]
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning,
Y. Wada, K. Kaneda, D. Saito, and K. Sugiura, “Polos: Multimodal Metric Learning from Human Feedback for Image Captioning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024
2024
-
[59]
Hierarchical Feature Fusion Network for Salient Object Detection,
X. Li, D. Song, and Y. Dong, “Hierarchical Feature Fusion Network for Salient Object Detection,”IEEE Transactions on Image Processing, vol. 29, pp. 9165– 9175, 2020
2020
-
[60]
DenseFuse: A Fusion Approach to Infrared and Visible Images,
H. Li and X. Wu, “DenseFuse: A Fusion Approach to Infrared and Visible Images,”IEEE Transactions on Image Processing, vol. 28, pp. 2614–2623, 2018
2018
-
[61]
Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affec- tive Computing,
S. Mai, H. Hu, and S. Xing, “Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affec- tive Computing,” inAnnual Meeting of the Association for Computational Linguistics, 2019
2019
-
[62]
A hierarchical feature fusion frame- work for adaptive visual tracking,
A. Makris, D. I. Kosmopoulos, S. J. Perantonis, and S. Theodoridis, “A hierarchical feature fusion frame- work for adaptive visual tracking,”Image Vis. Com- put., vol. 29, pp. 594–606, 2011
2011
-
[63]
Model level fusion of edge histogram descriptors and gabor wavelets for landmine detection with ground penetrating radar,
O. Missaoui, H. Frigui, and P. D. Gader, “Model level fusion of edge histogram descriptors and gabor wavelets for landmine detection with ground penetrating radar,” 2010 IEEE International Geoscience and Remote Sens- ing Symposium, pp. 3378–3381, 2010
2010
-
[64]
Towards Raw Sensor Fu- sion in 3D Object Detection,
A. R¨ ovid and V. Remeli, “Towards Raw Sensor Fu- sion in 3D Object Detection,”2019 IEEE 17th World Symposium on Applied Machine Intelligence and In- formatics (SAMI), pp. 293–298, 2019
2019
-
[65]
Design of a Low-Level Radar and Time-of-Flight Sen- sor Fusion Framework,
J. Steinbaeck, C. Steger, G. Holweg, and N. Druml, “Design of a Low-Level Radar and Time-of-Flight Sen- sor Fusion Framework,”2018 21st Euromicro Confer- ence on Digital System Design (DSD), pp. 268–275, 2018
2018
-
[66]
Guided Deep Decoder: Unsupervised Image Pair Fusion,
T. Uezato, D. Hong, N. Yokoya, and W. He, “Guided Deep Decoder: Unsupervised Image Pair Fusion,” in European Conference on Computer Vision, 2020
2020
-
[67]
Decision-Level Data Fusion in Quality Control and Predictive Main- tenance,
Y. Wei, D. Wu, and J. P. Terpenny, “Decision-Level Data Fusion in Quality Control and Predictive Main- tenance,”IEEE Transactions on Automation Science and Engineering, vol. 18, pp. 184–194, 2021
2021
-
[68]
ViLT: Vision and lan- guage transformer without convolution or region su- pervision
W. Kim, B. Son, and I. Kim, “ViLT: Vision and lan- guage transformer without convolution or region su- pervision.”
-
[69]
VLMo: Unified Vision-Language Pre-Training with Mixture of Modal- ity Experts,
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, and F. Wei, “VLMo: Unified Vision-Language Pre-Training with Mixture of Modal- ity Experts,” 2022
2022
-
[70]
BLIP: Boot- strapping language image pre-training for unified vi- sion language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Boot- strapping language image pre-training for unified vi- sion language understanding and generation,” inPro- ceedings of the 39th International Conference on Ma- chine Learning, pp. 12888–12900, PMLR, 2022
2022
-
[71]
BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” 2023
2023
-
[72]
InstructBLIP: Towards General-purpose Vision-Language Models with Instruc- tion Tuning,
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “InstructBLIP: Towards General-purpose Vision-Language Models with Instruc- tion Tuning,” 2023
2023
-
[73]
Minigpt-4: Enhancing vision-language understanding with advanced large language models,
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,”arXiv preprint arXiv:2304.10592, 2023
2023 arXiv
-
[74]
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision,
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “SimVLM: Simple Visual Language Model Pretraining with Weak Supervision,” 2022
2022
-
[75]
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution,
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin, “Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution,” 2024
2024
-
[76]
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localiza- tion, Text Reading, and Beyond,
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-VL: A Versatile Vision-Language Model for Understanding, Localiza- tion, Text Reading, and Beyond,” 2023
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.