Pith. sign in

REVIEW 3 major objections 6 minor 221 references

Data Transformation Strategies to Remove Heterogeneity

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Survey organizes data transformation strategies by format conflict

desk verdict A useful survey of data-to-text and data-to-graph transformations that over-claims its scope and needs cleanup before it can be a reliable reference. read the letter →

arxiv 2507.12677 v1 pith:4E3OZBAC submitted 2025-07-16 cs.LG cs.AI

classification cs.LGcs.AI
keywords dataheterogeneitytransformationformatconflictdata-to-textdata-to-graphtable-to-textknowledgegraphsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that data heterogeneity should be tackled at the format level, not only at the schema or data-value level, and that data transformation deserves a place alongside data integration and data cleaning as a first-class remedy. It claims that format conflicts among table, text, image, video, and numerical data can be systematically handled by converting data into text, graph, or visual formats, and it organizes the main methods for each conversion path. The reason to care is practical: training and using modern AI models requires matching input formats, and choosing the wrong transformation can destroy information. The paper's contribution is an organizing map: definitions of heterogeneity by conflict type, a strategy taxonomy, and a list of datasets and open problems per transformation route.

What carries the argument

The organizing device is a two-dimensional scheme: conflict factors (schema, data, format, domain) determine the remedy family (integration, type conversion, cleaning, transformation), and the source-target format pair determines the strategy category. Within that scheme, the paper maps each transformation route to its technical families and to benchmark datasets; tables, text, images, videos, and numbers are the sources, while text and graph are the principal targets. This scheme lets the paper assign every surveyed method a location, which is what carries the survey's argument that the area is coherent rather than a list of isolated tasks.

What would settle it

A reader could test the central claim by checking whether recent transformation methods—especially large-language-model-based table-to-text and image-to-text systems from after 2022—fit inside the paper's categories; if a substantial share of current methods cannot be assigned to any family, the claim of systematic coverage would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that heterogeneity has four root causes—schema conflict, data conflict, format conflict, and domain conflict—and that format conflict, although the most common in AI data preparation, has been under-served by surveys. It contends that data transformation is the appropriate remedy for format conflict, and that transformation strategies can be classified by target format: data-to-text (table-to-text, text-to-text, image-to-text, video-to-text) and data-to-graph (text-to-graph, image-to-graph, video-to-graph). Within each route it distinguishes further technical families, such as template-based versus attention-based sequence-to-sequence table generation, statistical versus learning-based keyword and topic extraction, rule-based versus deep-learning versus hybrid image captioning, and similarity-based versus convolution-based video graphs. The paper also asserts that the main risks in these transformations are information loss and hallucination, and that current graph transformations mostly serve as input to graph neural networks rather than as human-understandable knowledge structures.

Load-bearing premise

The survey's usefulness depends on the selected papers and datasets being representative and comprehensive for each transformation category, yet the paper does not describe a systematic search or selection protocol for its coverage.

Editorial extensions

If this is right

  • If format conflict is treated as a distinct heterogeneity type, data transformation becomes a legitimate step in data preparation pipelines rather than an afterthought.
  • Practitioners can select a baseline technique from the paper's taxonomy according to source and target format, and know which datasets to train or evaluate on.
  • The open challenges the paper names—hallucination in data-to-text, limited scalability of text-to-graph, and narrow use of image/video-to-graph as graph-neural-network input—define near-term research targets.
  • Deep-learning-based transformation is asserted to reduce human effort and information loss compared with manual conversion, implying that transformation models should be evaluated on fidelity rather than fluency alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy predates the large-language-model era, appearing to cover BERT and GPT models but not the instruction-tuned and multimodal generation models that now dominate data preparation, so extending the categories to those methods is an obvious next step the paper leaves implicit.
  • If transformation is truly a first-class remedy for format conflict, evaluation benchmarks should measure information preservation across transformations, not only downstream task performance.
  • The format-conflict category suggests a testable hypothesis: standardizing data at the format layer could reduce the need for schema and data cleaning downstream in AI pipelines.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey addresses data heterogeneity caused by format conflicts. It introduces a conflict-factor taxonomy (schema, data, format, and domain conflicts) and then organizes data transformation methods into data-to-text and data-to-graph categories, covering table-to-text, text-to-text, image-to-text, video-to-text, text-to-graph, image-to-graph, and video-to-graph. Each section summarizes representative methods, datasets, and challenges. The abstract and Section 1.2 additionally claim coverage of table, text, image, video, and numerical data transformed into text, graph, and visual formats, and describe the survey as comprehensive and systematic.

Significance. The useful core of the survey is the conflict-factor framing and the structured presentation of transformation methods with dataset tables (Tables 1-8). As an entry point to classical and early deep-learning data-to-text and data-to-graph methods, the paper has value, and the per-section challenge discussions are a strength. The significance for publication, however, is conditional: the claimed comprehensive scope is not matched by the implemented table of contents, and the absence of a search/selection methodology makes the 'comprehensive' label unverifiable. If the scope claims are reconciled with the actual coverage, the survey could be a useful reference for practitioners.

major comments (3)
  1. [Abstract; §1.2; §2.3; §5] The central claim of the paper is that it 'systematically categorizes strategies for converting prevalent data formats, including table, text, image, video, and numerical data, into formats ... specifically text, graph, and visual formats' (§1.2). The implemented table of contents, however, contains no data-to-visual section and no numerical-source transformation section. Section 5 itself narrows the actual coverage to 'data-to-text and data-to-graph,' omitting the visual target entirely, and numerical data appears only as dataset entries (e.g., numericNLG in Table 1) and as a challenge in §3.1.4. The promised source/target matrix is therefore not delivered. Please either add the missing sections or revise the abstract, §1.2, §2.3, and Section 3's opening paragraph to state the actual scope.
  2. [§1.1, §1.2] The survey calls itself 'comprehensive' and identifies as a gap the absence of 'the latest AI technology' (L3), but it provides no literature-search or selection protocol and no inclusion criteria. As a result, the comprehensiveness claim is not verifiable. The coverage also appears front-loaded: table-to-text (§3.1) centers on attention-based sequence-to-sequence models and does not discuss instruction-tuned or few-shot LLM-based transformation, and the BERT/GPT discussion in §3.2.3 ends at the 2018 models. Please add a methodology subsection describing the search and selection process, or soften the 'comprehensive' and 'latest' claims to match the actual chronological and topical coverage.
  3. [§2.3; §3; §5] Section 2.3 explicitly lists 'numerical format' among the source data formats that data transformation addresses, but no section or subsection is devoted to transforming numerical data into text or graph. The only numerical-adjacent content is the numericNLG dataset row (Table 1) and a brief acknowledgement in §3.1.4 that 'research on understanding numeric data is a critical challenge.' If numerical data is within scope, a dedicated treatment is needed; otherwise, remove it from the format list in §2.3 and the abstract.
minor comments (6)
  1. [§1.1.2] The statement that ChatGPT was introduced by OpenAI in November 2021 is incorrect; ChatGPT was released on November 30, 2022. Please correct the date.
  2. [§1.1.3; §2.2] Unresolved LaTeX citation commands appear in the prose: 'textciteyin2017local' in §1.1.3, and 'citeteplan2002fundamentals, cai2023mbrain' and 'citekim1991classifying' in §2.2. These should be replaced with proper references or removed.
  3. [§2.1] The first two paragraphs of §2.1 duplicate the introductory paragraphs of §2 almost verbatim, so the schema-conflict subsection begins with repeated text; this appears to be a copy-paste error that should be fixed.
  4. [Table 5; Table 8] Several table entries contain typos: Table 8 row 'ster 2019' should presumably be 'Jester 2019'; Table 5 'Googlee Billion Word' should be 'One Billion Word' (or 'Google Billion Word'); Table 8 'AG 2020' should be expanded to 'Action Genome' for readability.
  5. [Table 5] Several rows in Table 5 have missing reference entries (CORD-19, PubMed, NYT based new dataset, HudongBaike, CNCSM), making the dataset list difficult to verify; please add citations or explain their absence.
  6. [Metadata] The ACM Reference Format block and copyright line say '2018' while the manuscript is submitted in 2025; this looks like a template placeholder that should be updated.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the survey's taxonomy and summaries are drawn from external literature, and the single minor self-citation is not load-bearing.

full rationale

This is a literature survey rather than a derivation or empirical prediction, so the standard circularity patterns do not apply. The central claim is organizational: the paper categorizes data transformation strategies into data-to-text (Section 3) and data-to-graph (Section 4), with subsections for table, text, image, and video sources; these categories are labels applied to external methods and datasets, not quantities derived from the paper's own inputs. There is no fitted parameter later called a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz justified solely by a self-citation. The only overlapping-author reference is [66] (Hong et al., 2023), cited in Section 2.2 as one of several data-cleaning overviews and in Section 1.1.1 as part of a list supporting the scarcity of format-conflict transformation surveys; that citation is illustrative and does not force the taxonomy or any substantive conclusion, so it is not load-bearing. The paper does have completeness and consistency limitations: Section 1.2 promises conversion into 'visual formats' and lists 'numerical data' as a source, but Section 5 states the survey is 'categorized them into data-to-text and data-to-graph,' and no visual-target or numerical-source section exists. It also contains unresolved citation remnants such as 'textciteyin2017local,' 'citeteplan2002fundamentals,' and 'citekim1991classifying,' and states ChatGPT was introduced in November 2021. These are correctness and completeness risks, not circularity: they do not show that any claim is equivalent to its own input. Accordingly, no circular step can be exhibited, and the circularity score is low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey rests on the validity of its conflict taxonomy, the premise that data transformation is an effective remedy for format conflicts, and the assumption that the selected papers represent the field. No free parameters or invented entities are introduced.

assumptions (3)
  • domain assumption The taxonomy of heterogeneity into schema, data, format, and domain conflicts is a valid and complete framework.
    Section 2 defines these conflict types and uses them to delimit the survey; if this taxonomy is not exhaustive or well-defined, the survey's scope is unclear.
  • domain assumption Data transformation can remove format conflicts without unacceptable information loss.
    Section 1.1.3 and the discussion invoke this premise, citing Malik et al. [110]; the survey does not validate it empirically.
  • ad hoc to paper The surveyed methods accurately represent the state of the art at the time of writing.
    The survey claims comprehensiveness but includes no explicit inclusion or exclusion criteria and omits many recent advances (e.g., after 2023), making this a stated but unsupported assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data Transformation Strategies to Remove Heterogeneity." pith.science (2026). https://pith.science/paper/4E3OZBAC

@misc{pith2026250712677,
  author       = {Pith},
  title        = {Pith review of: Data Transformation Strategies to Remove Heterogeneity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4E3OZBAC}},
  note         = {Machine review of arXiv:2507.12677}
}
read the original abstract

Data heterogeneity is a prevalent issue, stemming from various conflicting factors, making its utilization complex. This uncertainty, particularly resulting from disparities in data formats, frequently necessitates the involvement of experts to find resolutions. Current methodologies primarily address conflicts related to data structures and schemas, often overlooking the pivotal role played by data transformation. As the utilization of artificial intelligence (AI) continues to expand, there is a growing demand for a more streamlined data preparation process, and data transformation becomes paramount. It customizes training data to enhance AI learning efficiency and adapts input formats to suit diverse AI models. Selecting an appropriate transformation technique is paramount in preserving crucial data details. Despite the widespread integration of AI across various industries, comprehensive reviews concerning contemporary data transformation approaches are scarce. This survey explores the intricacies of data heterogeneity and its underlying sources. It systematically categorizes and presents strategies to address heterogeneity stemming from differences in data formats, shedding light on the inherent challenges associated with each strategy.

Figures

Figures reproduced from arXiv: 2507.12677 by the authors.

Figure 1
Figure 1. Strategies to remove data heterogeneity according to conflict factors • In order to clarify semantic ambiguities, we introduce a classification of heterogeneity rooted in factors causing data conflicts, along with definitions of heterogeneous data and their respective solutions. • Based on their technical similarity, we categorize data transformation strategies for removing heterogeneity caused by format conflicts i… view at source ↗
Figure 2
Figure 2. Strategies of table-to-text transformation. Table-to-text is classified into (a) template-based method and (b) attention-based [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Strategies of text-to-text. (a) is keyword extraction, (b) is topic extraction, and (c) is summarization. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Strategies of image-to-text. (a) is rule-based, (b) is deep learning, and (c) is a hybrid strategy. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Strategies of video-to-text transformation. Video-to-text is classified into (a) mathematical model-based method and (b) deep [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Strategies of text-to-graph transformation. Text-to-graph transformation strategies include (a) Named Entity Recognition (NER) [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]
Figure 7
Figure 7. Figure 7: Strategies of image-to-graph. (a) is pixel distance-based method and (b) is deep learning-based method. [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Strategies of video-to-graph. (a) is similarity-based method and (b) is convolution-based model method. [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

221 extracted references · 67 canonical work pages

  1. [1]

    Kiran Adnan and Rehan Akbar. 2019. Limitations of information extraction methods and techniques for heterogeneous unstructured big data. International Journal of Engineering Business Management 11 (2019), 1847979019890771

  2. [2]

    Gabor Angeli, Julie Tibshirani, Jean Wu, and Christopher D Manning. 2014. Combining distant and partial supervision for relation extraction. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . IEEE, Doha, Qatar, 1556–1567

  3. [3]

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017. Localizing moments in video with natural language. In Proceedings of the IEEE international conference on computer vision . IEEE, Venice, Italy, 5803–5812. Manuscript submitted to ACM 26 Yoo et al

  4. [4]

    Batool Armouty and Sara Tedmori. 2019. Automated keyword extraction using support vector machine from Arabic news documents. In 2019 IEEE Jordan International Joint Conference on Electrical Engineering and Information Technology (JEEIT) . IEEE, Amman, Jordan, 342–346

  5. [5]

    Anurag Arnab, Chen Sun, and Cordelia Schmid. 2021. Unified Graph Structured Models for Video Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Virtual, 8117–8126

  6. [6]

    Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41, 2 (feb 2019), 423–443

  7. [7]

    Junwei Bao, Duyu Tang, Nan Duan, Zhao Yan, Yuanhua Lv, Ming Zhou, and Tiejun Zhao. 2018. Table-to-text: Describing table region with natural language. In Thirty-Second AAAI Conference on Artificial Intelligence . AAAI Press, New Orleans, Lousiana, USA, 5020–5027

  8. [8]

    Carlo Batini and Maurizio Lenzerini. 1984. A Methodology for Data Schema Integration in the Entity Relationship Model. IEEE Transactions on Software Engineering SE-10, 6 (1984), 650–664

Show all 221 references
  1. [9]

    Baumgardner, Larry L

    Marion F. Baumgardner, Larry L. Biehl, and David A. Landgrebe. 2015. 220 Band AVIRIS Hyperspectral Image Data Set: June 12, 1992 Indian Pine Test Site 3. https://purr.purdue.edu/publications/1947/1

  2. [10]

    Christian Bizer, Jens Lehmann, Georgi Kobilarov, Sören Auer, Christian Becker, Richard Cyganiak, and Sebastian Hellmann. 2009. Dbpedia-a crystallization point for the web of data. Journal of web semantics 7, 3 (2009), 154–165

  3. [11]

    Jens Bleiholder and Felix Naumann. 2009. Data fusion. ACM computing surveys (CSUR) 41, 1 (2009), 1–41

  4. [12]

    Yazid Bounab, Mourad Oussalah, and Ahlam Ferdenache. 2020. Reconciling Image Captioning and User’s Comments for Urban Tourism. In 2020 Tenth International Conference on Image Processing Theory, Tools and Applications (IPTA) . IEEE, Paris, France, 1–6

  5. [13]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  6. [14]

    Chen Cai, Kim-Hui Yap, and Suchen Wang. 2022. Attribute Conditioned Fashion Image Captioning. In 2022 IEEE International Conference on Image Processing (ICIP). IEEE, United Arab Emirates, 1921–1925

  7. [15]

    Ricardo Campos, Vítor Mangaravite, Arian Pasquali, Alípio Mário Jorge, Célia Nunes, and Adam Jatowt. 2018. A text feature based automatic keyword extraction method for single documents. In European conference on information retrieval . Springer, Grenoble, France, 684–691

  8. [16]

    Juan Cao, Tian Xia, Jintao Li, Yongdong Zhang, and Sheng Tang. 2009. A density-based method for adaptive LDA model selection. Neurocomputing 72, 7-9 (2009), 1775–1781

  9. [17]

    Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R Hruschka, and Tom M Mitchell. 2010. Toward an architecture for never-ending language learning. In Twenty-Fourth AAAI conference on artificial intelligence . AAAI Press, Atlanta, Georgia, USA, 1306–1313

  10. [18]

    Asli Celikyilmaz, Antoine Bosselut, Xiaodong He, and Yejin Choi. 2018. Deep Communicating Agents for Abstractive Summarization. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol...

  11. [19]

    Kai-Wei Chang, Wen-tau Yih, Bishan Yang, and Christopher Meek. 2014. Typed tensor decomposition of knowledge bases for relation extraction. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . ACL, Doha, Qatar, 1568–1579

  12. [20]

    Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005 (2013), 6 pages

  13. [21]

    David Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies . ACL, Portland, Oregon, USA, 190–200

  14. [22]

    David L Chen and Raymond J Mooney. 2008. Learning to sportscast: a test of grounded language acquisition. InProceedings of the 25th international conference on Machine learning . JMLR.org, Helsinki, Finland, 128–135

  15. [23]

    Liwei Chen, Yansong Feng, Songfang Huang, Yong Qin, and Dongyan Zhao. 2014. Encoding relation requirements for relation extraction via joint inference. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Bal...

  16. [24]

    Penghe Chen, Yu Lu, Vincent W Zheng, Xiyang Chen, and Boda Yang. 2018. Knowedu: A system to construct knowledge graph for education. Ieee Access 6 (2018), 31553–31563

  17. [25]

    Xinlei Chen and C Lawrence Zitnick. 2015. Mind’s eye: A recurrent visual representation for image caption generation. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Boston, MA, USA, 2422–2431

  18. [26]

    Jianpeng Cheng and Mirella Lapata. 2016. Neural Summarization by Extracting Sentences and Words. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Berlin, Germany, 484–494

  19. [27]

    NAVER-NAVER Cloud. 2023. CLOVA. https://clova.ai/

  20. [28]

    Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2019. Show, control and tell: A framework for generating controllable and grounded captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Long Beach, CA, USA, 8307–8316

  21. [29]

    Firefiles.ai Corp. 2023. Firefiles.ai | AI notetaker to transcribe, summarize, search, and analze vioce conversations. https://fireflies.ai/

  22. [30]

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. 2018. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European Conference on ...

  23. [31]

    Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990. Indexing by latent semantic analysis. Journal of the American society for information science 41, 6 (1990), 391–407

  24. [32]

    Chaorui Deng, Ning Ding, Mingkui Tan, and Qi Wu. 2020. Length-controllable image captioning. In Computer Vision–ECCV 2020: 16th European Conference, August 23–28, 2020, Proceedings, Part XIII 16 . Springer, Glasgow, UK, 712–729

  25. [33]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition . IEEE, Miami, Florida, USA, 248–255

  26. [34]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  27. [35]

    Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision ...

  28. [36]

    Xin Luna Dong and Divesh Srivastava. 2013. Big data integration. In 2013 IEEE 29th international conference on data engineering (ICDE) . IEEE, Brisbane, QLD, 1245–1248

  29. [37]

    Xiaoyu Duan, Shi Ying, Hailong Cheng, Wanli Yuan, and Xiang Yin. 2021. OILog: An online incremental log keyword extraction approach based on MDP-LSTM neural network. Information Systems 95 (2021), 101618

  30. [38]

    Greg Durrett, Taylor Berg-Kirkpatrick, and Dan Klein. 2016. Learning-Based Single-Document Summarization with Compression and Anaphoricity Constraints. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ACL, Berl...

  31. [39]

    Elmagarmid, Panagiotis G

    Ahmed K. Elmagarmid, Panagiotis G. Ipeirotis, and Vassilios S. Verykios. 2007. Duplicate Record Detection: A Survey. IEEE Transactions on Knowledge and Data Engineering 19, 1 (2007), 1–16

  32. [40]

    ELSA. 2023. ELSA Speak. https://elsaspeak.com/en/

  33. [41]

    Günes Erkan and Dragomir R Radev. 2004. Lexrank: Graph-based lexical centrality as salience in text summarization.Journal of artificial intelligence research 22 (2004), 457–479

  34. [42]

    Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision 88, 2 (2010), 303–338

  35. [43]

    Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth. 2009. Describing objects by their attributes. In 2009 IEEE conference on computer vision and pattern recognition. IEEE, Miami, Florida, USA, 1778–1785

  36. [44]

    Dhomas Hatta Fudholi, Yurio Windiatmoko, Nurdi Afrianto, Prastyo Eko Susanto, Magfirah Suyuti, Ahmad Fathan Hidayatullah, and Ridho Rahmadi. 2021. Image Captioning with Attention for Smart Local Tourism using EfficientNet. IOP Conference Series: Materials Science and Engineeri...

  37. [45]

    Fan Ge and Li Kuang. 2021. Keywords guided method name generation. In 2021 IEEE/ACM 29th International Conference on Program Comprehension (ICPC). IEEE, Madrid, Spain, 196–206

  38. [46]

    Sebastian Gehrmann, Yuntian Deng, and Alexander Rush. 2018. Bottom-Up Abstractive Summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . ACL, Brussels, Belgium, 4098–4109

  39. [47]

    Shima Gerani, Yashar Mehdad, Giuseppe Carenini, Raymond Ng, and Bita Nejat. 2014. Abstractive summarization of product reviews using discourse structure. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). ACL, Doha, Qatar, 1602–1613

  40. [48]

    Heng Gong, Wei Bi, Xiaocheng Feng, Bing Qin, Xiaojiang Liu, and Ting Liu. 2020. Enhancing content planning for table-to-text generation with data understanding and verification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings...

  41. [49]

    Heng Gong, Xiaocheng Feng, Bing Qin, and Ting Liu. 2019. Table-to-Text Generation with Effective Hierarchical Encoder on Three Dimensions (Row, Column and Time). In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Join...

  42. [50]

    Heng Gong, Yawei Sun, Xiaocheng Feng, Bing Qin, Wei Bi, Xiaojiang Liu, and Ting Liu. 2020. Tablegpt: Few-shot table-to-text generation with table structure reconstruction and content matching. In Proceedings of the 28th International Conference on Computational Linguistics . I...

  43. [51]

    something something

    Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The" something something" video database for learning and evaluating visual common sense....

  44. [52]

    Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. 2018. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE C...

  45. [53]

    Jiuxiang Gu, Shafiq Joty, Jianfei Cai, Handong Zhao, Xu Yang, and Gang Wang. 2019. Unpaired Image Captioning via Scene Graph Alignments. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Seoul, South Korea, 10323–10332. Manuscript submit...

  46. [54]

    Jiatao Gu, Zhengdong Lu, Hang Li, and Victor O.K. Li. 2016. Incorporating Copying Mechanism in Sequence-to-Sequence Learning. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Berlin, Germany, 1631–1640

  47. [55]

    Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko

  48. [56]

    Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou, and Yoshua Bengio. 2016. Pointing the Unknown Words. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Berlin, Germany, 140–149

  49. [57]

    Alon Halevy, Anand Rajaraman, and Joann Ordille. 2006. Data integration: The teenage years. In Proceedings of the 32nd international conference on Very large data bases. ACM, Seoul, South Korea, 9–16

  50. [58]

    Xianpei Han and Le Sun. 2014. Semantic consistency: A local subspace based method for distant supervised relation extraction. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics . ACL, Baltimore, Maryland, 718–724

  51. [59]

    Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision. IEEE, Cambridge, MA, USA, 2961–2969

  52. [60]

    Joseph M Hellerstein. 2008. Quantitative data cleaning for large databases. United Nations Economic Commission for Europe (UNECE) 25 (2008), 1–42

  53. [61]

    James Hendler. 2014. Data integration for heterogenous datasets. Big data 2, 4 (2014), 205–215

  54. [62]

    Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. Advances in neural information processing systems 28 (2015), 9 pages

  55. [63]

    Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013. Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research 47 (2013), 853–899

  56. [64]

    Thomas Hofmann. 1999. Probabilistic latent semantic indexing. In Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval . ACM Press, Berkeley, California, USA, 50–57

  57. [65]

    Danfeng Hong, Lianru Gao, Jing Yao, Bing Zhang, Antonio Plaza, and Jocelyn Chanussot. 2021. Graph Convolutional Networks for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensing 59, 7 (2021), 5966–5978

  58. [66]

    Hyein Hong, Sangbong Yoo, Yejin Jin, and Yun Jang. 2023. How Can We Improve Data Quality for Machine Learning? A Visual Analytics System using Data and Process-driven Strategies. In 2023 IEEE 16th Pacific Visualization Symposium (PacificVis) . IEEE, Seoul, South Korea, 112–121

  59. [67]

    Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga

    MD. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019. A Comprehensive Survey of Deep Learning for Image Captioning. ACM Comput. Surv. 51, 6, Article 118 (feb 2019), 36 pages

  60. [68]

    Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018. A unified model for extractive and abstractive summarization using inconsistency loss. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics . ACL, Melbour...

  61. [69]

    Anette Hulth. 2003. Improved automatic keyword extraction given more linguistic knowledge. In Proceedings of the 2003 conference on Empirical methods in natural language processing . ACL, Sapporo, Japan, 216–223

  62. [70]

    Apple Inc. 2023. Siri - Apple. https://www.apple.com/

  63. [71]

    Masaru Isonuma, Toru Fujino, Junichiro Mori, Yutaka Matsuo, and Ichiro Sakata. 2017. Extractive summarization using multi-task learning with document classification. In Proceedings of the 2017 Conference on empirical methods in natural language processing . ACL, Copenhagen, De...

  64. [72]

    Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J Black. 2013. Towards understanding action recognition. In Proceedings of the IEEE international conference on computer vision . IEEE, Cambridge, MA, USA, 3192–3199

  65. [73]

    Heng Ji and Ralph Grishman. 2011. Knowledge base population: Successful approaches and challenges. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies . ACL, Portland, Oregon, USA, 1148–1158

  66. [74]

    Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles. 2020. Action genome: Actions as compositions of spatio-temporal scene graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Seattle, WA, USA, 10236–10247

  67. [75]

    Qing-Yuan Jiang, Yi He, Gen Li, Jian Lin, Lei Li, and Wu-Jun Li. 2019. SVD: A large-scale short video dataset for near-duplicate video retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision . IEEE, Seoul, South Korea, 5281–5289

  68. [76]

    Xiaotian Jiang, Quan Wang, Peng Li, and Bin Wang. 2016. Relation extraction with multi-instance multi-label convolutional neural networks. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers . The COLING 2016 Organizi...

  69. [77]

    Ian T Jolliffe. 2002. Principal component analysis for special types of data . Springer

  70. [78]

    Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Boston, MA, USA, 3128–3137

  71. [79]

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017), 22 pages

  72. [80]

    Daniel A Keim, Florian Mansmann, Jörn Schneidewind, and Hartmut Ziegler. 2006. Challenges in visual data analysis. In Tenth International Conference on Information Visualisation (IV’06) . IEEE, London, England, 9–16. Manuscript submitted to ACM Data Transformation Strategies t...

  73. [81]

    Vinit Khetani, Yatin Gandhi, Saurabh Bhattacharya, Samir N Ajani, and Suresh Limkar. 2023. Cross-Domain Analysis of ML and DL: Evaluating their Impact in Diverse Domains. International Journal of Intelligent Systems and Applications in Engineering 11, 7s (2023), 253–262

  74. [82]

    Suhyeon Kim, Haecheong Park, and Junghye Lee. 2020. Word2vec-based latent semantic analysis (W2V-LSA) for topic modeling: A study on blockchain technology trend analysis. Expert Systems with Applications 152 (2020), 113401

  75. [83]

    Taejin Kim, Yeoil Yun, and Namgyu Kim. 2021. Deep learning-based knowledge graph generation for COVID-19. Sustainability 13, 4 (2021), 2276

  76. [84]

    Won Kim and Jungyun Seo. 1991. Classifying schematic and data heterogeneity in multidatabase systems. Computer 24, 12 (1991), 12–18

  77. [85]

    Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision . IEEE, Venice, Italy, 706–715

  78. [86]

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of com...

  79. [87]

    Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: a large video database for human motion recognition. In 2011 International conference on computer vision . IEEE, Barcelona, 2556–2563

  80. [88]

    Speakeasy Labs. 2023. Speak - The speaking app that actually talks. https://www.speak.com/

  81. [89]

    Rémi Lebret, David Grangier, and Michael Auli. 2016. Neural Text Generation from Structured Data with Application to the Biography Domain. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . ACL, Austin, Texas, 1203–1213

  82. [90]

    Rémi Lebret, Pedro O Pinheiro, and Ronan Collobert. 2014. Simple image description generator via a linear phrase-based approach. arXiv preprint arXiv:1412.8419 (2014), 7 pages

  83. [91]

    Jie Lei, Liwei Wang, Yelong Shen, Dong Yu, Tamara Berg, and Mohit Bansal. 2020. MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . ACL, Online, 2603–2614

  84. [92]

    Maurizio Lenzerini. 2002. Data Integration: A Theoretical Perspective. ACM, New York, NY, USA, 233–246

  85. [93]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58t...

  86. [94]

    Chengxi Li, Sagar Gandhi, and Brent Harrison. 2019. End-to-end let’s play commentary generation using multi-modal video representations. In Proceedings of the 14th International Conference on the Foundations of Digital Games . ACM, San Luis Obispo, California, USA, 1–7

  87. [95]

    Chenliang Li, Weiran Xu, Si Li, and Sheng Gao. 2018. Guiding generation for abstractive text summarization based on key information guide network. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langua...

  88. [96]

    Siming Li, Girish Kulkarni, Tamara Berg, Alexander Berg, and Yejin Choi. 2011. Composing simple image descriptions using web-scale n-grams. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning . ACL, Portland, Oregon, USA, 220–228

  89. [97]

    Percy Liang, Michael I Jordan, and Dan Klein. 2009. Learning semantic correspondences with less supervision. InProceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP . ACL...

  90. [98]

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision . Springer, Zurich, Switzerland, 740–755

  91. [99]

    Yankai Lin, Shiqi Shen, Zhiyuan Liu, Huanbo Luan, and Maosong Sun. 2016. Neural relation extraction with selective attention over instances. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Berlin, German...

  92. [100]

    Jue Liu, Zhuocheng Lu, and Wei Du. 2019. Combining enterprise knowledge graph and news sentiment analysis for stock price prediction. In Proceedings of the 52nd Hawaii International Conference on System Sciences . University of Hawaii at Manoa, Association for Information Syst...

  93. [101]

    Tianyu Liu, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, and Zhifang Sui. 2019. Hierarchical encoder with auxiliary supervision for neural table-to-text generation: Learning better representation for tables. In Proceedings of the AAAI Conference on Artificial Intelligence ...

  94. [102]

    Tianyu Liu, Kexiang Wang, Lei Sha, Baobao Chang, and Zhifang Sui. 2018. Table-to-text generation by structure-aware seq2seq learning. In Thirty-Second AAAI Conference on Artificial Intelligence . AAAI Press, New Orleans, Louisiana, USA, 4881–4888

  95. [103]

    Tianyu Liu, Xin Zheng, Baobao Chang, and Zhifang Sui. 2021. Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric View. Proceedings of the AAAI Conference on Artificial Intelligence 35, 15 (May 2021), 13415–13423

  96. [104]

    Hans Peter Luhn. 1957. A statistical approach to mechanized encoding and searching of literary information. IBM Journal of research and development 1, 4 (1957), 309–317

  97. [105]

    Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . ACL, Lisbon, Portugal, 1412–1421

  98. [106]

    Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li. 2015. Multimodal convolutional neural networks for matching image and sentence. In Proceedings of the IEEE international conference on computer vision . IEEE, Santiago, Chile, 2623–2631

  99. [107]

    Shuming Ma, Pengcheng Yang, Tianyu Liu, Peng Li, Jie Zhou, and Xu Sun. 2019. Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics . ACL, Florence, Italy, 2047–2...

  100. [108]

    Jayant Madhavan, Philip A Bernstein, and Erhard Rahm. 2001. Generic schema matching with cupid. InIn Proceedings of the International Conference on Very Large Data Bases , Vol. 1. Morgan Kaufmann Publishers Inc., Roma, Italy, 49–58

  101. [109]

    Farzaneh Mahdisoltani, Guillaume Berger, Waseem Gharbieh, David Fleet, and Roland Memisevic. 2018. On the effectiveness of task granularity for transfer learning. arXiv preprint arXiv:1804.09235 (2018), 20 pages

  102. [110]

    Kaleem Razzaq Malik, Tauqir Ahmad, Muhammad Farhan, Muhammad Aslam, Sohail Jabbar, Shehzad Khalid, and Mucheol Kim. 2016. Big-data: transformation from heterogeneous data to semantically-enriched simplified data. Multimedia Tools and Applications 75 (2016), 12727–12747

  103. [111]

    Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L Yuille. 2014. Explain images with multimodal recurrent neural networks. arXiv preprint arXiv:1410.1090 (2014), 9 pages

  104. [112]

    Mohammed Maree and Mohammed Belkhatir. 2015. Addressing semantic heterogeneity through multiple knowledge base assisted merging of domain-specific ontologies. Knowledge-Based Systems 73 (2015), 199–211

  105. [113]

    Martin, C

    D. Martin, C. Fowlkes, D. Tal, and J. Malik. 2001. A Database of Human Segmented Natural Images and its Application to Evaluating Segmentation Algorithms and Measuring Ecological Statistics. In Proc. 8th Int’l Conf. Computer Vision , Vol. 2. IEEE, Vancouver, BC, Canada, 416–423

  106. [114]

    Rebecca Mason and Eugene Charniak. 2014. Nonparametric method for data-driven image captioning. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics . ACL, Baltimore, Maryland, 592–598

  107. [115]

    Joanna Materzynska, Guillaume Berger, Ingo Bax, and Roland Memisevic. 2019. The jester dataset: A large-scale video dataset of human gestures. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops . IEEE, Montreal, BC, Canada, 0–0

  108. [116]

    Yutaka Matsuo and Mitsuru Ishizuka. 2004. Keyword extraction from a single document using word co-occurrence statistical information. International Journal on Artificial Intelligence Tools 13, 01 (2004), 157–169

  109. [117]

    Tong Meng, Xuyang Jing, Zheng Yan, and Witold Pedrycz. 2020. A survey on machine learning for data fusion. Information Fusion 57 (2020), 115–129

  110. [118]

    Remya RK Menon, Deepthy Joseph, and MR Kaimal. 2017. Semantics-based topic inter-relationship extraction. Journal of Intelligent & Fuzzy Systems 32, 4 (2017), 2941–2951

  111. [119]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016), 13 pages

  112. [120]

    Yishu Miao, Edward Grefenstette, and Phil Blunsom. 2017. Discovering discrete latent topics with neural variational inference. In International Conference on Machine Learning . PMLR, Sydney, Australia, 2410–2419

  113. [121]

    Microsoft. 2023. Cortana. https://www.microsoft.com/en-us/cortana

  114. [122]

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision ....

  115. [123]

    Rada Mihalcea and Paul Tarau. 2004. Textrank: Bringing order into text. In Proceedings of the 2004 conference on empirical methods in natural language processing. ACL, Barcelona, Spain, 404–411

  116. [124]

    Margaret Mitchell, Jesse Dodge, Amit Goyal, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa Mensch, Alexander Berg, Tamara Berg, and Hal Daumé III. 2012. Midge: Generating image descriptions from computer vision detections. In Proceedings of the 13th Conference of the Europea...

  117. [125]

    Nicholas Moratelli, Manuele Barraco, Davide Morelli, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2023. Fashion-Oriented Image Captioning with External Knowledge Retrieval and Fully Attentive Gates. Sensors 23, 3 (2023), 16 pages

  118. [126]

    Lili Mou, Yiping Song, Rui Yan, Ge Li, Lu Zhang, and Zhi Jin. 2016. Sequence to Backward and Forward Sequences: A Content-Introducing Approach to Generative Short-Text Conversation. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: ...

  119. [127]

    Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017. Summarunner: A recurrent neural network based sequence model for extractive summarization of documents. In Thirty-first AAAI conference on artificial intelligence . AAAI Press, San Francisco, California USA, 3075–3081

  120. [128]

    Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang. 2016. Abstractive Text Summarization using Sequence-to- sequence RNNs and Beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning . ACL, Berlin, Germany, 280–290

  121. [129]

    Shashi Narayan, Joshua Maynez, Jakub Adamek, Daniele Pighin, Blaz Bratanic, and Ryan McDonald. 2020. Stepwise Extractive Summarization and Planning with Structured Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) ....

  122. [130]

    Shashi Narayan, Nikos Papasarantopoulos, Shay B Cohen, and Mirella Lapata. 2017. Neural extractive summarization with side information. arXiv preprint arXiv:1704.04530 (2017), 9 pages

  123. [131]

    Felix Naumann, Alexander Bilke, Jens Bleiholder, and Melanie Weis. 2006. Data Fusion in Three Steps: Resolving Schema, Tuple, and Value Inconsistencies. IEEE Data Eng. Bull. 29, 2 (2006), 21–31

  124. [132]

    Ani Nenkova and Lucy Vanderwende. 2005. The impact of frequency on summarization. Microsoft Research, Redmond, Washington, Tech. Rep. MSR-TR-2005 101 (2005), 9 pages

  125. [133]

    National Institute of Standards and Techonlogy (NIST). 2002. Document Understanding Conference - publications data. https://duc.nist.gov/data. html Manuscript submitted to ACM Data Transformation Strategies to Remove Heterogeneity 31

  126. [134]

    National Institute of Standards and Techonlogy (NIST). 2021. Text Analysis Conference - Past TAC Data. https://tac.nist.gov/data/index.html

  127. [135]

    Naoya Okumura and Takao Miura. 2015. Automatic labelling of documents based on ontology. In2015 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing (PACRIM). IEEE, Victoria, BC, Canada, 34–39

  128. [136]

    Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011. Im2text: Describing images using 1 million captioned photographs. Advances in neural information processing systems 24 (2011), 1143–1151

  129. [137]

    Tatsuro Oya, Yashar Mehdad, Giuseppe Carenini, and Raymond Ng. 2014. A template-based abstractive meeting summarization: Leveraging summary and source text relationships. In Proceedings of the 8th International Natural Language Generation Conference (INLG) . ACL, Philadelphia,...

  130. [138]

    Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles. 2020. Spatio-Temporal Graph for Video Captioning With Knowledge Distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IE...

  131. [139]

    Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. ToTTo: A Controlled Table-To-Text Generation Dataset. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . ACL, Onlin...

  132. [140]

    Maria Pershina, Bonan Min, Wei Xu, and Ralph Grishman. 2014. Infusion of labeled data into distant supervision for relation extraction. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics . ACL, Baltimore, Maryland, 732–738

  133. [141]

    Venkatesh Babu

    Nikita Prabhu and R. Venkatesh Babu. 2015. Attribute-Graph: A Graph Based Approach to Image Ranking. In 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, Santiago, Chile, 1071–1079

  134. [142]

    Ratish Puduppully, Li Dong, and Mirella Lapata. 2019. Data-to-text Generation with Entity Modeling. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . ACL, Florence, Italy, 2023–2035

  135. [143]

    Longhua Qian, Guodong Zhou, Fang Kong, Qiaoming Zhu, and Peide Qian. 2008. Exploiting constituent dependencies for tree kernel-based semantic relation extraction. In Proceedings of the 22nd International Conference on Computational Linguistics (Coling 2008) . Coling 2008 Organ...

  136. [144]

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. https://openai.com/research/language-unsupervised

  137. [145]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21, 1, Article 140 (jan 2020), 67 pages

  138. [146]

    Erhard Rahm, Hong Hai Do, et al. 2000. Data cleaning: Problems and current approaches. IEEE Data Eng. Bull. 23, 4 (2000), 3–13

  139. [147]

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 (2022), 27 pages

  140. [148]

    Juan Ramos et al. 2003. Using tf-idf to determine word relevance in document queries. In Proceedings of the first instructional conference on machine learning, Vol. 242. Citeseer, irtual Conference, 29–48

  141. [149]

    Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting Image Annotations Using Amazon’s Mechanical Turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk . ACL, Los Angeles, 139–147

  142. [150]

    Clement Rebuffel, Marco Roberti, Laure Soulier, Geoffrey Scoutheeten, Rossella Cancelliere, and Patrick Gallinari. 2022. Controlling Hallucinations at Word Level in Data-to-Text Generation. Data Min. Knowl. Discov. 36, 1 (jan 2022), 318–354

  143. [151]

    Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics 1 (2013), 25–36

  144. [152]

    Sebastian Riedel, Limin Yao, and Andrew McCallum. 2010. Modeling relations and their mentions without labeled text. InJoint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, Athens, Greece, 148–163

  145. [153]

    Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014. Coherent multi-sentence video description with variable level of detail. In German conference on pattern recognition . Springer, Munster, Germany, 184–195

  146. [154]

    Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. 2015. A dataset for movie description. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Boston, MA, USA, 3202–3212

  147. [155]

    Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele. 2013. Translating video content to natural language descriptions. In Proceedings of the IEEE international conference on computer vision . IEEE, Sydney, NSW, Australia, 433–440

  148. [156]

    Maya Rotmensch, Yoni Halpern, Abdulhakim Tlimat, Steven Horng, and David Sontag. 2017. Learning a health knowledge graph from electronic medical records. Scientific reports 7, 1 (2017), 1–11

  149. [157]

    Rush, Sumit Chopra, and Jason Weston

    Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A Neural Attention Model for Abstractive Sentence Summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . ACL, Lisbon, Portugal, 379–389

  150. [158]

    Yasubumi Sakakibara, Kazuo Misue, and Takeshi Koshiba. 1993. Text classification and keyword extraction by learning decision trees. InProceedings of 9th IEEE Conference on Artificial Intelligence for Applications . IEEE, Orlando, FL, USA, 466

  151. [159]

    Liu, and Christopher D

    Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . ACL, Vancouver, Canada, 1073–1083

  152. [160]

    Lei Sha, Lili Mou, Tianyu Liu, Pascal Poupart, Sujian Li, Baobao Chang, and Zhifang Sui. 2018. Order-planning neural text generation from structured data. In Thirty-Second AAAI Conference on Artificial Intelligence . AAAI Press, New Orleans, Louisiana, USA, 5414–5421. Manuscri...

  153. [161]

    David I Shuman, Sunil K Narang, Pascal Frossard, Antonio Ortega, and Pierre Vandergheynst. 2013. The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains. IEEE signal processing magazine 30, 3 (2013), 83–98

  154. [162]

    Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision . Springer, Amsterdam, Netherlands, 510–526

  155. [163]

    Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem. 2021. VLG-Net: Video-Language Graph Matching Network for Video Grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops . IEEE, Montreal, BC, Canada, 3224–3234

  156. [164]

    Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012), 7 pages

  157. [165]

    The standalone DAVIS. 2017. The 2017 davis challenge on video object segmentation. https://davischallenge.org/davis2017/code.html

  158. [166]

    Yixuan Su, Zaiqiao Meng, Simon Baker, and Nigel Collier. 2021. Few-Shot Table-to-Text Generation with Prototype Memory. InFindings of the Association for Computational Linguistics: EMNLP 2021 . ACL, Punta Cana, Dominican Republic, 910–917

  159. [167]

    Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura. 2021. Towards Table-to-Text Generation with Numerical Reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internati...

  160. [168]

    Le Sun and Xianpei Han. 2014. A feature-enriched tree kernel for relation extraction. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics. ACL, Baltimore, Maryland, 61–67

  161. [169]

    Tsunehiko Tanaka and Edgar Simo-Serra. 2021. LoL-V2T: Large-Scale Esports Video Description Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, Nashville, TN, USA, 4557–4566

  162. [170]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca: A Strong, Replicable Instruction-Following Model. https://crfm.stanford.edu/2023/03/13/alpaca.html

  163. [171]

    Common Crawl team founded by Gil Elbaz. 2007. Common Crawl. https://commoncrawl.org/

  164. [172]

    Radu Timofte, Eirikur Agustsson, Luc Van Gool, Ming-Hsuan Yang, and Lei Zhang. 2017. Ntire 2017 challenge on single image super-resolution: Methods and results. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops . IEEE, Honolulu, HI, USA...

  165. [173]

    Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015. Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070 (2015), 7 pages

  166. [174]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023), 27 pages

  167. [175]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023), 77 pages

  168. [176]

    Yuen-Hsien Tseng, Chi-Jen Lin, Hsiu-Han Chen, and Yu-I Lin. 2006. Toward generic title generation for clustered documents. In Asia Information Retrieval Symposium. Springer, Singapore, 145–157

  169. [177]

    Diego Valsesia, Giulia Fracastoro, and Enrico Magli. 2020. Deep Graph-Convolutional Image Denoising. IEEE Transactions on Image Processing 29 (2020), 8226–8237

  170. [178]

    Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, 11 (2008), 2579–2605

  171. [179]

    Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2015. Sequence to sequence-video to text. In Proceedings of the IEEE international conference on computer vision . IEEE, Santiago, Chile, 4534–4542

  172. [180]

    Yashaswi Verma and CV Jawahar. 2014. Im2Text and Text2Im: Associating Images and Texts for Cross-Modal Retrieval. InBMVC, Vol. 1. Citeseer, Nottingham, 2

  173. [181]

    Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Boston, MA, USA, 3156–3164

  174. [182]

    Sheng Wan, Chen Gong, Ping Zhong, Shirui Pan, Guangyu Li, and Jian Yang. 2021. Hyperspectral Image Classification With Context-Aware Dynamic Graph Convolutional Network. IEEE Transactions on Geoscience and Remote Sensing 59, 1 (2021), 597–612

  175. [183]

    Lidong Wang. 2017. Heterogeneous data and big data analytics. Automatic Control and Information Sciences 3, 1 (2017), 8–15

  176. [184]

    Peng Wang, Junyang Lin, An Yang, Chang Zhou, Yichang Zhang, Jingren Zhou, and Hongxia Yang. 2021. Sketch and Refine: Towards Faithful and Informative Table-to-Text Generation. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021 . ACL, Online, 4831–4843

  177. [185]

    Qingyun Wang, Xiaoman Pan, Lifu Huang, Boliang Zhang, Zhiying Jiang, Heng Ji, and Kevin Knight. 2018. Describing a Knowledge Base. In Proceedings of the 11th International Conference on Natural Language Generation . ACL, Tilburg University, The Netherlands, 10–21

  178. [186]

    Sarma, Michael M

    Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E. Sarma, Michael M. Bronstein, and Justin M. Solomon. 2019. Dynamic Graph CNN for Learning on Point Clouds. ACM Trans. Graph. 38, 5, Article 146 (oct 2019), 12 pages

  179. [187]

    Y Wang and J Zhang. 2017. Keyword extraction from online product reviews based on bi-directional LSTM recurrent neural network. In 2017 IEEE International Conference on Industrial Engineering and Engineering Management (IEEM) . IEEE, Singapore, 2241–2245

  180. [188]

    Zhenyi Wang, Xiaoyang Wang, Bang An, Dong Yu, and Changyou Chen. 2020. Towards Faithful Neural Table-to-Text Generation with Content- Matching Constraints. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . ACL, Online, 1072–1086

  181. [189]

    Christian Wartena, Rogier Brussee, and Wout Slakhorst. 2010. Keyword extraction using word co-occurrence. In 2010 Workshops on Database and Expert Systems Applications. IEEE, Bilbao, Spain, 54–58. Manuscript submitted to ACM Data Transformation Strategies to Remove Heterogeneity 33

  182. [190]

    Xiangpeng Wei, Yue Hu, Luxi Xing, Yipeng Wang, and Li Gao. 2019. Translating with bilingual topic knowledge for neural machine translation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. AAAI Press, Honolulu, Hawaii, USA, 7257–7264

  183. [191]

    Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In Proceedings of the Twelfth Language Resources and Evaluation Conferenc...

  184. [192]

    Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017. Challenges in Data-to-Document Generation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . ACL, Copenhagen, Denmark, 2253–2263

  185. [193]

    Yi-fang Brook Wu, Quanzhi Li, Razvan Stefan Bot, and Xin Chen. 2005. Domain-specific keyphrase extraction. In Proceedings of the 14th ACM international conference on Information and knowledge management . ACM, Bremen, Germany, 283–284

  186. [194]

    Chen Xing, Wei Wu, Yu Wu, Jie Liu, Yalou Huang, Ming Zhou, and Wei-Ying Ma. 2017. Topic aware neural response generation. InProceedings of the AAAI Conference on Artificial Intelligence , Vol. 31. AAAI Press, San Francisco, California, USA, 3351–3357

  187. [195]

    Jian Xu, Sunkyu Kim, Min Song, Minbyul Jeong, Donghyeon Kim, Jaewoo Kang, Justin F Rousseau, Xin Li, Weijia Xu, Vetle I Torvik, et al. 2020. Building a PubMed knowledge graph. Scientific data 7, 1 (2020), 1–15

  188. [196]

    Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Las Vegas, NV, USA, 5288–5296

  189. [197]

    Semih Yagcioglu, Erkut Erdem, Aykut Erdem, and Ruket Cakıcı. 2015. A distributed representation based query expansion approach for image captioning. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Confe...

  190. [198]

    Fei Yan and Krystian Mikolajczyk. 2015. Deep correlation for matching images and text. In Proceedings of the IEEE conference on computer vision and pattern recognition. IEEE, Boston, MA, USA, 3441–3450

  191. [199]

    Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning. In Proceedings of the IEEE/CVF Conference on Computer V...

  192. [200]

    Yezhou Yang, Ching Teo, Hal Daumé III, and Yiannis Aloimonos. 2011. Corpus-guided sentence generation of natural images. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing . ACL, Edinburgh, Scotland, UK, 444–454

  193. [201]

    Wen-tau Yih, Joshua Goodman, and Vitor R Carvalho. 2006. Finding advertising keywords on web pages. In Proceedings of the 15th international conference on World Wide Web. ACM, Edinburgh, Scotland, 213–222

  194. [202]

    Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016. Image captioning with semantic attention. In Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, Las Vegas, NV, USA, 4651–4659

  195. [203]

    Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics 2 (2014), 67–78

  196. [204]

    Youper. 2023. Youper: Artificial Intelligence For Mental Health Care. https://www.youper.ai/

  197. [205]

    Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2022. A Survey of Knowledge-Enhanced Text Generation. ACM Comput. Surv. 54, 11s, Article 227 (nov 2022), 38 pages

  198. [206]

    Dmitry Zelenko, Chinatsu Aone, and Anthony Richardella. 2003. Kernel Methods for Relation Extraction. J. Mach. Learn. Res. 3, null (mar 2003), 1083–1106

  199. [207]

    Wenyuan Zeng, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2017. Incorporating Relation Paths in Neural Relation Extraction. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . ACL, Copenhagen, Denmark, 1768–1777

  200. [208]

    Junchao Zhang and Yuxin Peng. 2019. Object-Aware Aggregation With Bidirectional Temporal Graph for Video Captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Long Beach, CA, USA, 8319–8328

  201. [209]

    Jingran Zhang, Fumin Shen, Xing Xu, and Heng Tao Shen. 2020. Temporal Reasoning Graph for Activity Recognition. IEEE Transactions on Image Processing 29 (2020), 5491–5506

  202. [210]

    Jingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao, Jie Shao, and Xiaofeng Zhu. 2021. Video Representation Learning with Graph Contrastive Augmentation. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). ACM, New York, NY, USA, 3043–3051

  203. [211]

    Wei Zhang, Yue Ying, Pan Lu, and Hongyuan Zha. 2020. Learning Long-and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. AAAI Press, Ne...

  204. [212]

    Guoping Zhao, Mingyu Zhang, Yaxian Li, Jiajun Liu, Bingqing Zhang, and Ji-Rong Wen. 2021. Pyramid regional graph representation learning for content-based video retrieval. Information Processing & Management 58, 3 (2021), 102488

  205. [213]

    Zixu Zhao, Yueming Jin, and Pheng-Ann Heng. 2021. Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Montreal, QC, Canada, 9960–9969

  206. [214]

    GuoDong Zhou, Jian Su, Jie Zhang, and Min Zhang. 2005. Exploring Various Knowledge in Relation Extraction. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05) . ACL, Ann Arbor, Michigan, 427–434

  207. [215]

    Luowei Zhou, Chenliang Xu, and Jason J Corso. 2018. Towards automatic learning of procedures from web instructional videos. In Thirty-Second AAAI Conference on Artificial Intelligence . AAAI Press, New Orleans, LA, USA, 7590–7598. Manuscript submitted to ACM 34 Yoo et al

  208. [216]

    Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong. 2018. End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, Salt Lake City, UT, USA, 8739–8748

  209. [217]

    Qixian Zhou, Xiaodan Liang, Ke Gong, and Liang Lin. 2018. Adaptive temporal encoding network for video instance-level human parsing. In Proceedings of the ACM international conference on Multimedia . ACM, Seoul, South Korea, 1527–1535

  210. [218]

    Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. 2018. Neural Document Summarization by Jointly Learning to Score and Select Sentences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)...

  211. [219]

    Shangchen Zhou, Jiawei Zhang, Wangmeng Zuo, and Chen Change Loy. 2020. Cross-Scale Internal Graph Neural Network for Image Super- Resolution. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. ...

  212. [220]

    Thapliyal, William Yang Wang, and Radu Soricut

    Wanrong Zhu, Bo Pang, Ashish V. Thapliyal, William Yang Wang, and Radu Soricut. 2022. End-to-end Dense Video Captioning as Sequence Generation. In Proceedings of the 29th International Conference on Computational Linguistics . International Committee on Computational Linguisti...

  213. [2013]

    In Proceedings of the IEEE international conference on computer vision

    Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition. In Proceedings of the IEEE international conference on computer vision . IEEE, Cambridge, MA, USA, 2712–2719

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.