Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

Bridging the Data Provenance Gap Across Text, Speech and Video

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A multimodal audit of nearly 4,000 datasets argues that dataset licenses drastically understate the legal constraints on AI training data, because over 80% of the content in widely used text, speech, and video datasets carries…

desk verdict First modality-spanning provenance audit; useful reference, but 'over 80%' overstates speech (78.6%) and the manual source coding needs validation before the headline numbers are taken as fact. read the letter →

arxiv 2412.17847 v2 pith:IFMWDFZY submitted 2024-12-19 cs.AI cs.CLcs.CYcs.LGcs.MM

classification cs.AIcs.CLcs.CYcs.LGcs.MM
keywords dataprovenancedatasetlicensingtrainingauditmultimodaldatasetssourcerestrictionsnon-commercialgeographicalrepresentationlinguistic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the public record of how AI training data may be used is systematically misleading. Auditing nearly 4,000 text, speech, and video datasets released between 1990 and 2024, the authors find that while fewer than a third of datasets carry restrictive licenses, the underlying source content is far more constrained: 99.8% of text tokens, 78% of speech hours, and 99% of video hours carry non-commercial restrictions that come from the sources — websites, social media platforms, or generative models — the data was drawn from. Because those source-level restrictions are usually dropped when datasets are re-packaged, a permissive dataset license does not tell practitioners what they may actually use. A sympathetic reader would care because the finding puts a measurable legal shadow over the common practice of assembling large training corpora from web crawls, YouTube, and synthetic model outputs, and the released audit lets individual developers trace a dataset's chain of provenance for themselves.

What carries the argument

The load-bearing instrument is a manual provenance-tracing protocol that codes each dataset on two axes at once. The dataset license is categorized as Commercial, Non-commercial/Academic, or Unspecified, while each underlying source — every website, platform, or model the content came from — is coded as Unrestricted, Unspecified, Source Closed, or Model Closed, following a defined taxonomy that covers terms of service, acceptable-use policies, and anti-crawling clauses. A dataset's overall terms status is set to the strictest of its sources, and the crosstab of license against terms is then reported both by dataset count and weighted by tokens or hours, which is what produces the headline mismatch figures.

What would settle it

Take a random sample of a few hundred datasets from the released audit, have a second independent team re-annotate source restrictions from the same lineage, and compute agreement; low agreement on the Source Closed and Model Closed categories would undercut the 80%-plus mismatch claim. A separate check: recompute the 99.8% text, 78% speech, and 99% video figures with the single largest collection per modality removed (for example the roughly 370k-hour YouTube-sourced speech corpus), since one re-coded giant could shift the ecosystem totals.

Watch

Extended reading notes

Core claim

The paper's central claim is a mismatch between two layers of restriction: dataset licenses and source terms. Counting datasets, only 25% of text, 33% of speech, and 32% of video datasets are licensed non-commercially; weighted by content, those figures are 21%, 26%, and 33%. But when the authors trace each dataset back to its original sources and classify the sources' licenses or terms of service, 99.8% of text tokens, 78% of speech hours, and 99% of video hours carry some non-commercial restriction at the source level, and the license-versus-terms mismatch covers 79% of text tokens, 55% of speech hours, and 65% of video hours. The paper further claims that since 2019 multimodal training has overwhelmingly moved to web-crawled, synthetic, and social-media sources (YouTube alone supplies roughly 71% of video data and 69% of speech data), and that although the absolute number of languages and countries represented keeps rising, the relative dispersion measured by Gini coefficients has not significantly improved since 2013 — Western concentration persists despite diversification at the margins.

Load-bearing premise

The central percentages depend on the manual classification of each dataset's sources as Unrestricted, Unspecified, Source Closed, or Model Closed, and the paper reports no inter-annotator agreement or independent validation of that coding, so a systematic bias in it would move every headline number.

Editorial extensions

If this is right

  • Developers who filter training data by dataset license alone will systematically misjudge the restrictions on the content, since source-level terms bind more than 80% of content in each modality.
  • Because the strictest-source rule intensifies at the collection level, large re-packaged text collections hide commercially usable subsets that practitioners cannot easily extract.
  • The concentration of speech and video on a single video platform means platform terms, not dataset licenses, are the effective gatekeepers of most audio-visual training data.
  • If the source coding is correct, the permissive commons for multimodal training reduces to the small residual of content that is both commercially licensed and sourced from unrestricted sources — under 1% of text and video content and about 5% of speech hours.
  • Relative geographical and linguistic representation has been flat for a decade, so adding more languages and countries at the margins does not by itself reduce concentration.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One implication the paper leaves implicit: if source terms bind, then a 'clean' multimodal training run under current terms would have to be assembled almost entirely from a short list of explicitly permissive sources, and the paper's released tables make that set enumerable.
  • A testable extension: apply the same two-axis coding to datasets released after April 2024 to see whether the license-versus-source mismatch grows as synthetic outputs and short-video platforms enter the training mix.
  • The mismatch framing points to a fork the paper deliberately does not resolve: either source terms are largely unenforceable against training (shifting risk to copyright law itself), or a large fraction of existing multimodal training is already in breach of terms — future litigation will pick the branch.
  • Because the volume-weighted percentages are driven by a handful of giant collections, a sensitivity analysis that recomputes the headline figures with the largest collection per modality removed would show how robust the ecosystem-level claims actually are.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a large-scale manual audit of 3,916 public text, speech, and video datasets released between 1990 and 2024, tracing sourcing trends, license and source-term restrictions, and geographic and linguistic representation. The authors report three headline findings: (1) multimodal training data increasingly comes from web-crawled, social media, and synthetic sources; (2) although fewer than one-third of datasets carry restrictive licenses, over 80% of the underlying source content carries non-commercial restrictions; and (3) absolute language and country counts have risen since 2013 but relative measures of geographic and multilingual inequality have not significantly improved. The paper includes extensive appendix tables, a detailed taxonomy, and a promised public release of the audit data and code. The central quantitative claims are descriptive measurements rather than fitted-model outputs, but the main restriction percentages depend entirely on manually assigned source-level labels whose reliability is not reported.

Significance. If the headline measurement is correct, the paper provides an important and policy-relevant result: the effective legal constraints on widely used training data are much stronger than dataset licenses alone suggest. The work is the first multimodal provenance audit at this scale, and the appendix tables and attribution cards are a substantial community resource. The paper is appropriately cautious in describing licenses and terms as signals rather than enforceable legal determinations, and it avoids overfitting by presenting descriptive statistics rather than model-based claims. The empirical contribution would be strengthened by the promised release of annotations and code, which is an explicit strength if delivered. However, the central 'over 80%' claim is not yet fully supported: it is internally inconsistent with the paper's own Table 3 for speech, and the underlying manual classifications have no reported reliability validation.

major comments (3)
  1. [Section 2, 'Annotation Features & Methodology'; Tables 3 and 4] The central quantitative claims—99.8% of text tokens, 78.6% of speech hours, and 99.1% of video hours carrying source restrictions—are direct tabulations of the manually assigned labels in Table 2, yet the paper reports no inter-annotator agreement, no gold-standard validation, and no sensitivity analysis for ambiguous cases. A systematic bias in coding (for example, treating 'no information found' as Restricted rather than Unspecified) would directly change the headline percentages and the abstract's 'over 80%' conclusion. The manuscript states that 'All annotations and analysis code will be made publicly available on release,' but no link or commit hash is given, so independent verification is not currently possible. Please add a reliability study, such as dual annotation with agreement statistics on a representative subsample or an audit of ambiguous cases, and provide the data and code artifact at submission time.
  2. [Abstract and Section 3.2 (Table 3)] The abstract claims that 'over 80% of the source content in widely-used text, speech, and video datasets carry non-commercial restrictions,' and the introduction states 'over 80% of content from each modality.' Table 3 reports 78.6% for speech hours, and Section 3.2 itself says '78%' for speech. The claim is therefore internally inconsistent with the paper's own appendix data. The wording should be corrected to 'roughly 80%,' 'over 78%,' or the computation should be adjusted so that the abstract matches the reported tables.
  3. [Section 3.3, Figure 4] The claim that geographical and linguistic representation 'has not significantly improved' rests on Gini coefficients with 95% confidence intervals and statements about significance at the p = 0.05 level, but the method for computing these intervals and tests is not described anywhere in the paper or appendix, and no code is provided. Without knowing whether the intervals come from a bootstrap, jackknife, or analytic approach, the 'not significant' conclusion is not verifiable. Please specify the estimator and test procedure, or downgrade the claim to a descriptive statement about observed trends.
minor comments (6)
  1. [Section 2.1 vs Table 1] Section 2.1 reports 3,713 text datasets from 108 collections, while Table 1 reports 3,717 text datasets; these counts should be reconciled.
  2. [References] References [11] and [12] are the same paper by Buolamwini and Gebru, and the Common Voice citation appears twice, once as [13] and again as [20]; these duplicates should be merged.
  3. [Table 3 caption] The caption says the table is a breakdown 'across datasets' but the cells are percentages of total tokens or hours; clarify that the units are shares of content, not dataset counts, to avoid confusion with Table 4.
  4. [Figure 2 caption and Section 3.2] The phrase 'bare restrictions' appears in the Figure 2 caption and in the body text; this should be 'bear restrictions'.
  5. [Introduction, finding 2] The phrase 'undocumented restrictions in the dataset's sources' is imprecise: the source restrictions are often documented in terms of service, but not in the dataset license; consider wording such as 'source-level restrictions not reflected in dataset licenses.'
  6. [Section 3.1] The comparison of average synthetic and natural dataset lengths (1,756 vs 1,065 tokens) would benefit from sample sizes and a measure of dispersion to support the word 'notably.'

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the reported percentages are new manual measurements, and the inherited taxonomy from the authors' prior audit is not load-bearing.

full rationale

The paper's central claims are measurements from manual annotations, not derivations from fitted parameters or self-referential definitions. The headline restriction percentages (e.g., 99.8% text, 78.6% speech, 99.1% video in Table 3) are computed by classifying each dataset source into Unrestricted, Unspecified, Source Closed, or Model Closed per Table 2, then aggregating by token or hour counts. Nothing in Table 2 is defined in terms of the headline result, and no parameter is fitted to a subset of the data and then 'predicted' on a closely related quantity; the annotation labels are the input and the source-restriction percentages are the output, which is the normal structure of a descriptive audit rather than a circular derivation. The adoption of the prior taxonomy and annotation process from the authors' own work, Longpre et al. [123], is a methodological inheritance and is explicitly acknowledged; it does not function as an unverified uniqueness theorem that forces the conclusion, and the current paper extends the taxonomy to source terms for speech and video. Other self-citations such as [122]-[125] provide context and framing, but the load-bearing evidence is the new dataset-level annotation, which is independently checkable once released. The most serious concerns are correctness and robustness issues rather than circularity: no inter-annotator agreement or gold-standard validation is reported for the manual source coding, the abstract's 'over 80%' claim sits uneasily with the 78.6% speech figure in Table 3, and the annotation code is promised rather than linked, as the paper states that 'all annotations and analysis code will be made publicly available on release.' These limitations affect reliability and verifiability, but they do not make the derivation circular.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is an empirical audit rather than a derivation, so there are no fitted parameters or invented entities. The load-bearing inputs are the curated dataset sample and the manual annotation decisions, listed as domain assumptions.

assumptions (3)
  • domain assumption The curated dataset list (from HuggingFace, surveys, and expert review) is representative of widely used public datasets
    Section 2 "Scope & Dataset Selection" defines inclusion by popularity and expert supplementation; analyses generalize to the ecosystem.
  • domain assumption Manual source-term annotations follow the Table 2 taxonomy consistently without measurable disagreement
    Section 2 "Annotation Features & Methodology" describes expert annotation but reports no inter-annotator agreement.
  • ad hoc to paper The Gini coefficient confidence intervals and p-value statements in Figure 4 are computed by an unspecified but valid method
    Section 3.3 reports 95% CIs and p=0.05 significance tests but does not describe the bootstrap or test procedure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging the Data Provenance Gap Across Text, Speech and Video." pith.science (2026). https://pith.science/paper/IFMWDFZY

@misc{pith2026241217847,
  author       = {Pith},
  title        = {Pith review of: Bridging the Data Provenance Gap Across Text, Speech and Video},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFMWDFZY}},
  note         = {Machine review of arXiv:2412.17847}
}
read the original abstract

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities--popular text, speech, and video datasets--from their detailed sourcing trends and use restrictions to their geographical and linguistic representation. Our manual analysis covers nearly 4000 public datasets between 1990-2024, spanning 608 languages, 798 sources, 659 organizations, and 67 countries. We find that multimodal machine learning applications have overwhelmingly turned to web-crawled, synthetic, and social media platforms, such as YouTube, for their training sets, eclipsing all other sources since 2019. Secondly, tracing the chain of dataset derivations we find that while less than 33% of datasets are restrictively licensed, over 80% of the source content in widely-used text, speech, and video datasets, carry non-commercial restrictions. Finally, counter to the rising number of languages and geographies represented in public AI training datasets, our audit demonstrates measures of relative geographical and multilingual representation have failed to significantly improve their coverage since 2013. We believe the breadth of our audit enables us to empirically examine trends in data sourcing, restrictions, and Western-centricity at an ecosystem-level, and that visibility into these questions are essential to progress in responsible AI. As a contribution to ongoing improvements in dataset transparency and responsible use, we release our entire multimodal audit, allowing practitioners to trace data provenance across text, speech, and video.

Figures

Figures reproduced from arXiv: 2412.17847 by the authors.

Figure 1
Figure 1. The cumulative size of data (log-scale tokens for text, hours for speech/video) from each [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The distribution of restrictions from dataset [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The geographical distribution of countries (world maps) and continents (table) represented [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The cumulative totals (left) of languages and countries represented in the data over time, and [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: The distribution of creator organizations by modality. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: The distribution of dataset sizes for each modality. Most text data collections are between [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]
Figure 7
Figure 7. Figure 7: The task distribution of datasets, across modalities. Post-training text and video datasets [PITH_FULL_IMAGE:figures/full_fig_p022_7.png]
Figure 6
Figure 6. Figure 6: However, for Figure 2 we draw the distinction between collection and dataset metrics, as [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.

  2. TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.

  3. BRoverbs -- Measuring how much LLMs understand Portuguese proverbs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.

Reference graph

Works this paper leans on

294 extracted references · 22 canonical work pages · cited by 3 Pith papers

  1. [1]

    Untitled review,

    E. B. Wilson, “Untitled review,” The American Economic Review, vol. 4, no. 2, pp. 442– 444, 1914, ISSN : 00028282. [Online]. Available: http://www.jstor.org/stable/ 1804762 (visited on 09/26/2024)

  2. [2]

    On the measurement of inequality,

    A. B. Atkinson et al., “On the measurement of inequality,”Journal of economic theory, vol. 2, no. 3, pp. 244–263, 1970

  3. [3]

    A survey of video datasets for human action and activity recognition,

    J. M. Chaquet, E. J. Carmona, and A. Fernández-Caballero, “A survey of video datasets for human action and activity recognition,” Computer Vision and Image Understanding, vol. 117, no. 6, pp. 633–659, 2013, ISSN : 1077-3142. DOI: 10.1016/j.cviu.2013.01.013 . [Online]. Available: http://dx.doi.org/10.1016/j.cviu.2013.01.013

  4. [8]

    No classification without representation: Assessing geodiversity issues in open data sets for the developing world,

    S. Shankar, Y . Halpern, E. Breck, J. Atwood, J. Wilson, and D. Sculley, “No classification without representation: Assessing geodiversity issues in open data sets for the developing world,” arXiv preprint arXiv:1711.08536, 2017

  5. [9]

    Playing hard exploration games by watching youtube,

    Y . Aytar, T. Pfaff, D. Budden, T. Paine, Z. Wang, and N. de Freitas, “Playing hard exploration games by watching youtube,” in Advances in Neural Information Process- ing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31, Curran Associates, Inc., 2018. [Online]. Available: https : / / proceedings . n...

  6. [10]

    Data statements for natural language processing: Toward mitigating system bias and enabling better science,

    E. M. Bender and B. Friedman, “Data statements for natural language processing: Toward mitigating system bias and enabling better science,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 587–604, 2018. DOI: 10.1162/tacl_a_00041. [Online]. Available: https://aclanthology.org/Q18-1041

  7. [12]

    Gender shades: Intersectional accuracy disparities in commer- cial gender classification,

    J. Buolamwini and T. Gebru, “Gender shades: Intersectional accuracy disparities in commer- cial gender classification,” in Proceedings of the 1st Conference on Fairness, Accountability and Transparency, S. A. Friedler and C. Wilson, Eds., ser. Proceedings of Machine Learning Research, vol. 81, PMLR, 2018, pp. 77–91. [Online]. Available:https://proceedings...

  8. [13]

    Common voice: A massively-multilingual speech corpus,

    R. Ardila, M. Branson, K. Davis, et al., “Common voice: A massively-multilingual speech corpus,” arXiv preprint arXiv:1912.06670, 2019

Show all 294 references
  1. [14]

    Does object recognition work for everyone?

    T. De Vries, I. Misra, C. Wang, and L. Van der Maaten, “Does object recognition work for everyone?” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 52–59

  2. [15]

    Visual to text: Survey of image and video captioning,

    S. Li, Z. Tao, K. Li, and Y . Fu, “Visual to text: Survey of image and video captioning,”IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 3, no. 4, pp. 297–312,

  3. [16]

    Mundane content on social media: Creation, circulation, and the copyright problem,

    J. Meese and J. Hagedorn, “Mundane content on social media: Creation, circulation, and the copyright problem,” Social Media+ Society, vol. 5, no. 2, p. 2 056 305 119 839 190, 2019. 11 The Data Provenance Initiative, 2024

  4. [17]

    Model cards for model reporting,

    M. Mitchell, S. Wu, A. Zaldivar, et al., “Model cards for model reporting,” in Proceedings of the conference on fairness, accountability, and transparency, 2019, pp. 220–229

  5. [18]

    Moments in time dataset: One million videos for event understanding,

    M. Monfort, A. Andonian, B. Zhou, et al., “Moments in time dataset: One million videos for event understanding,” IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 2, pp. 502–508, 2019

  6. [19]

    Social data: Biases, methodological pitfalls, and ethical boundaries,

    A. Olteanu, C. Castillo, F. Diaz, and E. Kıcıman, “Social data: Biases, methodological pitfalls, and ethical boundaries,” Frontiers in big data, vol. 2, p. 13, 2019

  7. [20]

    Common voice: A massively-multilingual speech corpus,

    R. Ardila, M. Branson, K. Davis, et al., “Common voice: A massively-multilingual speech corpus,” English, in Proceedings of the Twelfth Language Resources and Evaluation Confer- ence, N. Calzolari, F. Béchet, P. Blache,et al., Eds., Marseille, France: European Language Resourc...

  8. [21]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: h...

  9. [22]

    Semantic visual navigation by watching youtube videos,

    M. Chang, A. Gupta, and S. Gupta, “Semantic visual navigation by watching youtube videos,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, Curran Associates, Inc., 2020, pp. 4283–

  10. [23]

    The pile: An 800gb dataset of diverse text for language modeling,

    L. Gao, S. Biderman, S. Black, et al., “The pile: An 800gb dataset of diverse text for language modeling,” arXiv preprint arXiv:2101.00027, 2020

  11. [24]

    Henighan, J

    T. Henighan, J. Kaplan, M. Katz, et al., Scaling laws for autoregressive generative modeling,

  12. [25]

    The state and fate of linguistic diversity and inclusion in the nlp world,

    P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, “The state and fate of linguistic diversity and inclusion in the nlp world,” arXiv preprint arXiv:2004.09095, 2020

  13. [26]

    Scaling laws for neural language models,

    J. Kaplan, S. McCandlish, T. Henighan, et al., “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361, 2020

  14. [27]

    Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages,

    A. Kunchukuttan, D. Kakwani, S. Golla, A. Bhattacharyya, M. M. Khapra, P. Kumar,et al., “Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages,” arXiv preprint arXiv:2005.00085, 2020

  15. [28]

    Beyond “i agree

    E. P. Robinson and Y . Zhu, “Beyond “i agree”: Users’ understanding of web site terms of service,” Social media+ society, vol. 6, no. 1, p. 2 056 305 119 897 321, 2020

  16. [29]

    The new legal landscape for text mining and machine learning,

    M. J. Sag, “The new legal landscape for text mining and machine learning,” in Journal of the Copyright Society of the USA, 2020

  17. [30]

    Y . Zhu, X. Li, C. Liu,et al., A comprehensive study of deep video action recognition, 2020. arXiv: 2012.06567 [cs.CV] . [Online]. Available: https://arxiv.org/abs/ 2012.06567

  18. [31]

    Masakhaner: Named entity recognition for african languages,

    D. I. Adelani, J. Abbott, G. Neubig, et al., “Masakhaner: Named entity recognition for african languages,” Transactions of the Association for Computational Linguistics, vol. 9, pp. 1116–1131, 2021

  19. [32]

    How might we create better benchmarks for speech recognition?

    A. Aksënova, D. van Esch, J. Flynn, and P. Golik, “How might we create better benchmarks for speech recognition?” In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, K. Church, M. Liberman, and V . Kordoni, Eds., Online: Association for Computational ...

  20. [33]

    Xls-r: Self-supervised cross-lingual speech representa- tion learning at scale,

    A. Babu, C. Wang, A. Tjandra, et al., “Xls-r: Self-supervised cross-lingual speech representa- tion learning at scale,” arXiv preprint arXiv:2111.09296, 2021

  21. [34]

    Addressing “documentation debt

    J. Bandy and N. Vincent, “Addressing “documentation debt” in machine learning research: A retrospective datasheet for bookcorpus,” arXiv preprint arXiv:2105.05241, 2021. 12 The Data Provenance Initiative, 2024

  22. [35]

    Multimodal datasets: Misogyny, pornography, and malignant stereotypes,

    A. Birhane, V . U. Prabhu, and E. Kahembwe, “Multimodal datasets: Misogyny, pornography, and malignant stereotypes,” arXiv preprint arXiv:2110.01963, 2021

  23. [36]

    Quality at a glance: An audit of web-crawled multilingual datasets,

    I. Caswell, J. Kreutzer, L. Wang, et al., “Quality at a glance: An audit of web-crawled multilingual datasets,” arXiv preprint arXiv:2103.12028, 2021

  24. [37]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus,

    J. Dodge, M. Sap, A. Marasovi´c, et al., “Documenting large webtext corpora: A case study on the colossal clean crawled corpus,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 1286–1305

  25. [38]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov,et al., An image is worth 16x16 words: Transformers for image recognition at scale, 2021. arXiv: 2010.11929 [cs.CV]

  26. [39]

    Datasheets for datasets,

    T. Gebru, J. Morgenstern, B. Vecchione,et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021

  27. [40]

    What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus,

    A. S. Luccioni and J. D. Viviano, “What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus,” 2021. arXiv: 2105.02732 [cs.CL]

  28. [41]

    Understanding gender and racial disparities in image recognition models,

    R. Mahadev and A. Chakravarti, “Understanding gender and racial disparities in image recognition models,” arXiv preprint arXiv:2107.09211, 2021

  29. [42]

    Automatic speech recognition: A survey,

    M. Malik, M. K. Malik, K. Mehmood, and I. Makhdoom, “Automatic speech recognition: A survey,” Multimedia Tools and Applications, vol. 80, pp. 9411–9457, 2021

  30. [43]

    Monfort, S

    M. Monfort, S. Jin, A. Liu, et al., Spoken Moments: Learning Joint Audio-Visual Representa- tions from Video Descriptions, arXiv:2105.04489 [cs, eess], 2021. DOI: 10.48550/arXiv. 2105.04489. [Online]. Available: http://arxiv.org/abs/2105.04489 (visited on 05/02/2024)

  31. [44]

    Data and its (dis) contents: A survey of dataset development and use in machine learning research,

    A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis) contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11, 2021

  32. [45]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, et al., “Learning transferable visual models from natural language supervision,” arXiv preprint arXiv:2103.00020, 2021

  33. [46]

    Changing the world by changing the data,

    A. Rogers, “Changing the world by changing the data,” inProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online: Association for Computati...

  34. [47]

    “Everyone wants to do the model work, not the data work

    N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo, ““Everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI,” in CHI, ser. CHI ’21, Yokohama, Japan: Association for Computing Machinery, 2021, ISBN : 9781450380966....

  35. [48]

    Multitask prompted training enables zero-shot task generalization,

    V . Sanh, A. Webson, C. Raffel,et al., “Multitask prompted training enables zero-shot task generalization,” ICLR 2022, 2021. [Online]. Available: https://arxiv.org/abs/ 2110.08207

  36. [49]

    Finetuned language models are zero-shot learners,

    J. Wei, M. Bosma, V . Zhao,et al., “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations, 2021

  37. [50]

    Challenges in detoxifying language models,

    J. Welbl, A. Glaese, J. Uesato,et al., “Challenges in detoxifying language models,” inFindings of the Association for Computational Linguistics: EMNLP 2021, 2021, pp. 2447–2469

  38. [51]

    Detoxifying language models risks marginalizing minority voices,

    A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein, “Detoxifying language models risks marginalizing minority voices,” inProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...

  39. [52]

    Masader: Metadata sourcing for arabic text and speech data resources,

    Z. Alyafeai, M. Masoud, M. Ghaleb, and M. S. Al-shaibani, “Masader: Metadata sourcing for arabic text and speech data resources,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2022, pp. 6340–6351

  40. [53]

    Quantifying memo- rization across neural language models,

    N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying memo- rization across neural language models,” 2022. arXiv: 2202.07646 [cs.LG]. 13 The Data Provenance Initiative, 2024

  41. [54]

    Elizalde, S

    B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, Clap: Learning audio concepts from natural language supervision, 2022. arXiv: 2206.04769 [cs.SD]

  42. [55]

    Dataset geography: Mapping language data to language users,

    F. Faisal, Y . Wang, and A. Anastasopoulos, “Dataset geography: Mapping language data to language users,” in Proceedings of the 60th Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), 2022, pp. 3381–3411

  43. [56]

    The flores-101 evaluation benchmark for low- resource and multilingual machine translation,

    N. Goyal, C. Gao, V . Chaudhary, et al., “The flores-101 evaluation benchmark for low- resource and multilingual machine translation,” Transactions of the Association for Computa- tional Linguistics, vol. 10, pp. 522–538, 2022

  44. [57]

    Training compute-optimal large language models,

    J. Hoffmann, S. Borgeaud, A. Mensch, et al., “Training compute-optimal large language models,” arXiv preprint arXiv:2203.15556, 2022

  45. [58]

    Leakage and the reproducibility crisis in ml-based science,

    S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in ml-based science,” arXiv preprint arXiv:2207.07048, 2022

  46. [59]

    Quality at a glance: An audit of web-crawled multilingual datasets,

    J. Kreutzer, I. Caswell, L. Wang, et al., “Quality at a glance: An audit of web-crawled multilingual datasets,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 50–72, 2022

  47. [60]

    The bigscience roots corpus: A 1.6tb composite multilingual dataset,

    H. Laurençon, L. Saulnier, T. Wang,et al., “The bigscience roots corpus: A 1.6tb composite multilingual dataset,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35, Curran Associates, Inc., 20...

  48. [61]

    McMillan-Major, Z

    A. McMillan-Major, Z. Alyafeai, S. Biderman, et al., Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,

  49. [62]

    Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,

    A. McMillan-Major, Z. Alyafeai, S. Biderman, et al., “Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,” arXiv preprint arXiv:2201.10066, 2022

  50. [63]

    Moctezuma, T

    D. Moctezuma, T. Ramírez-delReal, G. Ruiz, and O. González-Chávez, Video captioning: A comparative review of where we are and which could be the route , 2022. arXiv: 2204. 05976 [cs.CV]. [Online]. Available: https://arxiv.org/abs/2204.05976

  51. [64]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, et al., “Training language models to follow instructions with human feedback,” arXiv preprint arXiv:2203.02155, 2022. [Online]. Available: https: //arxiv.org/abs/2203.02155

  52. [65]

    Hierarchical Text-Conditional Image Generation with CLIP Latents,

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv: arXiv:2204.06125, 2022. DOI: 10.48550/ arXiv.2204.06125. arXiv: 2204.06125 [cs]

  53. [66]

    Rando, D

    J. Rando, D. Paleka, D. Lindner, L. Heim, and F. Tramèr,Red-teaming the stable diffusion safety filter, 2022. arXiv:2210.04610 [cs.AI]. [Online]. Available:https://arxiv. org/abs/2210.04610

  54. [67]

    Make-A-Video: Text-to-Video Generation without Text-Video Data,

    U. Singer, A. Polyak, T. Hayes, et al., “Make-A-Video: Text-to-Video Generation without Text-Video Data,” arXiv: arXiv:2209.14792, 2022. arXiv:2209.14792 [cs]. [Online]. Available: http://arxiv.org/abs/2209.14792

  55. [68]

    Bigssl: Exploring the frontier of large-scale semi- supervised learning for automatic speech recognition,

    Y . Zhang, D. S. Park, W. Han, et al., “Bigssl: Exploring the frontier of large-scale semi- supervised learning for automatic speech recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1519–1532, 2022, ISSN : 1941-0484. DOI: 10.1109/ jstsp.2...

  56. [69]

    Survey of video object detection algorithms based on deep learning,

    L. Zheng, T. Zhou, R. Jiang, and Y . Peng, “Survey of video object detection algorithms based on deep learning,” in Proceedings of the 2021 4th International Conference on Algorithms, Computing and Artificial Intelligence, ser. ACAI ’21, Sanya, China: Association for Com- puti...

  57. [70]

    Scaling laws for generative mixed-modal language models,

    A. Aghajanyan, L. Yu, A. Conneau,et al., “Scaling laws for generative mixed-modal language models,” in International Conference on Machine Learning, PMLR, 2023, pp. 265–279

  58. [71]

    Into the laions den: Investigating hate in multimodal datasets,

    A. Birhane, V . Prabhu, S. Han, V . N. Boddeti, and A. S. Luccioni, “Into the laions den: Investigating hate in multimodal datasets,”arXiv preprint arXiv:2311.03449, 2023

  59. [72]

    Blattmann, T

    A. Blattmann, T. Dockhorn, S. Kulal, et al., Stable video diffusion: Scaling latent video diffu- sion models to large datasets, 2023. arXiv: 2311.15127 [cs.CV]. [Online]. Available: https://arxiv.org/abs/2311.15127

  60. [73]

    Bommasani, K

    R. Bommasani, K. Klyman, S. Longpre, et al., The foundation model transparency index,

  61. [74]

    Quantifying mem- orization across neural language models,

    N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramèr, and C. Zhang, “Quantifying mem- orization across neural language models,” in The Eleventh International Conference on Learning Representations, OpenReview, 2023

  62. [75]

    Extracting training data from diffusion models,

    N. Carlini, J. Hayes, M. Nasr, et al., “Extracting training data from diffusion models,” in32nd USENIX Security Symposium (USENIX Security 23), Anaheim, CA: USENIX Association, 2023, pp. 5253–5270, ISBN : 978-1-939133-37-3. [Online]. Available: https : / / www . usenix.org/con...

  63. [76]

    S. H. Cen, A. Hopkins, A. Ilyas, A. Madry, I. Struckman, and L. Videgaray Caso, AI Supply Chains, 2023. [Online]. Available: http://dx.doi.org/10.2139/ssrn.4789403

  64. [77]

    Gender bias in hiring: An analysis of the impact of amazon’s recruiting algorithm,

    X. Chang, “Gender bias in hiring: An analysis of the impact of amazon’s recruiting algorithm,” Advances in Economics, Management and Political Sciences, vol. 23, pp. 134–140, 2023. DOI: 10.54254/2754-1169/23/20230367

  65. [78]

    Can language models be instructed to protect personal information?

    Y . Chen, E. Mendes, S. Das, W. Xu, and A. Ritter, “Can language models be instructed to protect personal information?” en, 2023

  66. [79]

    Dialect corpora from youtube,

    S. Coats, “Dialect corpora from youtube,” Language and linguistics in a complex world , 2023

  67. [80]

    Ai image training dataset found to include child sexual abuse imagery,

    E. David, “Ai image training dataset found to include child sexual abuse imagery,”The Verge, 2023, 7:57 AM PST. [Online]. Available: https://www.theverge.com/2023/12/ 20/24009418/generative-ai-image-laion-csam-google-stability- stanford

  68. [81]

    What’s in my big data?

    Y . Elazar, A. Bhagia, I. H. Magnusson, et al., “What’s in my big data?” In The Twelfth International Conference on Learning Representations, 2023

  69. [82]

    Esser, J

    P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis,Structure and content- guided video synthesis with diffusion models, 2023. arXiv: 2302.03011 [cs.CV]. [On- line]. Available: https://arxiv.org/abs/2302.03011

  70. [83]

    Datacomp: In search of the next generation of multi- modal datasets,

    S. Y . Gadre, G. Ilharco, A. Fang,et al., “Datacomp: In search of the next generation of multi- modal datasets,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36, Curran Associates, Inc., 2...

  71. [84]

    Foundation models and fair use,

    P. Henderson, X. Li, D. Jurafsky, T. Hashimoto, M. A. Lemley, and P. Liang, “Foundation models and fair use,” arXiv preprint arXiv:2303.15715, 2023

  72. [85]

    Understanding catastrophic forgetting in language models via implicit inference,

    S. Kotha, J. M. Springer, and A. Raghunathan, “Understanding catastrophic forgetting in language models via implicit inference,” arXiv preprint arXiv:2309.10105, 2023

  73. [86]

    Harnessing large- language models to generate private synthetic text,

    A. Kurakin, N. Ponomareva, U. Syed, L. MacDermed, and A. Terzis, “Harnessing large- language models to generate private synthetic text,” 2023. arXiv:2306.01684 [cs.LG]

  74. [87]

    Platypus: Quick, cheap, and powerful refinement of llms,

    A. N. Lee, C. J. Hunter, and N. Ruiz, “Platypus: Quick, cheap, and powerful refinement of llms,” NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023

  75. [88]

    Talkin”bout ai generation: Copyright and the generative-ai supply chain,

    K. Lee, A. F. Cooper, and J. Grimmelmann, “Talkin”bout ai generation: Copyright and the generative-ai supply chain,” arXiv preprint arXiv:2309.08133, 2023

  76. [89]

    Yodas: Youtube- oriented dataset for audio and speech,

    X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watanabe, “Yodas: Youtube- oriented dataset for audio and speech,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, 2023, pp. 1–8. 15 The Data Provenance Initiative, 2024

  77. [90]

    H. Liu, C. Li, Y . Li, and Y . J. Lee, Improved baselines with visual instruction tuning, 2023

  78. [91]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, 2023

  79. [92]

    Longpre, G

    S. Longpre, G. Yauney, E. Reif, et al., A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023. arXiv: 2305.13169 [cs.CL]

  80. [93]

    Discit ergo est: Training data provenance and fair use,

    R. Mahari and S. Longpre, “Discit ergo est: Training data provenance and fair use,”Robert Mahari and Shayne Longpre, Discit ergo est: Training Data Provenance And Fair Use, Dynamics of Generative AI (ed. Thibault Schrepel & Volker Stocker), Network Law Review, Winter, 2023

  81. [94]

    Mahari, L

    R. Mahari, L. Shayne, L. Donewald, A. Polozov, A. ’. Pentland, and A. Lipsitz, Comment to US copyright office on data provenance and copyright, 2023

  82. [95]

    Marion, A

    M. Marion, A. Üstün, L. Pozzobon, A. Wang, M. Fadaee, and S. Hooker, When less is more: Investigating data pruning for pretraining llms at scale, 2023. arXiv: 2309.04564 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2309.04564

  83. [96]

    Silo language models: Isolating legal risk in a nonparametric datastore,

    S. Min, S. Gururangan, E. Wallace, et al., “Silo language models: Isolating legal risk in a nonparametric datastore,” in NeurIPS 2023 Workshop on Distribution Shifts: New Frontiers with Foundation Models, 2023

  84. [97]

    Licensed to learn: Mitigating copyright infringement liability of generative ai systems through contracts,

    F. Morton-Park, “Licensed to learn: Mitigating copyright infringement liability of generative ai systems through contracts,” Notre Dame Journal on Emerging Technology, vol. 5, p. 64, 2023

  85. [98]

    Crosslingual generalization through multitask finetuning,

    N. Muennighoff, T. Wang, L. Sutawika,et al., “Crosslingual generalization through multitask finetuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15 991–16 111

  86. [99]

    The RefinedWeb dataset for falcon LLM: Outperforming curated corpora with web data, and web data only,

    G. Penedo, Q. Malartic, D. Hesslow, et al., “The RefinedWeb dataset for falcon LLM: Outperforming curated corpora with web data, and web data only,” 2023. arXiv: 2306. 01116 [cs.CL]

  87. [100]

    Reproducing whisper-style training using an open-source toolkit and publicly available data,

    Y . Peng, J. Tian, B. Yan,et al., “Reproducing whisper-style training using an open-source toolkit and publicly available data,” in 2023 IEEE Automatic Speech Recognition and Under- standing Workshop (ASRU), IEEE, 2023, pp. 1–8

  88. [101]

    Porgali, V

    B. Porgali, V . Albiero, J. Ryda, C. C. Ferrer, and C. Hazirbas,The casual conversations v2 dataset, 2023. arXiv: 2303.04838 [cs.CV] . [Online]. Available: https://arxiv. org/abs/2303.04838

  89. [102]

    Pozzobon, B

    L. Pozzobon, B. Ermis, P. Lewis, and S. Hooker, Goodtriever: Adaptive toxicity mitiga- tion with retrieval-augmented models , 2023. arXiv: 2310 . 07589 [cs.AI]. [Online]. Available: https://arxiv.org/abs/2310.07589

  90. [103]

    End-to-end speech recognition: A survey,

    R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schlüter, and S. Watanabe, “End-to-end speech recognition: A survey,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023

  91. [104]

    Robust speech recognition via large-scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning, PMLR, 2023, pp. 28 492–28 518

  92. [105]

    Direct pref- erence optimization: Your language model is secretly a reward model,

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct pref- erence optimization: Your language model is secretly a reward model,” arXiv preprint arXiv:2305.18290, 2023

  93. [106]

    Self-supervised learning for videos: A survey,

    M. C. Schiappa, Y . S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,” ACM Computing Surveys , vol. 55, no. 13s, pp. 1–37, 2023, ISSN : 1557-7341. DOI: 10 . 1145/3577925. [Online]. Available: http://dx.doi.org/10.1145/3577925

  94. [107]

    Detecting personal information in training corpora: An analysis,

    N. Subramani, S. Luccioni, J. Dodge, and M. Mitchell, “Detecting personal information in training corpora: An analysis,” in Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023), Toronto, Canada: Association for Computational Linguistics, 2023

  95. [108]

    Gemini: A family of highly capable multimodal models,

    G. Team, R. Anil, S. Borgeaud, et al., “Gemini: A family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023. 16 The Data Provenance Initiative, 2024

  96. [109]

    Youtube-asl: A large-scale, open-domain ameri- can sign language-english parallel corpus,

    D. Uthus, G. Tanzer, and M. Georg, “Youtube-asl: A large-scale, open-domain ameri- can sign language-english parallel corpus,” in Advances in Neural Information Process- ing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36, Curran As...

  97. [110]

    Tag your fish in the broken net: A responsible web framework for protecting online privacy and copyright,

    D. Zhang, B. Xia, Y . Liu,et al., “Tag your fish in the broken net: A responsible web framework for protecting online privacy and copyright,” 2023. arXiv:2310.07915 [cs.NI]

  98. [111]

    C. Zhu, Q. Jia, W. Chen, Y . Guo, and Y . Liu,Deep learning for video-text retrieval: A review,

  99. [112]

    Ahmadian, B

    Aakanksha, A. Ahmadian, B. Ermis,et al., The multilingual alignment prism: Aligning global and local preferences to reduce harm , 2024. arXiv: 2406.18682 [cs.CL] . [Online]. Available: https://arxiv.org/abs/2406.18682

  100. [113]

    D. I. Adelani, J. Ojo, I. A. Azime, et al., Irokobench: A new benchmark for african languages in the age of large language models , 2024. arXiv: 2406 . 03368 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2406.03368

  101. [114]

    A survey on data selection for language models,

    A. Albalak, Y . Elazar, S. M. Xie,et al., “A survey on data selection for language models,” arXiv preprint arXiv:2402.16827, 2024

  102. [115]

    Video generation models as world simulators,

    T. Brooks, B. Peebles, C. Holmes,et al., “Video generation models as world simulators,” 2024. [Online]. Available: https : / / openai . com / research / video - generation - models-as-world-simulators

  103. [116]

    Nvidia sued for scraping youtube after 404 media investigation,

    S. Cole, “Nvidia sued for scraping youtube after 404 media investigation,” 404 Media,

  104. [117]

    [Online]

    arXiv: 2302.12552 [cs.CV] . [Online]. Available: https://arxiv.org/ abs/2302.12552

  105. [118]

    Datacomp: In search of the next generation of multimodal datasets,

    S. Y . Gadre, G. Ilharco, A. Fang, et al., “Datacomp: In search of the next generation of multimodal datasets,” Advances in Neural Information Processing Systems, vol. 36, 2024

  106. [119]

    Klyman, Acceptable use policies for foundation models , 2024

    K. Klyman, Acceptable use policies for foundation models , 2024. arXiv: 2409.09041 [cs.CY]. [Online]. Available: https://arxiv.org/abs/2409.09041

  107. [120]

    Best practices and lessons learned on synthetic data,

    R. Liu, J. Wei, F. Liu, et al., “Best practices and lessons learned on synthetic data,” 2024. arXiv: 2404.07503 [cs.CL]

  108. [121]

    Datasets for large language models: A compre- hensive survey,

    Y . Liu, J. Cao, C. Liu, K. Ding, and L. Jin, “Datasets for large language models: A compre- hensive survey,”arXiv preprint arXiv:2402.18041, 2024

  109. [122]

    The responsible foundation model development cheatsheet: A review of tools & resources,

    S. Longpre, S. Biderman, A. Albalak, et al., “The responsible foundation model development cheatsheet: A review of tools & resources,” arXiv preprint arXiv:2406.16746, 2024

  110. [123]

    A large-scale audit of dataset licensing and attribution in AI,

    S. Longpre, R. Mahari, A. Chen,et al., “A large-scale audit of dataset licensing and attribution in AI,” Nature Machine Intelligence, vol. 6, no. 8, pp. 975–987, 2024. DOI: 10/gt8f5p. arXiv: 2310.16787 [cs]

  111. [124]

    Nvlm: Open frontier-class multimodal llms,

    W. Dai, N. Lee, B. Wang,et al., “Nvlm: Open frontier-class multimodal llms,”arXiv preprint, 2024

  112. [125]

    Data authenticity, consent, & provenance for ai are all broken: What will it take to fix them?

    S. Longpre, R. Mahari, N. Obeng-Marnu, et al., “Data authenticity, consent, & provenance for ai are all broken: What will it take to fix them?” arXiv preprint arXiv:2404.12691, 2024

  113. [126]

    Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages,

    H. Lovenia, R. Mahendra, S. M. Akbar, et al., “Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages,” arXiv preprint arXiv:2406.10118, 2024

  114. [127]

    Mauran, What was Sora trained on? Creatives demand answers

    C. Mauran, What was Sora trained on? Creatives demand answers. https://mashable. com / article / openai - sora - ai - video - generator - training - data, [Accessed 28-09-2024], 2024. 17 The Data Provenance Initiative, 2024

  115. [128]

    Topics, au- thors, and institutions in large language model research: Trends from 17k arxiv papers,

    R. Movva, S. Balachandar, K. Peng, G. Agostini, N. Garg, and E. Pierson, “Topics, au- thors, and institutions in large language model research: Trends from 17k arxiv papers,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computation...

  116. [129]

    [Online]

    OpenAI, Hello gpt-4o: We’re announcing gpt-4o, our new flagship model that can reason across audio, vision, and text in real time.2024. [Online]. Available: https://openai. com/index/hello-gpt-4o/

  117. [130]

    Data, data everywhere: A guide for pretraining dataset construction,

    J. Parmar, S. Prabhumoye, J. Jennings, et al., “Data, data everywhere: A guide for pretraining dataset construction,” arXiv preprint 2407.06380, 2024

  118. [131]

    Consent in crisis: The rapid decline of the ai data commons,

    S. Longpre, R. Mahari, A. Lee, et al., “Consent in crisis: The rapid decline of the ai data commons,” arXiv preprint arXiv:2407.14933, 2024

  119. [132]

    Anatomy of industrial scale multilingual asr,

    F. M. Ramirez, L. Chkhetiani, A. Ehrenberg, et al., “Anatomy of industrial scale multilingual asr,” arXiv preprint arXiv:2404.09841, 2024

  120. [133]

    Romanou, N

    A. Romanou, N. Foroutan, A. Sotnikova, et al., Include: Evaluating multilingual language understanding with regional knowledge, 2024. arXiv: 2411.19799 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2411.19799

  121. [134]

    Singh, F

    S. Singh, F. Vargus, D. Dsouza,et al., Aya dataset: An open-access collection for multilingual instruction tuning, 2024. arXiv: 2402.06619 [cs.CL]

  122. [135]

    Openai sued over using youtube videos without creators’ consent,

    S. Skolnik, “Openai sued over using youtube videos without creators’ consent,”Bloomberg Law, 2024. [Online]. Available: https : / / news . bloomberglaw . com / litigation / openai - sued - over - using - youtube - videos - without - creators-consent

  123. [136]

    Dolma: An open corpus of three trillion tokens for language model pretraining research,

    L. Soldaini, R. Kinney, A. Bhagia, et al., “Dolma: An open corpus of three trillion tokens for language model pretraining research,” arXiv preprint arXiv:2402.00159, 2024

  124. [137]

    Aya model: An instruction finetuned open-access multilingual language model,

    A. Üstün, V . Aryabumi, Z.-X. Yong, et al., “Aya model: An instruction finetuned open-access multilingual language model,” arXiv preprint arXiv:2402.07827, 2024

  125. [138]

    Scaling speech technology to 1,000+ languages,

    V . Pratap, A. Tjandra, B. Shi,et al., “Scaling speech technology to 1,000+ languages,”Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024

  126. [139]

    X. Yang, W. Liang, and J. Zou, Navigating dataset documentations in ai: A large-scale analysis of dataset cards on hugging face, 2024. arXiv: 2401.13822 [cs.LG]. [Online]. Available: https://arxiv.org/abs/2401.13822

  127. [140]

    Commercial

    Z. Zheng, X. Peng, T. Yang, et al., Open-sora: Democratizing efficient video production for all, 2024. [Online]. Available: https://github.com/hpcaitech/Open-Sora. 18 The Data Provenance Initiative, 2024 LABEL DEFINITION MODEL CLOSED A model used to generate part or all of the...

  128. [141]

    Chalkidis, I

    I. Chalkidis, I. Androutsopoulos, and N. Aletras, Neural legal judgment prediction in english,

  129. [142]

    Chalkidis, M

    I. Chalkidis, M. Fergadiotis, P. Malakasiotis, and I. Androutsopoulos,Large-scale multi-label text classification on EU legislation, 2019. arXiv: 1906.02192[cs]. [Online]. Available: http://arxiv.org/abs/1906.02192 (visited on 05/29/2024)

  130. [143]

    W. Chen, J. Chen, P. Qin, X. Yan, and W. Y . Wang,Semantically conditioned dialog response generation via hierarchical disentangled self-attention, 2019. arXiv: 1905.12866[cs]. [Online]. Available: http://arxiv.org/abs/1905.12866 (visited on 05/01/2024)

  131. [144]

    Look before you hop: Conversational question answering over knowledge graphs using judicious context expansion,

    P. Christmann, R. S. Roy, A. Abujabal, J. Singh, and G. Weikum, “Look before you hop: Conversational question answering over knowledge graphs using judicious context expansion,” in Proceedings of the 28th ACM International Conference on Information and Knowledge Management, 20...

  132. [145]

    Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models,

    W. Wang and Y . Yang, “Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models,” arXiv preprint arXiv:2403.06098, 2024

  133. [146]

    A corpus of regional american language from YouTube,

    S. Coats, “A corpus of regional american language from YouTube,” in Proceed- ings of the Digital Humanities in the Nordic Countries 5th Conference , 2019. [On- line]. Available: https : / / www . semanticscholar . org / paper / A - Corpus - of - Regional - American - Language ...

  134. [147]

    Toyota smarthome: Real-world activities of daily living,

    S. Das, R. Dai, M. Koperski, et al., “Toyota smarthome: Real-world activities of daily living,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , Seoul, Korea (South): IEEE, 2019, pp. 833–842, ISBN : 978-1-72814-803-8. DOI: 10/ghfjc7. [Online]. Available: h...

  135. [148]

    The ATIS spoken language systems pilot corpus,

    C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, “The ATIS spoken language systems pilot corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Hid- den Valley, Pennsylvania, June 24-27,1990, 1990. DOI: 10/cz3442. [Online]. Available: https://aclantholo...

  136. [149]

    SWITCHBOARD: Telephone speech corpus for research and development,

    J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in [Proceedings] ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing, San Francisco, CA, USA: IEEE, 1992, 517–520 vol.1, ISBN : ...

  137. [150]

    Garofolo, John S., Lamel, Lori F., Fisher, William M., et al., TIMIT acoustic-phonetic continuous speech corpus, Artwork Size: 715776 KB Pages: 715776 KB, 1993. DOI: 10. 35111/17GK- BN40. [Online]. Available: https://catalog.ldc.upenn.edu/ LDC93S1 (visited on 05/01/2024)

  138. [151]

    Can prosody aid the automatic classification of dialog acts in conversational speech?

    E. Shriberg, R. Bates, A. Stolcke, et al., “Can prosody aid the automatic classification of dialog acts in conversational speech?” Language and Speech, vol. 41 ( Pt 3-4), pp. 443–492, 1998, ISSN : 0023-8309. DOI: 10.1177/002383099804100410

  139. [152]

    Dialogue act modeling for automatic tagging and recognition of conversational speech,

    A. Stolcke, K. Ries, N. Coccaro, et al., “Dialogue act modeling for automatic tagging and recognition of conversational speech,”Computational Linguistics, vol. 26, no. 3, pp. 339–373, 2000, ISSN : 0891-2017, 1530-9312. DOI: 10/dqmv4j . arXiv: cs/0006023 . [Online]. Available: ...

  140. [153]

    Corpus of spontaneous japanese: Its design and evaluation,

    K. Maekawa, “Corpus of spontaneous japanese: Its design and evaluation,” in Proceedings of the ISCA/IEEE Workshop on Spontaneous Speech Processing and Recognition, 2003, paper MMO2. [Online]. Available: https://www.isca- archive.org/sspr_2003/ maekawa03_sspr.html (visited on 0...

  141. [154]

    The fisher corpus: A resource for the next generations of speech-to-text,

    C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text,” in Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04), M. T. Lino, M. F. Xavier, F. Ferreira, R. Costa, and R. Silva, ...

  142. [155]

    Lander, T, CSLU: 22 languages corpus , 2005. DOI: 10 . 35111 / ZKN2 - 5X88. [On- line]. Available: https://catalog.ldc.upenn.edu/LDC2005S26 (visited on 05/29/2024)

  143. [156]

    Pang and L

    B. Pang and L. Lee, Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales , 2005. arXiv: cs / 0506075. [Online]. Available: http : //arxiv.org/abs/cs/0506075 (visited on 05/29/2024). 38 The Data Provenance Initiative, 2024

  144. [157]

    The AMI meeting corpus: A pre-announcement,

    J. Carletta, S. Ashby, S. Bourban, et al., “The AMI meeting corpus: A pre-announcement,” in Machine Learning for Multimodal Interaction, S. Renals and S. Bengio, Eds., vol. 3869, Series Title: Lecture Notes in Computer Science, Berlin, Heidelberg: Springer Berlin Heidelberg, 2...

  145. [158]

    Lander, T, CSLU: Foreign accented english release 1.2, Artwork Size: 1468006 KB Pages: 1468006 KB, 2007. DOI: 10 . 35111 / 0VWP - XN48. [Online]. Available: https : / / catalog.ldc.upenn.edu/LDC2007S08 (visited on 05/01/2024)

  146. [159]

    Measures of semantic sim- ilarity and relatedness in the biomedical domain,

    T. Pedersen, S. V . S. Pakhomov, S. Patwardhan, and C. G. Chute, “Measures of semantic sim- ilarity and relatedness in the biomedical domain,” Journal of Biomedical Informatics, vol. 40, no. 3, pp. 288–299, 2007, ISSN : 1532-0464. DOI: 10/fghjwr. [Online]. Available: https: //...

  147. [160]

    Actions in context,

    M. Marszalek, I. Laptev, and C. Schmid, “Actions in context,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL: IEEE, 2009, pp. 2929–2936, ISBN : 978-1-4244-3992-8. DOI: 10/d5bs7p. [Online]. Available: https://ieeexplore. ieee.org/document/5206557/...

  148. [161]

    What are they doing? : Collective activity classification using spatio-temporal relationship among people,

    Wongun Choi, K. Shahid, and S. Savarese, “What are they doing? : Collective activity classification using spatio-temporal relationship among people,” in 2009 IEEE 12th Interna- tional Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan: IEEE, 2009, pp. 1282–1...

  149. [162]

    Bradlow, ALLSSTAR: Archive of l1 and l2 scripted and spontaneous transcripts and recordings, 2010

    A. Bradlow, ALLSSTAR: Archive of l1 and l2 scripted and spontaneous transcripts and recordings, 2010. [Online]. Available: https : / / speechbox . linguistics . northwestern.edu/#!/?goto=allsstar (visited on 05/01/2024)

  150. [163]

    Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions,

    K. Ganesan, C. Zhai, and J. Han, “Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions,” in Proceedings of the 23rd International Conference on Computational Linguistics (Coling 2010), C.-R. Huang and D. Jurafsky, Eds., Beijing, China: C...

  151. [164]

    Semantic similarity and relatedness between clinical terms: An experimental study,

    S. Pakhomov, B. McInnes, T. Adam, Y . Liu, T. Pedersen, and G. B. Melton, “Semantic similarity and relatedness between clinical terms: An experimental study,” AMIA Annual Symposium Proceedings, vol. 2010, pp. 572–576, 2010,ISSN : 1942-597X. [Online]. Available: https://www.ncb...

  152. [165]

    Evaluation of topic identification methods on arabic corpora,

    M. Abbas, K. Smaïli, and D. Berkani, “Evaluation of topic identification methods on arabic corpora,” Journal of Digital Information Management, vol. 9, no. 5, 2011. [Online]. Available: https://www.dline.info/fpaper/jdim/v9i5/1.pdf (visited on 10/02/2024)

  153. [166]

    HMDB: A large video database for human motion recognition,

    H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: A large video database for human motion recognition,” in 2011 International Conference on Computer Vision , Barcelona, Spain: IEEE, 2011, pp. 2556–2563, ISBN : 978-1-4577-1102-2. DOI: 10/fxpf8k. [Online]. Availa...

  154. [167]

    Soomro, A

    K. Soomro, A. R. Zamir, and M. Shah, UCF101: A dataset of 101 human actions classes from videos in the wild, 2012. arXiv: 1212.0402[cs]. [Online]. Available: http:// arxiv.org/abs/1212.0402 (visited on 05/01/2024)

  155. [168]

    A thousand frames in just a few words: Lin- gual description of videos through latent topics and sparse object stitching,

    P. Das, C. Xu, R. F. Doell, and J. J. Corso, “A thousand frames in just a few words: Lin- gual description of videos through latent topics and sparse object stitching,” in 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA: IEEE, 2013, pp. 2634–...

  156. [169]

    Asgard: A portable architecture for multilin- gual dialogue systems,

    J. Liu, P. Pasupat, S. Cyphers, and J. Glass, “Asgard: A portable architecture for multilin- gual dialogue systems,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada: IEEE, 2013, pp. 8386–8390, ISBN : 978-1-4799- 0356-6. D...

  157. [170]

    P. Malo, A. Sinha, P. Takala, P. Korhonen, and J. Wallenius,Good debt or bad debt: Detecting semantic orientations in economic texts, 2013. arXiv: 1307.5336[cs,q-fin]. [Online]. Available: http://arxiv.org/abs/1307.5336 (visited on 05/29/2024)

  158. [171]

    Combining embedded accelerometers with computer vision for recognizing food preparation activities,

    S. Stein and S. J. McKenna, “Combining embedded accelerometers with computer vision for recognizing food preparation activities,” in Proceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing, ser. UbiComp ’13, New York, NY, USA: Associati...

  159. [172]

    Spatial pattern templates for recognition of objects with regu- lar structure,

    R. Tyleˇcek and R. Šára, “Spatial pattern templates for recognition of objects with regu- lar structure,” in Pattern Recognition, J. Weickert, M. Hein, and B. Schiele, Eds., Berlin, Heidelberg: Springer, 2013, pp. 364–374, ISBN : 978-3-642-40602-7. DOI: 10/ggwb5g

  160. [173]

    Bojanowski, R

    P. Bojanowski, R. Lajugie, F. Bach, et al., Weakly supervised action labeling in videos under ordering constraints, 2014. arXiv: 1407.1208[cs]. [Online]. Available: http: //arxiv.org/abs/1407.1208 (visited on 05/29/2024)

  161. [174]

    Creating summaries from user videos,

    M. Gygli, H. Grabner, H. Riemenschneider, and L. Van Gool, “Creating summaries from user videos,” in Computer Vision – ECCV 2014 , D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds., vol. 8695, Series Title: Lecture Notes in Computer Science, Cham: Springer International...

  162. [175]

    VideoStory: A new multimedia embedding for few-example recognition and translation of events,

    A. Habibian, T. Mensink, and C. G. Snoek, “VideoStory: A new multimedia embedding for few-example recognition and translation of events,” in Proceedings of the 22nd ACM international conference on Multimedia, Orlando Florida USA: ACM, 2014, pp. 17–26,ISBN : 978-1-4503-3063-3. ...

  163. [176]

    Large-scale video classification with convolutional neural networks,

    A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA: IEEE, 2014, pp. 1725– 1732, ISBN : 978-1-...

  164. [177]

    Free english and czech telephone speech corpus shared under the CC-BY-SA 3.0 license,

    M. Korvas, O. Plátek, O. Dušek, L. Žilka, and F. Jurˇcíˇcek, “Free english and czech telephone speech corpus shared under the CC-BY-SA 3.0 license,” in Proceedings of the Ninth In- ternational Conference on Language Resources and Evaluation (LREC’14), N. Calzolari, K. Choukri,...

  165. [178]

    The language of actions: Recovering the syntax and semantics of goal-directed human activities,

    H. Kuehne, A. Arslan, and T. Serre, “The language of actions: Recovering the syntax and semantics of goal-directed human activities,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA: IEEE, 2014, pp. 780–787, ISBN : 978-1-4799- 5118-5. DOI:...

  166. [179]

    StoryGraphs: Visualizing character interactions as a timeline,

    M. Tapaswi, M. Bauml, and R. Stiefelhagen, “StoryGraphs: Visualizing character interactions as a timeline,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 827–834. [Online]. Available: https://openaccess.thecvf. com/content_cvpr_201...

  167. [180]

    Extraction of relations between genes and diseases from text and large-scale data analysis: Implications for translational research,

    A. Bravo, J. Pinero, N. Queralt-Rosinach, M. Rautschka, and L. I. Furlong, “Extraction of relations between genes and diseases from text and large-scale data analysis: Implications for translational research,” BMC Bioinformatics, vol. 16, no. 1, p. 55, 2015, ISSN : 1471-2105. ...

  168. [181]

    Building large arabic multi-domain resources for sentiment analysis,

    H. ElSahar and S. R. El-Beltagy, “Building large arabic multi-domain resources for sentiment analysis,” in Computational Linguistics and Intelligent Text Processing, A. Gelbukh, Ed., Cham: Springer International Publishing, 2015, pp. 23–34, ISBN : 978-3-319-18117-2. DOI: 10/g6k58r

  169. [182]

    ActivityNet: A large-scale video benchmark for human activity understanding,

    F. C. Heilbron, V . Escorcia, B. Ghanem, and J. C. Niebles, “ActivityNet: A large-scale video benchmark for human activity understanding,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA: IEEE, 2015, pp. 961–970, ISBN : 978- 1-4673-69...

  170. [183]

    Librispeech: An ASR corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, Queensland, Australia: IEEE, 2015, pp. 5206–5210, IS...

  171. [184]

    Compositional semantic parsing on semi-structured tables,

    P. Pasupat and P. Liang, “Compositional semantic parsing on semi-structured tables,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C....

  172. [185]

    Rohrbach, M

    A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele,A dataset for movie description, 2015. arXiv: 1501.02530[cs] . [Online]. Available: http://arxiv.org/abs/1501. 02530 (visited on 05/01/2024)

  173. [186]

    An open/free database and benchmark for uyghur speaker recognition,

    A. Rozi, Dong Wang, Zhiyong Zhang, and T. F. Zheng, “An open/free database and benchmark for uyghur speaker recognition,” in 2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE), Sh...

  174. [187]

    A. M. Rush, S. Chopra, and J. Weston, A neural attention model for abstractive sentence summarization, 2015. arXiv: 1509.00685[cs]. [Online]. Available: http://arxiv. org/abs/1509.00685 (visited on 05/29/2024)

  175. [188]

    ImageNet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, et al., “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015, ISSN : 1573-

  176. [190]

    Wang and X

    D. Wang and X. Zhang, THCHS-30 : A free chinese speech corpus , 2015. arXiv: 1512. 01882[cs]. [Online]. Available: http://arxiv.org/abs/1512.01882 (visited on 05/01/2024)

  177. [191]

    Weston, A

    J. Weston, A. Bordes, S. Chopra, et al., Towards AI-complete question answering: A set of prerequisite toy tasks , 2015. arXiv: 1502 . 05698[cs , stat]. [Online]. Available: http://arxiv.org/abs/1502.05698 (visited on 05/29/2024)

  178. [192]

    TVSum: Summarizing web videos using titles,

    Yale Song, J. Vallmitjana, A. Stent, and A. Jaimes, “TVSum: Summarizing web videos using titles,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA: IEEE, 2015, pp. 5179–5187, ISBN : 978-1-4673-6964-0. DOI: 10/gfsj74. [Online]. Availabl...

  179. [193]

    Abu-El-Haija, N

    S. Abu-El-Haija, N. Kothari, J. Lee, et al., YouTube-8m: A large-scale video classification benchmark, 2016. arXiv: 1609 . 08675[cs]. [Online]. Available: http : / / arxiv . org/abs/1609.08675 (visited on 05/01/2024). 41 The Data Provenance Initiative, 2024

  180. [194]

    Alayrac, P

    J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien,Unsuper- vised learning from narrated instruction videos, 2016. arXiv: 1506.09215[cs]. [Online]. Available: http://arxiv.org/abs/1506.09215 (visited on 05/01/2024)

  181. [195]

    BRAD 1.0: Book reviews in arabic dataset,

    A. Elnagar and O. Einea, “BRAD 1.0: Book reviews in arabic dataset,” in2016 IEEE/ACS 13th International Conference of Computer Systems and Applications (AICCSA) , Agadir, Morocco: IEEE, 2016, pp. 1–8, ISBN : 978-1-5090-4320-0. DOI: 10/g6k6jm . [Online]. Available: http://ieeex...

  182. [196]

    Ibrahim, S

    M. Ibrahim, S. Muralidharan, Z. Deng, A. Vahdat, and G. Mori,A hierarchical deep temporal model for group activity recognition, 2016. arXiv: 1511.06040[cs]. [Online]. Available: http://arxiv.org/abs/1511.06040 (visited on 05/01/2024)

  183. [197]

    Jurczyk, M

    T. Jurczyk, M. Zhai, and J. D. Choi, SelQA: A new benchmark for selection-based question answering, 2016. arXiv: 1606.08513[cs]. [Online]. Available:http://arxiv.org/ abs/1606.08513 (visited on 10/02/2024)

  184. [198]

    I. A. El-khair, 1.5 billion words arabic corpus, 2016. arXiv: 1611.04033[cs]. [Online]. Available: http://arxiv.org/abs/1611.04033 (visited on 10/02/2024)

  185. [199]

    Lebret, D

    R. Lebret, D. Grangier, and M. Auli, Neural text generation from structured data with application to the biography domain, 2016. arXiv: 1603.07771[cs]. [Online]. Available: http://arxiv.org/abs/1603.07771 (visited on 05/29/2024)

  186. [200]

    Y . Li, Y . Song, L. Cao, et al. , TGIF: A new dataset and benchmark on animated GIF description, 2016. arXiv: 1604 . 02748[cs]. [Online]. Available: http : / / arxiv . org/abs/1604.02748 (visited on 05/01/2024)

  187. [201]

    R. Lowe, N. Pow, I. Serban, and J. Pineau, The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems, 2016. arXiv: 1506.08909[cs]. [Online]. Available: http://arxiv.org/abs/1506.08909 (visited on 10/02/2024)

  188. [202]

    Merity, C

    S. Merity, C. Xiong, J. Bradbury, and R. Socher,Pointer sentinel mixture models, 2016. arXiv: 1609.07843[cs]. [Online]. Available: http://arxiv.org/abs/1609.07843 (visited on 05/29/2024)

  189. [203]

    A benchmark dataset and evaluation methodology for video object segmentation,

    F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), ISSN: 1063-6919, 2016, pp. 724–732...

  190. [204]

    Rajpurkar, J

    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang,SQuAD: 100,000+ questions for machine comprehension of text, 2016. arXiv: 1606.05250[cs]. [Online]. Available: http:// arxiv.org/abs/1606.05250 (visited on 05/29/2024)

  191. [205]

    Rohrbach, A

    A. Rohrbach, A. Torabi, M. Rohrbach, et al. , Movie description , 2016. arXiv: 1605 . 03705[cs]. [Online]. Available: http : / / arxiv . org / abs / 1605 . 03705(vis- ited on 05/29/2024)

  192. [206]

    Recognizing fine-grained and composite activities using hand-centric features and script data,

    M. Rohrbach, A. Rohrbach, M. Regneri, et al., “Recognizing fine-grained and composite activities using hand-centric features and script data,” International Journal of Computer Vision, vol. 119, no. 3, pp. 346–373, 2016, ISSN : 0920-5691, 1573-1405. DOI: 10/f8w6kp. arXiv: 1502...

  193. [207]

    Shahroudy, J

    A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, NTU RGB+d: A large scale dataset for 3d human activity analysis , 2016. arXiv: 1604 . 02808[cs]. [Online]. Available: http : //arxiv.org/abs/1604.02808 (visited on 05/02/2024)

  194. [208]

    G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, Hollywood in homes: Crowdsourcing data collection for activity understanding , 2016. arXiv: 1604. 01753[cs]. [Online]. Available: http://arxiv.org/abs/1604.01753 (visited on 05/01/2024)

  195. [209]

    Tapaswi, Y

    M. Tapaswi, Y . Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler,MovieQA: Under- standing stories in movies through question-answering, 2016. arXiv: 1512.02902[cs]. [Online]. Available: http://arxiv.org/abs/1512.02902 (visited on 05/29/2024). 42 The Data Provenance...

  196. [210]

    MSR-VTT: A large video description dataset for bridging video and language,

    J. Xu, T. Mei, T. Yao, and Y . Rui, “MSR-VTT: A large video description dataset for bridging video and language,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), ISSN: 1063-6919, 2016, pp. 5288–5296. DOI: 10/ggv9gj. [Online]. Available: https://ieeex...

  197. [211]

    The value of semantic parse labeling for knowledge base question answering,

    W.-t. Yih, M. Richardson, C. Meek, M.-W. Chang, and J. Suh, “The value of semantic parse labeling for knowledge base question answering,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), K. Erk and N. A. Smith...

  198. [212]

    Zeng, T.-H

    K.-H. Zeng, T.-H. Chen, J. C. Niebles, and M. Sun, Title generation for user generated videos, 2016. arXiv: 1608.07068[cs]. [Online]. Available: http://arxiv.org/ abs/1608.07068 (visited on 05/01/2024)

  199. [213]

    Zhang, J

    X. Zhang, J. Zhao, and Y . LeCun,Character-level convolutional networks for text classifica- tion, 2016. arXiv: 1509.01626[cs]. [Online]. Available: http://arxiv.org/abs/ 1509.01626 (visited on 05/29/2024)

  200. [214]

    MARS: A video benchmark for large-scale person re- identification,

    L. Zheng, Z. Bie, Y . Sun, et al., “MARS: A video benchmark for large-scale person re- identification,” in Computer Vision – ECCV 2016 , B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., Cham: Springer International Publishing, 2016, pp. 868–884, ISBN : 978-3- 319-46466-4. DO...

  201. [215]

    H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline , 2017. arXiv: 1709 . 05522[cs]. [Online]. Available: http://arxiv.org/abs/1709.05522 (visited on 05/01/2024)

  202. [216]

    Chouldechova, Fair prediction with disparate impact: A study of bias in recidivism pre- diction instruments, 2017

    A. Chouldechova, Fair prediction with disparate impact: A study of bias in recidivism pre- diction instruments, 2017. arXiv: 1703.00056[cs,stat]. [Online]. Available: http: //arxiv.org/abs/1703.00056 (visited on 10/02/2024)

  203. [217]

    Frames: A corpus for adding memory to goal- oriented dialogue systems,

    L. El Asri, H. Schulz, S. Sharma, et al., “Frames: A corpus for adding memory to goal- oriented dialogue systems,” in Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, K. Jokinen, M. Stede, D. DeVault, and A. Louis, Eds., Saarbrücken, Germany: Associati...

  204. [218]

    Eric and C

    M. Eric and C. D. Manning, Key-value retrieval networks for task-oriented dialogue, 2017. arXiv: 1705.05414[cs] . [Online]. Available: http://arxiv.org/abs/1705. 05414 (visited on 05/02/2024)

  205. [219]

    D. F. Fouhey, W.-c. Kuo, A. A. Efros, and J. Malik, From lifestyle vlogs to everyday inter- actions, 2017. arXiv: 1712.02310[cs]. [Online]. Available: http://arxiv.org/ abs/1712.02310 (visited on 05/01/2024)

  206. [220]

    something something

    R. Goyal, S. E. Kahou, V . Michalski,et al., The "something something" video database for learning and evaluating visual common sense, 2017. arXiv: 1706.04261[cs]. [Online]. Available: http://arxiv.org/abs/1706.04261 (visited on 05/01/2024)

  207. [221]

    Ha and D

    D. Ha and D. Eck, A neural representation of sketch drawings , 2017. arXiv: 1704 . 03477[cs,stat] . [Online]. Available: http://arxiv.org/abs/1704.03477 (visited on 10/02/2024)

  208. [222]

    The THUMOS challenge on action recognition for videos

    H. Idrees, A. R. Zamir, Y .-G. Jiang, et al., “The THUMOS challenge on action recognition for videos "in the wild",” Computer Vision and Image Understanding, vol. 155, pp. 1–23, 2017, ISSN : 10773142. DOI: 10/f9rwnr. arXiv: 1604.06182[cs]. [Online]. Available: http://arxiv.org...

  209. [223]

    Ito and L

    K. Ito and L. Johnson, The LJ Speech Dataset , 2017. [Online]. Available: https : / / keithito.com/LJ-Speech-Dataset (visited on 05/01/2024)

  210. [224]

    S. Iyer, I. Konstas, A. Cheung, J. Krishnamurthy, and L. Zettlemoyer, Learning a neural semantic parser from user feedback, 2017. arXiv: 1704.08760[cs]. [Online]. Available: http://arxiv.org/abs/1704.08760 (visited on 10/02/2024). 43 The Data Provenance Initiative, 2024

  211. [225]

    Search-based neural structured learning for se- quential question answering,

    M. Iyyer, W.-t. Yih, and M. -W. Chang, “Search-based neural structured learning for se- quential question answering,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), R. Barzilay and M.-Y . Kan, Eds., Vancouver...

  212. [226]

    Joshi, E

    M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, TriviaQA: A large scale distantly su- pervised challenge dataset for reading comprehension, 2017. arXiv: 1705.03551[cs]. [Online]. Available: http://arxiv.org/abs/1705.03551 (visited on 05/29/2024)

  213. [227]

    W. Kay, J. Carreira, K. Simonyan,et al., The kinetics human action video dataset, 2017. arXiv: 1705.06950[cs]. [Online]. Available: http://arxiv.org/abs/1705.06950 (visited on 05/01/2024)

  214. [228]

    Korzinek, K

    D. Korzinek, K. Marasek, L. Brocki, and K. Wolk,Polish read speech corpus for speech tools and services, 2017. arXiv: 1706.00245[cs] . [Online]. Available: http://arxiv. org/abs/1706.00245 (visited on 05/29/2024)

  215. [229]

    G. Lai, Q. Xie, H. Liu, Y . Yang, and E. Hovy,RACE: Large-scale ReAding comprehension dataset from examinations, 2017. arXiv: 1704.04683[cs]. [Online]. Available: http: //arxiv.org/abs/1704.04683 (visited on 05/29/2024)

  216. [230]

    Lewis, D

    M. Lewis, D. Yarats, Y . N. Dauphin, D. Parikh, and D. Batra,Deal or no deal? end-to-end learning for negotiation dialogues, 2017. arXiv: 1706.05125[cs]. [Online]. Available: http://arxiv.org/abs/1706.05125 (visited on 05/29/2024)

  217. [231]

    Free linguistic and speech resources for tibetan,

    G. Li, H. Yu, T. F. Zheng, J. Yan, and S. Xu, “Free linguistic and speech resources for tibetan,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Kuala Lumpur: IEEE, 2017, pp. 733–736, ISBN : 978-1-5386- 1542-3. DOI...

  218. [232]

    W. Ling, D. Yogatama, C. Dyer, and P. Blunsom,Program induction by rationale generation : Learning to solve and explain algebraic word problems, 2017. arXiv: 1705.04146[cs]. [Online]. Available: http://arxiv.org/abs/1705.04146 (visited on 05/29/2024)

  219. [233]

    C. Liu, Y . Hu, Y . Li, S. Song, and J. Liu,PKU-MMD: A large scale benchmark for continuous multi-modal human action understanding , 2017. arXiv: 1703 . 07475[cs]. [Online]. Available: http://arxiv.org/abs/1703.07475 (visited on 05/01/2024)

  220. [234]

    Mrksic, D

    N. Mrksic, D. O. Seaghdha, T.-H. Wen, B. Thomson, and S. Young,Neural belief tracker: Data-driven dialogue state tracking, 2017. arXiv: 1606.03777[cs]. [Online]. Available: http://arxiv.org/abs/1606.03777 (visited on 05/01/2024)

  221. [235]

    Novikova, O

    J. Novikova, O. Dušek, and V . Rieser, The e2e dataset: New challenges for end-to-end generation, 2017. arXiv: 1706 . 09254[cs]. [Online]. Available: http : / / arxiv . org/abs/1706.09254 (visited on 05/29/2024)

  222. [236]

    A. See, P. J. Liu, and C. D. Manning, Get to the point: Summarization with pointer-generator networks, 2017. arXiv: 1704.04368[cs]. [Online]. Available: http://arxiv.org/ abs/1704.04368 (visited on 05/29/2024)

  223. [237]

    Sharghi, J

    A. Sharghi, J. S. Laurel, and B. Gong, Query-focused video summarization: Dataset, evalua- tion, and a memory network based approach, 2017. arXiv: 1707.04960[cs]. [Online]. Available: http://arxiv.org/abs/1707.04960 (visited on 05/01/2024)

  224. [238]

    A free kazakh speech database and a speech recognition baseline,

    Y . Shi, A. Hamdullah, Z. Tang, D. Wang, and T. F. Zheng, “A free kazakh speech database and a speech recognition baseline,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , Kuala Lumpur: IEEE, 2017, pp. 745–748, IS...

  225. [239]

    Welbl, N

    J. Welbl, N. F. Liu, and M. Gardner, Crowdsourcing multiple choice science questions, 2017. arXiv: 1707.06209[cs, stat]. [Online]. Available: http://arxiv.org/abs/ 1707.06209 (visited on 05/29/2024). 44 The Data Provenance Initiative, 2024

  226. [240]

    Yeung, O

    S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei, Every mo- ment counts: Dense detailed labeling of actions in complex videos , 2017. arXiv: 1507 . 05738[cs]. [Online]. Available: http://arxiv.org/abs/1507.05738 (visited on 05/02/2024)

  227. [241]

    Zhong, C

    V . Zhong, C. Xiong, and R. Socher,Seq2sql: Generating structured queries from natural lan- guage using reinforcement learning, 2017. arXiv: 1709.00103[cs]. [Online]. Available: http://arxiv.org/abs/1709.00103 (visited on 05/01/2024)

  228. [242]

    L. Zhou, C. Xu, and J. J. Corso, Towards automatic learning of procedures from web instructional videos, 2017. arXiv: 1703.09788[cs]. [Online]. Available: https:// arxiv.org/abs/1703.09788v3 (visited on 05/02/2024)

  229. [243]

    Bajaj, D

    P. Bajaj, D. Campos, N. Craswell, et al., MS MARCO: A human generated MAchine reading COmprehension dataset, 2018. arXiv: 1611 . 09268[cs]. [Online]. Available: http : //arxiv.org/abs/1611.09268 (visited on 05/29/2024)

  230. [244]

    J. A. Botha, M. Faruqui, J. Alex, J. Baldridge, and D. Das, Learning to split and rephrase from wikipedia edit history, 2018. arXiv: 1808.09468[cs]. [Online]. Available: http: //arxiv.org/abs/1808.09468 (visited on 10/02/2024)

  231. [245]

    Carreira, E

    J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman, A short note about kinetics-600, 2018. arXiv: 1808.01340[cs] . [Online]. Available: http://arxiv. org/abs/1808.01340 (visited on 05/01/2024)

  232. [246]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, et al., Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018. arXiv: 1803.05457[cs]. [Online]. Available: http://arxiv.org/abs/1803.05457 (visited on 05/29/2024)

  233. [247]

    Coucke, A

    A. Coucke, A. Saade, A. Ball, et al., Snips voice platform: An embedded spoken language un- derstanding system for private-by-design voice interfaces, 2018. arXiv: 1805.10190[cs]. [Online]. Available: http://arxiv.org/abs/1805.10190 (visited on 05/01/2024)

  234. [248]

    Damen, H

    D. Damen, H. Doughty, G. M. Farinella, et al. , Scaling egocentric vision: The EPIC- KITCHENS dataset, 2018. arXiv: 1804 . 02748[cs]. [Online]. Available: http : / / arxiv.org/abs/1804.02748 (visited on 05/01/2024)

  235. [249]

    J. Du, X. Na, X. Liu, and H. Bu, AISHELL-2: Transforming mandarin ASR research into industrial scale, 2018. arXiv: 1808.10583[cs]. [Online]. Available: http://arxiv. org/abs/1808.10583 (visited on 05/31/2024)

  236. [250]

    Faruqui and D

    M. Faruqui and D. Das, Identifying well-formed natural language questions, 2018. arXiv: 1808.09419[cs]. [Online]. Available: http://arxiv.org/abs/1808.09419 (visited on 05/29/2024)

  237. [251]

    Gorrell, K

    G. Gorrell, K. Bontcheva, L. Derczynski, E. Kochkina, M. Liakata, and A. Zubiaga, Ru- mourEval 2019: Determining rumour veracity and support for rumours, 2018. arXiv: 1809. 06683[cs]. [Online]. Available: http://arxiv.org/abs/1809.06683 (visited on 05/29/2024)

  238. [252]

    C. Gu, C. Sun, D. A. Ross, et al., AVA: A video dataset of spatio-temporally localized atomic visual actions, 2018. arXiv: 1705.08421[cs]. [Online]. Available: http://arxiv. org/abs/1705.08421 (visited on 05/29/2024)

  239. [253]

    MMQA: A multi-domain multi- lingual question-answering framework for english and hindi,

    D. Gupta, S. Kumari, A. Ekbal, and P. Bhattacharyya, “MMQA: A multi-domain multi- lingual question-answering framework for english and hindi,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), N. Calzolari, K. Choukri, C....

  240. [254]

    Gupta, R

    S. Gupta, R. Shah, M. Mohit, A. Kumar, and M. Lewis, Semantic parsing for task ori- ented dialog using hierarchical representations, 2018. arXiv: 1810.07942[cs]. [Online]. Available: http://arxiv.org/abs/1810.07942 (visited on 05/01/2024)

  241. [255]

    H. He, D. Chen, A. Balakrishnan, and P. Liang, Decoupling strategy and generation in negotiation dialogues, 2018. arXiv: 1808.09637[cs]. [Online]. Available: http:// arxiv.org/abs/1808.09637 (visited on 05/01/2024). 45 The Data Provenance Initiative, 2024

  242. [256]

    Localizing moments in video with temporal language,

    L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with temporal language,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, E...

  243. [257]

    TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

    F. Hernandez, V . Nguyen, S. Ghannay, N. Tomashenko, and Y . Estève, “TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Lecture Notes in Computer Science , vol. 11096, Springer, Cham, 2018, pp. 198–208. DOI: 10 . 1007/978- 3- 319-...

  244. [258]

    SciTaiL: A textual entailment dataset from science question answering,

    T. Khot, A. Sabharwal, and P. Clark, “SciTaiL: A textual entailment dataset from science question answering,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018, ISSN : 2374-3468, 2159-5399. DOI: 10/grm22d. [Online]. Available: https: / / ojs ....

  245. [259]

    S. Kim, I. Kang, and N. Kwak, Semantic sentence matching with densely-connected recurrent and co-attentive information, 2018. arXiv: 1805.11360[cs]. [Online]. Available: http: //arxiv.org/abs/1805.11360 (visited on 10/02/2024)

  246. [260]

    Crowd-sourced speech corpora for javanese, sundanese, sinhala, nepali, and bangladeshi bengali,

    O. Kjartansson, S. Sarin, K. Pipatsrisawat, M. Jansche, and L. Ha, “Crowd-sourced speech corpora for javanese, sundanese, sinhala, nepali, and bangladeshi bengali,” in6th Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU 2018), ISCA, 2018, pp. 52–55....

  247. [261]

    B. M. Lake and M. Baroni, Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks, 2018. arXiv: 1711.00350[cs]. [Online]. Available: http://arxiv.org/abs/1711.00350 (visited on 05/29/2024)

  248. [262]

    Event representations for automated story generation with deep neural nets,

    L. J. Martin, P. Ammanabrolu, X. Wang,et al., “Event representations for automated story generation with deep neural nets,” Proceedings of the AAAI Conference on Artificial In- telligence, vol. 32, no. 1, 2018, ISSN : 2374-3468, 2159-5399. DOI: 10/g6k72p . arXiv: 1706.01331[cs...

  249. [263]

    Mihaylov, P

    T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal,Can a suit of armor conduct electricity? a new dataset for open book question answering, 2018. arXiv: 1809.02789[cs]. [Online]. Available: http://arxiv.org/abs/1809.02789 (visited on 05/29/2024)

  250. [264]

    Moniz and L

    N. Moniz and L. Torgo, Multi-source social feedback of online news feeds , 2018. arXiv: 1801.07055[cs]. [Online]. Available: http://arxiv.org/abs/1801.07055 (visited on 10/02/2024)

  251. [265]

    Mrkši ´c and I

    N. Mrkši ´c and I. Vuli ´c, Fully statistical neural belief tracking , 2018. arXiv: 1805 . 11350[cs]. [Online]. Available: http : / / arxiv . org / abs / 1805 . 11350(vis- ited on 05/01/2024)

  252. [266]

    Nagrani, J

    A. Nagrani, J. S. Chung, and A. Zisserman, VoxCeleb: A large-scale speaker identifi- cation dataset , 2018. DOI: 10 . 21437 / Interspeech . 2017 - 950. arXiv: 1706 . 08612[cs]. [Online]. Available: http://arxiv.org/abs/1706.08612 (visited on 05/29/2024)

  253. [267]

    Narayan, S

    S. Narayan, S. B. Cohen, and M. Lapata, Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization, 2018. arXiv: 1808. 08745[cs]. [Online]. Available: http://arxiv.org/abs/1808.08745 (visited on 05/29/2024)

  254. [268]

    Royer, K

    A. Royer, K. Bousmalis, S. Gouws, et al., XGAN: Unsupervised image-to-image translation for many-to-many mappings, 2018. arXiv: 1711.05139[cs]. [Online]. Available: http: //arxiv.org/abs/1711.05139 (visited on 10/02/2024)

  255. [269]

    Saeidi, M

    M. Saeidi, M. Bartolo, P. Lewis, et al., Interpretation of natural language rules in conver- sational machine reading, 2018. arXiv: 1809.01494[cs,stat] . [Online]. Available: http://arxiv.org/abs/1809.01494 (visited on 10/02/2024). 46 The Data Provenance Initiative, 2024

  256. [270]

    A. Saha, R. Aralikatte, M. M. Khapra, and K. Sankaranarayanan, DuoRC: Towards complex language understanding with paraphrased reading comprehension , 2018. arXiv: 1804 . 07927[cs]. [Online]. Available: http://arxiv.org/abs/1804.07927 (visited on 05/29/2024)

  257. [271]

    Sanabria, O

    R. Sanabria, O. Caglayan, S. Palaskar, et al., How2: A large-scale dataset for multimodal language understanding, 2018. arXiv: 1811.00347[cs] . [Online]. Available: http: //arxiv.org/abs/1811.00347 (visited on 05/01/2024)

  258. [272]

    P. Shah, D. Hakkani-Tür, G. Tür, et al., Building a conversational agent overnight with dia- logue self-play, 2018. arXiv: 1801.04871[cs]. [Online]. Available: http://arxiv. org/abs/1801.04871 (visited on 05/01/2024)

  259. [273]

    Unsupervised abstractive meeting summarization with multi-sentence compression and budgeted submodular maximization,

    G. Shang, W. Ding, Z. Zhang, et al., “Unsupervised abstractive meeting summarization with multi-sentence compression and budgeted submodular maximization,” in Proceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. ...

  260. [274]

    G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, Actor and observer: Joint modeling of first and third-person videos, 2018. arXiv: 1804.09627[cs]. [Online]. Available: http://arxiv.org/abs/1804.09627 (visited on 05/01/2024)

  261. [275]

    Tafjord, P

    O. Tafjord, P. Clark, M. Gardner, W.-t. Yih, and A. Sabharwal, QuaRel: A dataset and models for answering questions about qualitative relationships, 2018. arXiv: 1811.08048[cs]. [Online]. Available: http://arxiv.org/abs/1811.08048 (visited on 10/02/2024)

  262. [276]

    The web as a knowledge-base for answering complex questions,

    A. Talmor and J. Berant, “The web as a knowledge-base for answering complex questions,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), M. Walker, H. Ji, ...

  263. [277]

    Thorne, A

    J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, FEVER: A large-scale dataset for fact extraction and VERification, 2018. arXiv: 1803.05355[cs]. [Online]. Available: http://arxiv.org/abs/1803.05355 (visited on 05/29/2024)

  264. [278]

    Vicol, M

    P. Vicol, M. Tapaswi, L. Castrejon, and S. Fidler, MovieGraphs: Towards understanding human-centric situations from videos, 2018. arXiv: 1712.06761[cs]. [Online]. Available: http://arxiv.org/abs/1712.06761 (visited on 05/29/2024)

  265. [279]

    AirDialogue: An environment for goal-oriented dialogue research,

    W. Wei, Q. Le, A. Dai, and J. Li, “AirDialogue: An environment for goal-oriented dialogue research,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds., Brussels, Belgium: As- soc...

  266. [280]

    Welbl, P

    J. Welbl, P. Stenetorp, and S. Riedel,Constructing datasets for multi-hop reading compre- hension across documents, 2018. arXiv: 1710.06481[cs]. [Online]. Available: http: //arxiv.org/abs/1710.06481 (visited on 10/02/2024)

  267. [281]

    Z. Wu, B. Ramsundar, E. N. Feinberg, et al., MoleculeNet: A benchmark for molecular machine learning, 2018. arXiv: 1703.00564[physics, stat]. [Online]. Available: http://arxiv.org/abs/1703.00564 (visited on 10/02/2024)

  268. [282]

    Z. Yang, P. Qi, S. Zhang, et al., HotpotQA: A dataset for diverse, explainable multi-hop question answering, 2018. arXiv: 1809 . 09600[cs]. [Online]. Available: http : / / arxiv.org/abs/1809.09600 (visited on 05/29/2024)

  269. [283]

    Amini, S

    A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi, MathQA: Towards interpretable math word problem solving with operation-based formalisms, 2019. arXiv: 1905.13319[cs] . [Online]. Available: http://arxiv.org/abs/1905. 13319 (visited on 05/29/2024)...

  270. [284]

    Balakrishnan, J

    A. Balakrishnan, J. Rao, K. Upasani, M. White, and R. Subba, Constrained decoding for neural NLG from compositional representations in task-oriented dialogue , 2019. arXiv: 1906.07220[cs]. [Online]. Available: http://arxiv.org/abs/1906.07220 (visited on 10/02/2024)

  271. [285]

    The spoken wikipedia corpus collection: Harvesting, alignment and an application to hyperlistening,

    T. Baumann, A. Köhn, and F. Hennig, “The spoken wikipedia corpus collection: Harvesting, alignment and an application to hyperlistening,”Language Resources and Evaluation, vol. 53, no. 2, pp. 303–329, 2019, ISSN : 1574-0218. DOI: 10/gq5xdf. [Online]. Available: https: //doi.or...

  272. [286]

    Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, PIQA: Reasoning about physical commonsense in natural language, 2019. arXiv: 1911.11641[cs]. [Online]. Available: http://arxiv.org/abs/1911.11641 (visited on 05/29/2024)

  273. [287]

    Byrne, K

    B. Byrne, K. Krishnamoorthi, C. Sankar, et al., Taskmaster-1: Toward a realistic and diverse dialog dataset, 2019. arXiv: 1909.05358[cs]. [Online]. Available: http://arxiv. org/abs/1909.05358 (visited on 05/01/2024)

  274. [288]

    Named entity dis- ambiguation using deep learning on graphs,

    A. Cetoli, M. Akbari, S. Bragaglia, A. D. O’Harney, and M. Sloan, “Named entity dis- ambiguation using deep learning on graphs,” in vol. 11438, 2019, pp. 78–86. DOI: 10 . 1007/978- 3- 030- 15719- 7_10 . arXiv: 1810.09164[cs] . [Online]. Available: http://arxiv.org/abs/1810.091...

  275. [290]

    [Online]

    arXiv: 1906.02059[cs] . [Online]. Available: http:// arxiv.org/abs/ 1906.02059 (visited on 10/02/2024)

  276. [294]

    Clark, K

    C. Clark, K. Lee, M. -W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, BoolQ: Exploring the surprising difficulty of natural yes/no questions , 2019. arXiv: 1905 . 10044[cs]. [Online]. Available: http : / / arxiv . org / abs / 1905 . 10044(vis- ited on 05/29/2024)

  277. [297]

    Dasigi, N

    P. Dasigi, N. F. Liu, A. Marasovi´c, N. A. Smith, and M. Gardner, Quoref: A reading com- prehension dataset with questions requiring coreferential reasoning, 2019. arXiv: 1908. 05803[cs]. [Online]. Available: http://arxiv.org/abs/1908.05803 (visited on 05/29/2024). 48 The Data...

  278. [298]

    MuST-c: A multilingual speech translation corpus,

    M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-c: A multilingual speech translation corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (...

  279. [299]

    Dinan, V

    E. Dinan, V . Logacheva, V . Malykh,et al., The second conversational intelligence challenge (ConvAI2), 2019. arXiv: 1902.00098[cs]. [Online]. Available:http://arxiv.org/ abs/1902.00098 (visited on 05/01/2024)

  280. [300]

    Dinan, S

    E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston, Wizard of wikipedia: Knowledge-powered conversational agents, 2019. arXiv: 1811 . 01241[cs]. [Online]. Available: http://arxiv.org/abs/1811.01241 (visited on 05/01/2024)

  281. [1405]

    [Online]

    DOI: 10/gcgk7w. [Online]. Available:https://doi.org/10.1007/s11263- 015-0816-y (visited on 05/29/2024)

  282. [2019]

    DOI: 10.1109/TETCI.2019.2892755

  283. [2020]

    [Online]

    arXiv: 2010.14701 [cs.LG] . [Online]. Available: https://arxiv.org/ abs/2010.14701

  284. [2023]

    arXiv: 2310.12941 [cs.LG]

  285. [2024]

    Available: https : / / www

    [Online]. Available: https : / / www . 404media . co / nvidia - sued - for - scraping-youtube-after-404-media-investigation/

  286. [4294]

    Available: https://proceedings.neurips.cc/paper_files/ paper/2020/file/2cd4e8a2ce081c3d7c32c3cde4312ef7-Paper.pdf

    [Online]. Available: https://proceedings.neurips.cc/paper_files/ paper/2020/file/2cd4e8a2ce081c3d7c32c3cde4312ef7-Paper.pdf

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.