REVIEW 3 major objections 6 minor 3 cited by
Bridging the Data Provenance Gap Across Text, Speech and Video
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A multimodal audit of nearly 4,000 datasets argues that dataset licenses drastically understate the legal constraints on AI training data, because over 80% of the content in widely used text, speech, and video datasets carries…
desk verdict First modality-spanning provenance audit; useful reference, but 'over 80%' overstates speech (78.6%) and the manual source coding needs validation before the headline numbers are taken as fact. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is a manual provenance-tracing protocol that codes each dataset on two axes at once. The dataset license is categorized as Commercial, Non-commercial/Academic, or Unspecified, while each underlying source — every website, platform, or model the content came from — is coded as Unrestricted, Unspecified, Source Closed, or Model Closed, following a defined taxonomy that covers terms of service, acceptable-use policies, and anti-crawling clauses. A dataset's overall terms status is set to the strictest of its sources, and the crosstab of license against terms is then reported both by dataset count and weighted by tokens or hours, which is what produces the headline mismatch figures.
What would settle it
Take a random sample of a few hundred datasets from the released audit, have a second independent team re-annotate source restrictions from the same lineage, and compute agreement; low agreement on the Source Closed and Model Closed categories would undercut the 80%-plus mismatch claim. A separate check: recompute the 99.8% text, 78% speech, and 99% video figures with the single largest collection per modality removed (for example the roughly 370k-hour YouTube-sourced speech corpus), since one re-coded giant could shift the ecosystem totals.
Extended reading notes
Core claim
The paper's central claim is a mismatch between two layers of restriction: dataset licenses and source terms. Counting datasets, only 25% of text, 33% of speech, and 32% of video datasets are licensed non-commercially; weighted by content, those figures are 21%, 26%, and 33%. But when the authors trace each dataset back to its original sources and classify the sources' licenses or terms of service, 99.8% of text tokens, 78% of speech hours, and 99% of video hours carry some non-commercial restriction at the source level, and the license-versus-terms mismatch covers 79% of text tokens, 55% of speech hours, and 65% of video hours. The paper further claims that since 2019 multimodal training has overwhelmingly moved to web-crawled, synthetic, and social-media sources (YouTube alone supplies roughly 71% of video data and 69% of speech data), and that although the absolute number of languages and countries represented keeps rising, the relative dispersion measured by Gini coefficients has not significantly improved since 2013 — Western concentration persists despite diversification at the margins.
Load-bearing premise
The central percentages depend on the manual classification of each dataset's sources as Unrestricted, Unspecified, Source Closed, or Model Closed, and the paper reports no inter-annotator agreement or independent validation of that coding, so a systematic bias in it would move every headline number.
Editorial extensions
If this is right
- Developers who filter training data by dataset license alone will systematically misjudge the restrictions on the content, since source-level terms bind more than 80% of content in each modality.
- Because the strictest-source rule intensifies at the collection level, large re-packaged text collections hide commercially usable subsets that practitioners cannot easily extract.
- The concentration of speech and video on a single video platform means platform terms, not dataset licenses, are the effective gatekeepers of most audio-visual training data.
- If the source coding is correct, the permissive commons for multimodal training reduces to the small residual of content that is both commercially licensed and sourced from unrestricted sources — under 1% of text and video content and about 5% of speech hours.
- Relative geographical and linguistic representation has been flat for a decade, so adding more languages and countries at the margins does not by itself reduce concentration.
Reading between the lines
- One implication the paper leaves implicit: if source terms bind, then a 'clean' multimodal training run under current terms would have to be assembled almost entirely from a short list of explicitly permissive sources, and the paper's released tables make that set enumerable.
- A testable extension: apply the same two-axis coding to datasets released after April 2024 to see whether the license-versus-source mismatch grows as synthetic outputs and short-video platforms enter the training mix.
- The mismatch framing points to a fork the paper deliberately does not resolve: either source terms are largely unenforceable against training (shifting risk to copyright law itself), or a large fraction of existing multimodal training is already in breach of terms — future litigation will pick the branch.
- Because the volume-weighted percentages are driven by a handful of giant collections, a sensitivity analysis that recomputes the headline figures with the largest collection per modality removed would show how robust the ecosystem-level claims actually are.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a large-scale manual audit of 3,916 public text, speech, and video datasets released between 1990 and 2024, tracing sourcing trends, license and source-term restrictions, and geographic and linguistic representation. The authors report three headline findings: (1) multimodal training data increasingly comes from web-crawled, social media, and synthetic sources; (2) although fewer than one-third of datasets carry restrictive licenses, over 80% of the underlying source content carries non-commercial restrictions; and (3) absolute language and country counts have risen since 2013 but relative measures of geographic and multilingual inequality have not significantly improved. The paper includes extensive appendix tables, a detailed taxonomy, and a promised public release of the audit data and code. The central quantitative claims are descriptive measurements rather than fitted-model outputs, but the main restriction percentages depend entirely on manually assigned source-level labels whose reliability is not reported.
Significance. If the headline measurement is correct, the paper provides an important and policy-relevant result: the effective legal constraints on widely used training data are much stronger than dataset licenses alone suggest. The work is the first multimodal provenance audit at this scale, and the appendix tables and attribution cards are a substantial community resource. The paper is appropriately cautious in describing licenses and terms as signals rather than enforceable legal determinations, and it avoids overfitting by presenting descriptive statistics rather than model-based claims. The empirical contribution would be strengthened by the promised release of annotations and code, which is an explicit strength if delivered. However, the central 'over 80%' claim is not yet fully supported: it is internally inconsistent with the paper's own Table 3 for speech, and the underlying manual classifications have no reported reliability validation.
major comments (3)
- [Section 2, 'Annotation Features & Methodology'; Tables 3 and 4] The central quantitative claims—99.8% of text tokens, 78.6% of speech hours, and 99.1% of video hours carrying source restrictions—are direct tabulations of the manually assigned labels in Table 2, yet the paper reports no inter-annotator agreement, no gold-standard validation, and no sensitivity analysis for ambiguous cases. A systematic bias in coding (for example, treating 'no information found' as Restricted rather than Unspecified) would directly change the headline percentages and the abstract's 'over 80%' conclusion. The manuscript states that 'All annotations and analysis code will be made publicly available on release,' but no link or commit hash is given, so independent verification is not currently possible. Please add a reliability study, such as dual annotation with agreement statistics on a representative subsample or an audit of ambiguous cases, and provide the data and code artifact at submission time.
- [Abstract and Section 3.2 (Table 3)] The abstract claims that 'over 80% of the source content in widely-used text, speech, and video datasets carry non-commercial restrictions,' and the introduction states 'over 80% of content from each modality.' Table 3 reports 78.6% for speech hours, and Section 3.2 itself says '78%' for speech. The claim is therefore internally inconsistent with the paper's own appendix data. The wording should be corrected to 'roughly 80%,' 'over 78%,' or the computation should be adjusted so that the abstract matches the reported tables.
- [Section 3.3, Figure 4] The claim that geographical and linguistic representation 'has not significantly improved' rests on Gini coefficients with 95% confidence intervals and statements about significance at the p = 0.05 level, but the method for computing these intervals and tests is not described anywhere in the paper or appendix, and no code is provided. Without knowing whether the intervals come from a bootstrap, jackknife, or analytic approach, the 'not significant' conclusion is not verifiable. Please specify the estimator and test procedure, or downgrade the claim to a descriptive statement about observed trends.
minor comments (6)
- [Section 2.1 vs Table 1] Section 2.1 reports 3,713 text datasets from 108 collections, while Table 1 reports 3,717 text datasets; these counts should be reconciled.
- [References] References [11] and [12] are the same paper by Buolamwini and Gebru, and the Common Voice citation appears twice, once as [13] and again as [20]; these duplicates should be merged.
- [Table 3 caption] The caption says the table is a breakdown 'across datasets' but the cells are percentages of total tokens or hours; clarify that the units are shares of content, not dataset counts, to avoid confusion with Table 4.
- [Figure 2 caption and Section 3.2] The phrase 'bare restrictions' appears in the Figure 2 caption and in the body text; this should be 'bear restrictions'.
- [Introduction, finding 2] The phrase 'undocumented restrictions in the dataset's sources' is imprecise: the source restrictions are often documented in terms of service, but not in the dataset license; consider wording such as 'source-level restrictions not reflected in dataset licenses.'
- [Section 3.1] The comparison of average synthetic and natural dataset lengths (1,756 vs 1,065 tokens) would benefit from sample sizes and a measure of dispersion to support the word 'notably.'
Circularity Check
No significant circularity: the reported percentages are new manual measurements, and the inherited taxonomy from the authors' prior audit is not load-bearing.
full rationale
The paper's central claims are measurements from manual annotations, not derivations from fitted parameters or self-referential definitions. The headline restriction percentages (e.g., 99.8% text, 78.6% speech, 99.1% video in Table 3) are computed by classifying each dataset source into Unrestricted, Unspecified, Source Closed, or Model Closed per Table 2, then aggregating by token or hour counts. Nothing in Table 2 is defined in terms of the headline result, and no parameter is fitted to a subset of the data and then 'predicted' on a closely related quantity; the annotation labels are the input and the source-restriction percentages are the output, which is the normal structure of a descriptive audit rather than a circular derivation. The adoption of the prior taxonomy and annotation process from the authors' own work, Longpre et al. [123], is a methodological inheritance and is explicitly acknowledged; it does not function as an unverified uniqueness theorem that forces the conclusion, and the current paper extends the taxonomy to source terms for speech and video. Other self-citations such as [122]-[125] provide context and framing, but the load-bearing evidence is the new dataset-level annotation, which is independently checkable once released. The most serious concerns are correctness and robustness issues rather than circularity: no inter-annotator agreement or gold-standard validation is reported for the manual source coding, the abstract's 'over 80%' claim sits uneasily with the 78.6% speech figure in Table 3, and the annotation code is promised rather than linked, as the paper states that 'all annotations and analysis code will be made publicly available on release.' These limitations affect reliability and verifiability, but they do not make the derivation circular.
Assumptions & free parameters
assumptions (3)
- domain assumption The curated dataset list (from HuggingFace, surveys, and expert review) is representative of widely used public datasets
- domain assumption Manual source-term annotations follow the Table 2 taxonomy consistently without measurable disagreement
- ad hoc to paper The Gini coefficient confidence intervals and p-value statements in Figure 4 are computed by an unspecified but valid method
Cite this review
Pith. "Pith review of Bridging the Data Provenance Gap Across Text, Speech and Video." pith.science (2026). https://pith.science/paper/IFMWDFZY
@misc{pith2026241217847,
author = {Pith},
title = {Pith review of: Bridging the Data Provenance Gap Across Text, Speech and Video},
year = {2026},
howpublished = {\url{https://pith.science/paper/IFMWDFZY}},
note = {Machine review of arXiv:2412.17847}
}
read the original abstract
Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities--popular text, speech, and video datasets--from their detailed sourcing trends and use restrictions to their geographical and linguistic representation. Our manual analysis covers nearly 4000 public datasets between 1990-2024, spanning 608 languages, 798 sources, 659 organizations, and 67 countries. We find that multimodal machine learning applications have overwhelmingly turned to web-crawled, synthetic, and social media platforms, such as YouTube, for their training sets, eclipsing all other sources since 2019. Secondly, tracing the chain of dataset derivations we find that while less than 33% of datasets are restrictively licensed, over 80% of the source content in widely-used text, speech, and video datasets, carry non-commercial restrictions. Finally, counter to the rising number of languages and geographies represented in public AI training datasets, our audit demonstrates measures of relative geographical and multilingual representation have failed to significantly improve their coverage since 2013. We believe the breadth of our audit enables us to empirically examine trends in data sourcing, restrictions, and Western-centricity at an ecosystem-level, and that visibility into these questions are essential to progress in responsible AI. As a contribution to ongoing improvements in dataset transparency and responsible use, we release our entire multimodal audit, allowing practitioners to trace data provenance across text, speech, and video.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 3 Pith papers
-
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
A new 8TB openly-licensed text corpus trains 7B LLMs that are competitive with Llama 1/2, showing that performant models need not depend on unlicensed web data.
-
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.
-
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
BRoverbs lets researchers test whether language models understand Portuguese proverbs; commercial models nearly master it, small models often guess randomly.
Reference graph
Works this paper leans on
-
[1]
Untitled review,
E. B. Wilson, “Untitled review,” The American Economic Review, vol. 4, no. 2, pp. 442– 444, 1914, ISSN : 00028282. [Online]. Available: http://www.jstor.org/stable/ 1804762 (visited on 09/26/2024)
1914
-
[2]
On the measurement of inequality,
A. B. Atkinson et al., “On the measurement of inequality,”Journal of economic theory, vol. 2, no. 3, pp. 244–263, 1970
1970
-
[3]
A survey of video datasets for human action and activity recognition,
J. M. Chaquet, E. J. Carmona, and A. Fernández-Caballero, “A survey of video datasets for human action and activity recognition,” Computer Vision and Image Understanding, vol. 117, no. 6, pp. 633–659, 2013, ISSN : 1077-3142. DOI: 10.1016/j.cviu.2013.01.013 . [Online]. Available: http://dx.doi.org/10.1016/j.cviu.2013.01.013
-
[8]
S. Shankar, Y . Halpern, E. Breck, J. Atwood, J. Wilson, and D. Sculley, “No classification without representation: Assessing geodiversity issues in open data sets for the developing world,” arXiv preprint arXiv:1711.08536, 2017
arXiv 2017
-
[9]
Playing hard exploration games by watching youtube,
Y . Aytar, T. Pfaff, D. Budden, T. Paine, Z. Wang, and N. de Freitas, “Playing hard exploration games by watching youtube,” in Advances in Neural Information Process- ing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31, Curran Associates, Inc., 2018. [Online]. Available: https : / / proceedings . n...
2018
-
[10]
E. M. Bender and B. Friedman, “Data statements for natural language processing: Toward mitigating system bias and enabling better science,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 587–604, 2018. DOI: 10.1162/tacl_a_00041. [Online]. Available: https://aclanthology.org/Q18-1041
-
[12]
Gender shades: Intersectional accuracy disparities in commer- cial gender classification,
J. Buolamwini and T. Gebru, “Gender shades: Intersectional accuracy disparities in commer- cial gender classification,” in Proceedings of the 1st Conference on Fairness, Accountability and Transparency, S. A. Friedler and C. Wilson, Eds., ser. Proceedings of Machine Learning Research, vol. 81, PMLR, 2018, pp. 77–91. [Online]. Available:https://proceedings...
2018
-
[13]
Common voice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, et al., “Common voice: A massively-multilingual speech corpus,” arXiv preprint arXiv:1912.06670, 2019
arXiv 1912
Show all 294 references
-
[14]
Does object recognition work for everyone?
T. De Vries, I. Misra, C. Wang, and L. Van der Maaten, “Does object recognition work for everyone?” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2019, pp. 52–59
2019
-
[15]
Visual to text: Survey of image and video captioning,
S. Li, Z. Tao, K. Li, and Y . Fu, “Visual to text: Survey of image and video captioning,”IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 3, no. 4, pp. 297–312,
-
[16]
Mundane content on social media: Creation, circulation, and the copyright problem,
J. Meese and J. Hagedorn, “Mundane content on social media: Creation, circulation, and the copyright problem,” Social Media+ Society, vol. 5, no. 2, p. 2 056 305 119 839 190, 2019. 11 The Data Provenance Initiative, 2024
2019
-
[17]
Model cards for model reporting,
M. Mitchell, S. Wu, A. Zaldivar, et al., “Model cards for model reporting,” in Proceedings of the conference on fairness, accountability, and transparency, 2019, pp. 220–229
2019
-
[18]
Moments in time dataset: One million videos for event understanding,
M. Monfort, A. Andonian, B. Zhou, et al., “Moments in time dataset: One million videos for event understanding,” IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 2, pp. 502–508, 2019
2019
-
[19]
Social data: Biases, methodological pitfalls, and ethical boundaries,
A. Olteanu, C. Castillo, F. Diaz, and E. Kıcıman, “Social data: Biases, methodological pitfalls, and ethical boundaries,” Frontiers in big data, vol. 2, p. 13, 2019
2019
-
[20]
Common voice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, et al., “Common voice: A massively-multilingual speech corpus,” English, in Proceedings of the Twelfth Language Resources and Evaluation Confer- ence, N. Calzolari, F. Béchet, P. Blache,et al., Eds., Marseille, France: European Language Resourc...
2020
-
[21]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: h...
2020
-
[22]
Semantic visual navigation by watching youtube videos,
M. Chang, A. Gupta, and S. Gupta, “Semantic visual navigation by watching youtube videos,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, Curran Associates, Inc., 2020, pp. 4283–
2020
-
[23]
The pile: An 800gb dataset of diverse text for language modeling,
L. Gao, S. Biderman, S. Black, et al., “The pile: An 800gb dataset of diverse text for language modeling,” arXiv preprint arXiv:2101.00027, 2020
2020 arXiv
-
[24]
Henighan, J
T. Henighan, J. Kaplan, M. Katz, et al., Scaling laws for autoregressive generative modeling,
-
[25]
The state and fate of linguistic diversity and inclusion in the nlp world,
P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, “The state and fate of linguistic diversity and inclusion in the nlp world,” arXiv preprint arXiv:2004.09095, 2020
2004 arXiv
-
[26]
Scaling laws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, et al., “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[27]
Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages,
A. Kunchukuttan, D. Kakwani, S. Golla, A. Bhattacharyya, M. M. Khapra, P. Kumar,et al., “Ai4bharat-indicnlp corpus: Monolingual corpora and word embeddings for indic languages,” arXiv preprint arXiv:2005.00085, 2020
2005 arXiv
-
[28]
Beyond “i agree
E. P. Robinson and Y . Zhu, “Beyond “i agree”: Users’ understanding of web site terms of service,” Social media+ society, vol. 6, no. 1, p. 2 056 305 119 897 321, 2020
2020
-
[29]
The new legal landscape for text mining and machine learning,
M. J. Sag, “The new legal landscape for text mining and machine learning,” in Journal of the Copyright Society of the USA, 2020
2020
-
[30]
Y . Zhu, X. Li, C. Liu,et al., A comprehensive study of deep video action recognition, 2020. arXiv: 2012.06567 [cs.CV] . [Online]. Available: https://arxiv.org/abs/ 2012.06567
2020 arXiv
-
[31]
Masakhaner: Named entity recognition for african languages,
D. I. Adelani, J. Abbott, G. Neubig, et al., “Masakhaner: Named entity recognition for african languages,” Transactions of the Association for Computational Linguistics, vol. 9, pp. 1116–1131, 2021
2021
-
[32]
How might we create better benchmarks for speech recognition?
A. Aksënova, D. van Esch, J. Flynn, and P. Golik, “How might we create better benchmarks for speech recognition?” In Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future, K. Church, M. Liberman, and V . Kordoni, Eds., Online: Association for Computational ...
2021 doi
-
[33]
Xls-r: Self-supervised cross-lingual speech representa- tion learning at scale,
A. Babu, C. Wang, A. Tjandra, et al., “Xls-r: Self-supervised cross-lingual speech representa- tion learning at scale,” arXiv preprint arXiv:2111.09296, 2021
2021 arXiv
-
[34]
Addressing “documentation debt
J. Bandy and N. Vincent, “Addressing “documentation debt” in machine learning research: A retrospective datasheet for bookcorpus,” arXiv preprint arXiv:2105.05241, 2021. 12 The Data Provenance Initiative, 2024
2021 arXiv
-
[35]
Multimodal datasets: Misogyny, pornography, and malignant stereotypes,
A. Birhane, V . U. Prabhu, and E. Kahembwe, “Multimodal datasets: Misogyny, pornography, and malignant stereotypes,” arXiv preprint arXiv:2110.01963, 2021
2021 arXiv
-
[36]
Quality at a glance: An audit of web-crawled multilingual datasets,
I. Caswell, J. Kreutzer, L. Wang, et al., “Quality at a glance: An audit of web-crawled multilingual datasets,” arXiv preprint arXiv:2103.12028, 2021
2021 arXiv
-
[37]
Documenting large webtext corpora: A case study on the colossal clean crawled corpus,
J. Dodge, M. Sap, A. Marasovi´c, et al., “Documenting large webtext corpora: A case study on the colossal clean crawled corpus,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 1286–1305
2021
-
[38]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov,et al., An image is worth 16x16 words: Transformers for image recognition at scale, 2021. arXiv: 2010.11929 [cs.CV]
2021 arXiv
-
[39]
Datasheets for datasets,
T. Gebru, J. Morgenstern, B. Vecchione,et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021
2021
-
[40]
What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus,
A. S. Luccioni and J. D. Viviano, “What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus,” 2021. arXiv: 2105.02732 [cs.CL]
2021 arXiv
-
[41]
Understanding gender and racial disparities in image recognition models,
R. Mahadev and A. Chakravarti, “Understanding gender and racial disparities in image recognition models,” arXiv preprint arXiv:2107.09211, 2021
2021 arXiv
-
[42]
Automatic speech recognition: A survey,
M. Malik, M. K. Malik, K. Mehmood, and I. Makhdoom, “Automatic speech recognition: A survey,” Multimedia Tools and Applications, vol. 80, pp. 9411–9457, 2021
2021
- [43]
-
[44]
Data and its (dis) contents: A survey of dataset development and use in machine learning research,
A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis) contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11, 2021
2021
-
[45]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, et al., “Learning transferable visual models from natural language supervision,” arXiv preprint arXiv:2103.00020, 2021
2021 arXiv
-
[46]
Changing the world by changing the data,
A. Rogers, “Changing the world by changing the data,” inProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online: Association for Computati...
2021 doi
-
[47]
“Everyone wants to do the model work, not the data work
N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo, ““Everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI,” in CHI, ser. CHI ’21, Yokohama, Japan: Association for Computing Machinery, 2021, ISBN : 9781450380966....
2021
-
[48]
Multitask prompted training enables zero-shot task generalization,
V . Sanh, A. Webson, C. Raffel,et al., “Multitask prompted training enables zero-shot task generalization,” ICLR 2022, 2021. [Online]. Available: https://arxiv.org/abs/ 2110.08207
2022 arXiv
-
[49]
Finetuned language models are zero-shot learners,
J. Wei, M. Bosma, V . Zhao,et al., “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations, 2021
2021
-
[50]
Challenges in detoxifying language models,
J. Welbl, A. Glaese, J. Uesato,et al., “Challenges in detoxifying language models,” inFindings of the Association for Computational Linguistics: EMNLP 2021, 2021, pp. 2447–2469
2021
-
[51]
Detoxifying language models risks marginalizing minority voices,
A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein, “Detoxifying language models risks marginalizing minority voices,” inProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2021
-
[52]
Masader: Metadata sourcing for arabic text and speech data resources,
Z. Alyafeai, M. Masoud, M. Ghaleb, and M. S. Al-shaibani, “Masader: Metadata sourcing for arabic text and speech data resources,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2022, pp. 6340–6351
2022
-
[53]
Quantifying memo- rization across neural language models,
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang, “Quantifying memo- rization across neural language models,” 2022. arXiv: 2202.07646 [cs.LG]. 13 The Data Provenance Initiative, 2024
2022 arXiv
-
[54]
Elizalde, S
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, Clap: Learning audio concepts from natural language supervision, 2022. arXiv: 2206.04769 [cs.SD]
2022 arXiv
-
[55]
Dataset geography: Mapping language data to language users,
F. Faisal, Y . Wang, and A. Anastasopoulos, “Dataset geography: Mapping language data to language users,” in Proceedings of the 60th Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), 2022, pp. 3381–3411
2022
-
[56]
The flores-101 evaluation benchmark for low- resource and multilingual machine translation,
N. Goyal, C. Gao, V . Chaudhary, et al., “The flores-101 evaluation benchmark for low- resource and multilingual machine translation,” Transactions of the Association for Computa- tional Linguistics, vol. 10, pp. 522–538, 2022
2022
-
[57]
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, et al., “Training compute-optimal large language models,” arXiv preprint arXiv:2203.15556, 2022
2022 arXiv
-
[58]
Leakage and the reproducibility crisis in ml-based science,
S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in ml-based science,” arXiv preprint arXiv:2207.07048, 2022
2022 arXiv
-
[59]
Quality at a glance: An audit of web-crawled multilingual datasets,
J. Kreutzer, I. Caswell, L. Wang, et al., “Quality at a glance: An audit of web-crawled multilingual datasets,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 50–72, 2022
2022
-
[60]
The bigscience roots corpus: A 1.6tb composite multilingual dataset,
H. Laurençon, L. Saulnier, T. Wang,et al., “The bigscience roots corpus: A 1.6tb composite multilingual dataset,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35, Curran Associates, Inc., 20...
2022
-
[61]
McMillan-Major, Z
A. McMillan-Major, Z. Alyafeai, S. Biderman, et al., Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,
-
[62]
Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,
A. McMillan-Major, Z. Alyafeai, S. Biderman, et al., “Documenting geographically and contextually diverse data sources: The bigscience catalogue of language data and resources,” arXiv preprint arXiv:2201.10066, 2022
2022 arXiv
-
[63]
Moctezuma, T
D. Moctezuma, T. Ramírez-delReal, G. Ruiz, and O. González-Chávez, Video captioning: A comparative review of where we are and which could be the route , 2022. arXiv: 2204. 05976 [cs.CV]. [Online]. Available: https://arxiv.org/abs/2204.05976
2022 arXiv
-
[64]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, et al., “Training language models to follow instructions with human feedback,” arXiv preprint arXiv:2203.02155, 2022. [Online]. Available: https: //arxiv.org/abs/2203.02155
2022 arXiv
- [65]
-
[66]
Rando, D
J. Rando, D. Paleka, D. Lindner, L. Heim, and F. Tramèr,Red-teaming the stable diffusion safety filter, 2022. arXiv:2210.04610 [cs.AI]. [Online]. Available:https://arxiv. org/abs/2210.04610
2022 arXiv
-
[67]
Make-A-Video: Text-to-Video Generation without Text-Video Data,
U. Singer, A. Polyak, T. Hayes, et al., “Make-A-Video: Text-to-Video Generation without Text-Video Data,” arXiv: arXiv:2209.14792, 2022. arXiv:2209.14792 [cs]. [Online]. Available: http://arxiv.org/abs/2209.14792
2022 arXiv
-
[68]
Bigssl: Exploring the frontier of large-scale semi- supervised learning for automatic speech recognition,
Y . Zhang, D. S. Park, W. Han, et al., “Bigssl: Exploring the frontier of large-scale semi- supervised learning for automatic speech recognition,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1519–1532, 2022, ISSN : 1941-0484. DOI: 10.1109/ jstsp.2...
2022
-
[69]
Survey of video object detection algorithms based on deep learning,
L. Zheng, T. Zhou, R. Jiang, and Y . Peng, “Survey of video object detection algorithms based on deep learning,” in Proceedings of the 2021 4th International Conference on Algorithms, Computing and Artificial Intelligence, ser. ACAI ’21, Sanya, China: Association for Com- puti...
2021
-
[70]
Scaling laws for generative mixed-modal language models,
A. Aghajanyan, L. Yu, A. Conneau,et al., “Scaling laws for generative mixed-modal language models,” in International Conference on Machine Learning, PMLR, 2023, pp. 265–279
2023
-
[71]
Into the laions den: Investigating hate in multimodal datasets,
A. Birhane, V . Prabhu, S. Han, V . N. Boddeti, and A. S. Luccioni, “Into the laions den: Investigating hate in multimodal datasets,”arXiv preprint arXiv:2311.03449, 2023
2023 arXiv
-
[72]
Blattmann, T
A. Blattmann, T. Dockhorn, S. Kulal, et al., Stable video diffusion: Scaling latent video diffu- sion models to large datasets, 2023. arXiv: 2311.15127 [cs.CV]. [Online]. Available: https://arxiv.org/abs/2311.15127
2023 arXiv
-
[73]
Bommasani, K
R. Bommasani, K. Klyman, S. Longpre, et al., The foundation model transparency index,
-
[74]
Quantifying mem- orization across neural language models,
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramèr, and C. Zhang, “Quantifying mem- orization across neural language models,” in The Eleventh International Conference on Learning Representations, OpenReview, 2023
2023
-
[75]
Extracting training data from diffusion models,
N. Carlini, J. Hayes, M. Nasr, et al., “Extracting training data from diffusion models,” in32nd USENIX Security Symposium (USENIX Security 23), Anaheim, CA: USENIX Association, 2023, pp. 5253–5270, ISBN : 978-1-939133-37-3. [Online]. Available: https : / / www . usenix.org/con...
2023
-
[76]
S. H. Cen, A. Hopkins, A. Ilyas, A. Madry, I. Struckman, and L. Videgaray Caso, AI Supply Chains, 2023. [Online]. Available: http://dx.doi.org/10.2139/ssrn.4789403
2023 doi
-
[77]
Gender bias in hiring: An analysis of the impact of amazon’s recruiting algorithm,
X. Chang, “Gender bias in hiring: An analysis of the impact of amazon’s recruiting algorithm,” Advances in Economics, Management and Political Sciences, vol. 23, pp. 134–140, 2023. DOI: 10.54254/2754-1169/23/20230367
2023 doi
-
[78]
Can language models be instructed to protect personal information?
Y . Chen, E. Mendes, S. Das, W. Xu, and A. Ritter, “Can language models be instructed to protect personal information?” en, 2023
2023
-
[79]
Dialect corpora from youtube,
S. Coats, “Dialect corpora from youtube,” Language and linguistics in a complex world , 2023
2023
-
[80]
Ai image training dataset found to include child sexual abuse imagery,
E. David, “Ai image training dataset found to include child sexual abuse imagery,”The Verge, 2023, 7:57 AM PST. [Online]. Available: https://www.theverge.com/2023/12/ 20/24009418/generative-ai-image-laion-csam-google-stability- stanford
2023
-
[81]
What’s in my big data?
Y . Elazar, A. Bhagia, I. H. Magnusson, et al., “What’s in my big data?” In The Twelfth International Conference on Learning Representations, 2023
2023
-
[82]
Esser, J
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis,Structure and content- guided video synthesis with diffusion models, 2023. arXiv: 2302.03011 [cs.CV]. [On- line]. Available: https://arxiv.org/abs/2302.03011
2023 arXiv
-
[83]
Datacomp: In search of the next generation of multi- modal datasets,
S. Y . Gadre, G. Ilharco, A. Fang,et al., “Datacomp: In search of the next generation of multi- modal datasets,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36, Curran Associates, Inc., 2...
2023
-
[84]
Foundation models and fair use,
P. Henderson, X. Li, D. Jurafsky, T. Hashimoto, M. A. Lemley, and P. Liang, “Foundation models and fair use,” arXiv preprint arXiv:2303.15715, 2023
2023 arXiv
-
[85]
Understanding catastrophic forgetting in language models via implicit inference,
S. Kotha, J. M. Springer, and A. Raghunathan, “Understanding catastrophic forgetting in language models via implicit inference,” arXiv preprint arXiv:2309.10105, 2023
2023 arXiv
-
[86]
Harnessing large- language models to generate private synthetic text,
A. Kurakin, N. Ponomareva, U. Syed, L. MacDermed, and A. Terzis, “Harnessing large- language models to generate private synthetic text,” 2023. arXiv:2306.01684 [cs.LG]
2023 arXiv
-
[87]
Platypus: Quick, cheap, and powerful refinement of llms,
A. N. Lee, C. J. Hunter, and N. Ruiz, “Platypus: Quick, cheap, and powerful refinement of llms,” NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023
2023
-
[88]
Talkin”bout ai generation: Copyright and the generative-ai supply chain,
K. Lee, A. F. Cooper, and J. Grimmelmann, “Talkin”bout ai generation: Copyright and the generative-ai supply chain,” arXiv preprint arXiv:2309.08133, 2023
2023 arXiv
-
[89]
Yodas: Youtube- oriented dataset for audio and speech,
X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watanabe, “Yodas: Youtube- oriented dataset for audio and speech,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, 2023, pp. 1–8. 15 The Data Provenance Initiative, 2024
2023
-
[90]
H. Liu, C. Li, Y . Li, and Y . J. Lee, Improved baselines with visual instruction tuning, 2023
2023
-
[91]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, 2023
2023
-
[92]
Longpre, G
S. Longpre, G. Yauney, E. Reif, et al., A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023. arXiv: 2305.13169 [cs.CL]
2023 arXiv
-
[93]
Discit ergo est: Training data provenance and fair use,
R. Mahari and S. Longpre, “Discit ergo est: Training data provenance and fair use,”Robert Mahari and Shayne Longpre, Discit ergo est: Training Data Provenance And Fair Use, Dynamics of Generative AI (ed. Thibault Schrepel & Volker Stocker), Network Law Review, Winter, 2023
2023
-
[94]
Mahari, L
R. Mahari, L. Shayne, L. Donewald, A. Polozov, A. ’. Pentland, and A. Lipsitz, Comment to US copyright office on data provenance and copyright, 2023
2023
-
[95]
Marion, A
M. Marion, A. Üstün, L. Pozzobon, A. Wang, M. Fadaee, and S. Hooker, When less is more: Investigating data pruning for pretraining llms at scale, 2023. arXiv: 2309.04564 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2309.04564
2023 arXiv
-
[96]
Silo language models: Isolating legal risk in a nonparametric datastore,
S. Min, S. Gururangan, E. Wallace, et al., “Silo language models: Isolating legal risk in a nonparametric datastore,” in NeurIPS 2023 Workshop on Distribution Shifts: New Frontiers with Foundation Models, 2023
2023
-
[97]
Licensed to learn: Mitigating copyright infringement liability of generative ai systems through contracts,
F. Morton-Park, “Licensed to learn: Mitigating copyright infringement liability of generative ai systems through contracts,” Notre Dame Journal on Emerging Technology, vol. 5, p. 64, 2023
2023
-
[98]
Crosslingual generalization through multitask finetuning,
N. Muennighoff, T. Wang, L. Sutawika,et al., “Crosslingual generalization through multitask finetuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15 991–16 111
2023
-
[99]
The RefinedWeb dataset for falcon LLM: Outperforming curated corpora with web data, and web data only,
G. Penedo, Q. Malartic, D. Hesslow, et al., “The RefinedWeb dataset for falcon LLM: Outperforming curated corpora with web data, and web data only,” 2023. arXiv: 2306. 01116 [cs.CL]
2023
-
[100]
Reproducing whisper-style training using an open-source toolkit and publicly available data,
Y . Peng, J. Tian, B. Yan,et al., “Reproducing whisper-style training using an open-source toolkit and publicly available data,” in 2023 IEEE Automatic Speech Recognition and Under- standing Workshop (ASRU), IEEE, 2023, pp. 1–8
2023
-
[101]
Porgali, V
B. Porgali, V . Albiero, J. Ryda, C. C. Ferrer, and C. Hazirbas,The casual conversations v2 dataset, 2023. arXiv: 2303.04838 [cs.CV] . [Online]. Available: https://arxiv. org/abs/2303.04838
2023 arXiv
-
[102]
Pozzobon, B
L. Pozzobon, B. Ermis, P. Lewis, and S. Hooker, Goodtriever: Adaptive toxicity mitiga- tion with retrieval-augmented models , 2023. arXiv: 2310 . 07589 [cs.AI]. [Online]. Available: https://arxiv.org/abs/2310.07589
2023 arXiv
-
[103]
End-to-end speech recognition: A survey,
R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schlüter, and S. Watanabe, “End-to-end speech recognition: A survey,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023
2023
-
[104]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning, PMLR, 2023, pp. 28 492–28 518
2023
-
[105]
Direct pref- erence optimization: Your language model is secretly a reward model,
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn, “Direct pref- erence optimization: Your language model is secretly a reward model,” arXiv preprint arXiv:2305.18290, 2023
2023 arXiv
-
[106]
Self-supervised learning for videos: A survey,
M. C. Schiappa, Y . S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,” ACM Computing Surveys , vol. 55, no. 13s, pp. 1–37, 2023, ISSN : 1557-7341. DOI: 10 . 1145/3577925. [Online]. Available: http://dx.doi.org/10.1145/3577925
2023 doi
-
[107]
Detecting personal information in training corpora: An analysis,
N. Subramani, S. Luccioni, J. Dodge, and M. Mitchell, “Detecting personal information in training corpora: An analysis,” in Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023), Toronto, Canada: Association for Computational Linguistics, 2023
2023
-
[108]
Gemini: A family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, et al., “Gemini: A family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023. 16 The Data Provenance Initiative, 2024
2023 arXiv
-
[109]
Youtube-asl: A large-scale, open-domain ameri- can sign language-english parallel corpus,
D. Uthus, G. Tanzer, and M. Georg, “Youtube-asl: A large-scale, open-domain ameri- can sign language-english parallel corpus,” in Advances in Neural Information Process- ing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36, Curran As...
2023
-
[110]
Tag your fish in the broken net: A responsible web framework for protecting online privacy and copyright,
D. Zhang, B. Xia, Y . Liu,et al., “Tag your fish in the broken net: A responsible web framework for protecting online privacy and copyright,” 2023. arXiv:2310.07915 [cs.NI]
2023 arXiv
-
[111]
C. Zhu, Q. Jia, W. Chen, Y . Guo, and Y . Liu,Deep learning for video-text retrieval: A review,
-
[112]
Ahmadian, B
Aakanksha, A. Ahmadian, B. Ermis,et al., The multilingual alignment prism: Aligning global and local preferences to reduce harm , 2024. arXiv: 2406.18682 [cs.CL] . [Online]. Available: https://arxiv.org/abs/2406.18682
2024 arXiv
-
[113]
D. I. Adelani, J. Ojo, I. A. Azime, et al., Irokobench: A new benchmark for african languages in the age of large language models , 2024. arXiv: 2406 . 03368 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2406.03368
2024 arXiv
-
[114]
A survey on data selection for language models,
A. Albalak, Y . Elazar, S. M. Xie,et al., “A survey on data selection for language models,” arXiv preprint arXiv:2402.16827, 2024
2024 arXiv
-
[115]
Video generation models as world simulators,
T. Brooks, B. Peebles, C. Holmes,et al., “Video generation models as world simulators,” 2024. [Online]. Available: https : / / openai . com / research / video - generation - models-as-world-simulators
2024
-
[116]
Nvidia sued for scraping youtube after 404 media investigation,
S. Cole, “Nvidia sued for scraping youtube after 404 media investigation,” 404 Media,
- [117]
-
[118]
Datacomp: In search of the next generation of multimodal datasets,
S. Y . Gadre, G. Ilharco, A. Fang, et al., “Datacomp: In search of the next generation of multimodal datasets,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[119]
Klyman, Acceptable use policies for foundation models , 2024
K. Klyman, Acceptable use policies for foundation models , 2024. arXiv: 2409.09041 [cs.CY]. [Online]. Available: https://arxiv.org/abs/2409.09041
2024 arXiv
-
[120]
Best practices and lessons learned on synthetic data,
R. Liu, J. Wei, F. Liu, et al., “Best practices and lessons learned on synthetic data,” 2024. arXiv: 2404.07503 [cs.CL]
2024 arXiv
-
[121]
Datasets for large language models: A compre- hensive survey,
Y . Liu, J. Cao, C. Liu, K. Ding, and L. Jin, “Datasets for large language models: A compre- hensive survey,”arXiv preprint arXiv:2402.18041, 2024
2024 arXiv
-
[122]
The responsible foundation model development cheatsheet: A review of tools & resources,
S. Longpre, S. Biderman, A. Albalak, et al., “The responsible foundation model development cheatsheet: A review of tools & resources,” arXiv preprint arXiv:2406.16746, 2024
2024 arXiv
-
[123]
A large-scale audit of dataset licensing and attribution in AI,
S. Longpre, R. Mahari, A. Chen,et al., “A large-scale audit of dataset licensing and attribution in AI,” Nature Machine Intelligence, vol. 6, no. 8, pp. 975–987, 2024. DOI: 10/gt8f5p. arXiv: 2310.16787 [cs]
2024 arXiv
-
[124]
Nvlm: Open frontier-class multimodal llms,
W. Dai, N. Lee, B. Wang,et al., “Nvlm: Open frontier-class multimodal llms,”arXiv preprint, 2024
2024
-
[125]
Data authenticity, consent, & provenance for ai are all broken: What will it take to fix them?
S. Longpre, R. Mahari, N. Obeng-Marnu, et al., “Data authenticity, consent, & provenance for ai are all broken: What will it take to fix them?” arXiv preprint arXiv:2404.12691, 2024
2024 arXiv
-
[126]
Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages,
H. Lovenia, R. Mahendra, S. M. Akbar, et al., “Seacrowd: A multilingual multimodal data hub and benchmark suite for southeast asian languages,” arXiv preprint arXiv:2406.10118, 2024
2024 arXiv
-
[127]
Mauran, What was Sora trained on? Creatives demand answers
C. Mauran, What was Sora trained on? Creatives demand answers. https://mashable. com / article / openai - sora - ai - video - generator - training - data, [Accessed 28-09-2024], 2024. 17 The Data Provenance Initiative, 2024
2024
-
[128]
Topics, au- thors, and institutions in large language model research: Trends from 17k arxiv papers,
R. Movva, S. Balachandar, K. Peng, G. Agostini, N. Garg, and E. Pierson, “Topics, au- thors, and institutions in large language model research: Trends from 17k arxiv papers,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computation...
2024
-
[129]
[Online]
OpenAI, Hello gpt-4o: We’re announcing gpt-4o, our new flagship model that can reason across audio, vision, and text in real time.2024. [Online]. Available: https://openai. com/index/hello-gpt-4o/
2024
-
[130]
Data, data everywhere: A guide for pretraining dataset construction,
J. Parmar, S. Prabhumoye, J. Jennings, et al., “Data, data everywhere: A guide for pretraining dataset construction,” arXiv preprint 2407.06380, 2024
2024 arXiv
-
[131]
Consent in crisis: The rapid decline of the ai data commons,
S. Longpre, R. Mahari, A. Lee, et al., “Consent in crisis: The rapid decline of the ai data commons,” arXiv preprint arXiv:2407.14933, 2024
2024 arXiv
-
[132]
Anatomy of industrial scale multilingual asr,
F. M. Ramirez, L. Chkhetiani, A. Ehrenberg, et al., “Anatomy of industrial scale multilingual asr,” arXiv preprint arXiv:2404.09841, 2024
2024 arXiv
-
[133]
Romanou, N
A. Romanou, N. Foroutan, A. Sotnikova, et al., Include: Evaluating multilingual language understanding with regional knowledge, 2024. arXiv: 2411.19799 [cs.CL]. [Online]. Available: https://arxiv.org/abs/2411.19799
2024 arXiv
-
[134]
Singh, F
S. Singh, F. Vargus, D. Dsouza,et al., Aya dataset: An open-access collection for multilingual instruction tuning, 2024. arXiv: 2402.06619 [cs.CL]
2024 arXiv
-
[135]
Openai sued over using youtube videos without creators’ consent,
S. Skolnik, “Openai sued over using youtube videos without creators’ consent,”Bloomberg Law, 2024. [Online]. Available: https : / / news . bloomberglaw . com / litigation / openai - sued - over - using - youtube - videos - without - creators-consent
2024
-
[136]
Dolma: An open corpus of three trillion tokens for language model pretraining research,
L. Soldaini, R. Kinney, A. Bhagia, et al., “Dolma: An open corpus of three trillion tokens for language model pretraining research,” arXiv preprint arXiv:2402.00159, 2024
2024 arXiv
-
[137]
Aya model: An instruction finetuned open-access multilingual language model,
A. Üstün, V . Aryabumi, Z.-X. Yong, et al., “Aya model: An instruction finetuned open-access multilingual language model,” arXiv preprint arXiv:2402.07827, 2024
2024 arXiv
-
[138]
Scaling speech technology to 1,000+ languages,
V . Pratap, A. Tjandra, B. Shi,et al., “Scaling speech technology to 1,000+ languages,”Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024
2024
-
[139]
X. Yang, W. Liang, and J. Zou, Navigating dataset documentations in ai: A large-scale analysis of dataset cards on hugging face, 2024. arXiv: 2401.13822 [cs.LG]. [Online]. Available: https://arxiv.org/abs/2401.13822
2024 arXiv
-
[140]
Commercial
Z. Zheng, X. Peng, T. Yang, et al., Open-sora: Democratizing efficient video production for all, 2024. [Online]. Available: https://github.com/hpcaitech/Open-Sora. 18 The Data Provenance Initiative, 2024 LABEL DEFINITION MODEL CLOSED A model used to generate part or all of the...
2024
-
[141]
Chalkidis, I
I. Chalkidis, I. Androutsopoulos, and N. Aletras, Neural legal judgment prediction in english,
-
[142]
Chalkidis, M
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, and I. Androutsopoulos,Large-scale multi-label text classification on EU legislation, 2019. arXiv: 1906.02192[cs]. [Online]. Available: http://arxiv.org/abs/1906.02192 (visited on 05/29/2024)
2019 arXiv
-
[143]
W. Chen, J. Chen, P. Qin, X. Yan, and W. Y . Wang,Semantically conditioned dialog response generation via hierarchical disentangled self-attention, 2019. arXiv: 1905.12866[cs]. [Online]. Available: http://arxiv.org/abs/1905.12866 (visited on 05/01/2024)
2019 arXiv
-
[144]
Look before you hop: Conversational question answering over knowledge graphs using judicious context expansion,
P. Christmann, R. S. Roy, A. Abujabal, J. Singh, and G. Weikum, “Look before you hop: Conversational question answering over knowledge graphs using judicious context expansion,” in Proceedings of the 28th ACM International Conference on Information and Knowledge Management, 20...
2019 arXiv
-
[145]
Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models,
W. Wang and Y . Yang, “Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models,” arXiv preprint arXiv:2403.06098, 2024
2024 arXiv
-
[146]
A corpus of regional american language from YouTube,
S. Coats, “A corpus of regional american language from YouTube,” in Proceed- ings of the Digital Humanities in the Nordic Countries 5th Conference , 2019. [On- line]. Available: https : / / www . semanticscholar . org / paper / A - Corpus - of - Regional - American - Language ...
2019
-
[147]
Toyota smarthome: Real-world activities of daily living,
S. Das, R. Dai, M. Koperski, et al., “Toyota smarthome: Real-world activities of daily living,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , Seoul, Korea (South): IEEE, 2019, pp. 833–842, ISBN : 978-1-72814-803-8. DOI: 10/ghfjc7. [Online]. Available: h...
2019
-
[148]
The ATIS spoken language systems pilot corpus,
C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, “The ATIS spoken language systems pilot corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Hid- den Valley, Pennsylvania, June 24-27,1990, 1990. DOI: 10/cz3442. [Online]. Available: https://aclantholo...
1990
-
[149]
SWITCHBOARD: Telephone speech corpus for research and development,
J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in [Proceedings] ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing, San Francisco, CA, USA: IEEE, 1992, 517–520 vol.1, ISBN : ...
1992
-
[150]
Garofolo, John S., Lamel, Lori F., Fisher, William M., et al., TIMIT acoustic-phonetic continuous speech corpus, Artwork Size: 715776 KB Pages: 715776 KB, 1993. DOI: 10. 35111/17GK- BN40. [Online]. Available: https://catalog.ldc.upenn.edu/ LDC93S1 (visited on 05/01/2024)
1993
-
[151]
Can prosody aid the automatic classification of dialog acts in conversational speech?
E. Shriberg, R. Bates, A. Stolcke, et al., “Can prosody aid the automatic classification of dialog acts in conversational speech?” Language and Speech, vol. 41 ( Pt 3-4), pp. 443–492, 1998, ISSN : 0023-8309. DOI: 10.1177/002383099804100410
1998 doi
-
[152]
Dialogue act modeling for automatic tagging and recognition of conversational speech,
A. Stolcke, K. Ries, N. Coccaro, et al., “Dialogue act modeling for automatic tagging and recognition of conversational speech,”Computational Linguistics, vol. 26, no. 3, pp. 339–373, 2000, ISSN : 0891-2017, 1530-9312. DOI: 10/dqmv4j . arXiv: cs/0006023 . [Online]. Available: ...
2000 arXiv
-
[153]
Corpus of spontaneous japanese: Its design and evaluation,
K. Maekawa, “Corpus of spontaneous japanese: Its design and evaluation,” in Proceedings of the ISCA/IEEE Workshop on Spontaneous Speech Processing and Recognition, 2003, paper MMO2. [Online]. Available: https://www.isca- archive.org/sspr_2003/ maekawa03_sspr.html (visited on 0...
2003
-
[154]
The fisher corpus: A resource for the next generations of speech-to-text,
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text,” in Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04), M. T. Lino, M. F. Xavier, F. Ferreira, R. Costa, and R. Silva, ...
2004
-
[155]
Lander, T, CSLU: 22 languages corpus , 2005. DOI: 10 . 35111 / ZKN2 - 5X88. [On- line]. Available: https://catalog.ldc.upenn.edu/LDC2005S26 (visited on 05/29/2024)
2005
-
[156]
Pang and L
B. Pang and L. Lee, Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales , 2005. arXiv: cs / 0506075. [Online]. Available: http : //arxiv.org/abs/cs/0506075 (visited on 05/29/2024). 38 The Data Provenance Initiative, 2024
2005 arXiv
-
[157]
The AMI meeting corpus: A pre-announcement,
J. Carletta, S. Ashby, S. Bourban, et al., “The AMI meeting corpus: A pre-announcement,” in Machine Learning for Multimodal Interaction, S. Renals and S. Bengio, Eds., vol. 3869, Series Title: Lecture Notes in Computer Science, Berlin, Heidelberg: Springer Berlin Heidelberg, 2...
2006
-
[158]
Lander, T, CSLU: Foreign accented english release 1.2, Artwork Size: 1468006 KB Pages: 1468006 KB, 2007. DOI: 10 . 35111 / 0VWP - XN48. [Online]. Available: https : / / catalog.ldc.upenn.edu/LDC2007S08 (visited on 05/01/2024)
2007
-
[159]
Measures of semantic sim- ilarity and relatedness in the biomedical domain,
T. Pedersen, S. V . S. Pakhomov, S. Patwardhan, and C. G. Chute, “Measures of semantic sim- ilarity and relatedness in the biomedical domain,” Journal of Biomedical Informatics, vol. 40, no. 3, pp. 288–299, 2007, ISSN : 1532-0464. DOI: 10/fghjwr. [Online]. Available: https: //...
2007
-
[160]
Actions in context,
M. Marszalek, I. Laptev, and C. Schmid, “Actions in context,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL: IEEE, 2009, pp. 2929–2936, ISBN : 978-1-4244-3992-8. DOI: 10/d5bs7p. [Online]. Available: https://ieeexplore. ieee.org/document/5206557/...
2009
-
[161]
What are they doing? : Collective activity classification using spatio-temporal relationship among people,
Wongun Choi, K. Shahid, and S. Savarese, “What are they doing? : Collective activity classification using spatio-temporal relationship among people,” in 2009 IEEE 12th Interna- tional Conference on Computer Vision Workshops, ICCV Workshops, Kyoto, Japan: IEEE, 2009, pp. 1282–1...
2009
-
[162]
Bradlow, ALLSSTAR: Archive of l1 and l2 scripted and spontaneous transcripts and recordings, 2010
A. Bradlow, ALLSSTAR: Archive of l1 and l2 scripted and spontaneous transcripts and recordings, 2010. [Online]. Available: https : / / speechbox . linguistics . northwestern.edu/#!/?goto=allsstar (visited on 05/01/2024)
2010
-
[163]
Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions,
K. Ganesan, C. Zhai, and J. Han, “Opinosis: A graph based approach to abstractive summa- rization of highly redundant opinions,” in Proceedings of the 23rd International Conference on Computational Linguistics (Coling 2010), C.-R. Huang and D. Jurafsky, Eds., Beijing, China: C...
2010
-
[164]
Semantic similarity and relatedness between clinical terms: An experimental study,
S. Pakhomov, B. McInnes, T. Adam, Y . Liu, T. Pedersen, and G. B. Melton, “Semantic similarity and relatedness between clinical terms: An experimental study,” AMIA Annual Symposium Proceedings, vol. 2010, pp. 572–576, 2010,ISSN : 1942-597X. [Online]. Available: https://www.ncb...
2010
-
[165]
Evaluation of topic identification methods on arabic corpora,
M. Abbas, K. Smaïli, and D. Berkani, “Evaluation of topic identification methods on arabic corpora,” Journal of Digital Information Management, vol. 9, no. 5, 2011. [Online]. Available: https://www.dline.info/fpaper/jdim/v9i5/1.pdf (visited on 10/02/2024)
2011
-
[166]
HMDB: A large video database for human motion recognition,
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: A large video database for human motion recognition,” in 2011 International Conference on Computer Vision , Barcelona, Spain: IEEE, 2011, pp. 2556–2563, ISBN : 978-1-4577-1102-2. DOI: 10/fxpf8k. [Online]. Availa...
2011
-
[167]
Soomro, A
K. Soomro, A. R. Zamir, and M. Shah, UCF101: A dataset of 101 human actions classes from videos in the wild, 2012. arXiv: 1212.0402[cs]. [Online]. Available: http:// arxiv.org/abs/1212.0402 (visited on 05/01/2024)
2012 arXiv
-
[168]
A thousand frames in just a few words: Lin- gual description of videos through latent topics and sparse object stitching,
P. Das, C. Xu, R. F. Doell, and J. J. Corso, “A thousand frames in just a few words: Lin- gual description of videos through latent topics and sparse object stitching,” in 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA: IEEE, 2013, pp. 2634–...
2013
-
[169]
Asgard: A portable architecture for multilin- gual dialogue systems,
J. Liu, P. Pasupat, S. Cyphers, and J. Glass, “Asgard: A portable architecture for multilin- gual dialogue systems,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada: IEEE, 2013, pp. 8386–8390, ISBN : 978-1-4799- 0356-6. D...
2013
-
[170]
P. Malo, A. Sinha, P. Takala, P. Korhonen, and J. Wallenius,Good debt or bad debt: Detecting semantic orientations in economic texts, 2013. arXiv: 1307.5336[cs,q-fin]. [Online]. Available: http://arxiv.org/abs/1307.5336 (visited on 05/29/2024)
2013 arXiv
-
[171]
Combining embedded accelerometers with computer vision for recognizing food preparation activities,
S. Stein and S. J. McKenna, “Combining embedded accelerometers with computer vision for recognizing food preparation activities,” in Proceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing, ser. UbiComp ’13, New York, NY, USA: Associati...
2013 doi
-
[172]
Spatial pattern templates for recognition of objects with regu- lar structure,
R. Tyleˇcek and R. Šára, “Spatial pattern templates for recognition of objects with regu- lar structure,” in Pattern Recognition, J. Weickert, M. Hein, and B. Schiele, Eds., Berlin, Heidelberg: Springer, 2013, pp. 364–374, ISBN : 978-3-642-40602-7. DOI: 10/ggwb5g
2013
-
[173]
Bojanowski, R
P. Bojanowski, R. Lajugie, F. Bach, et al., Weakly supervised action labeling in videos under ordering constraints, 2014. arXiv: 1407.1208[cs]. [Online]. Available: http: //arxiv.org/abs/1407.1208 (visited on 05/29/2024)
2014 arXiv
-
[174]
Creating summaries from user videos,
M. Gygli, H. Grabner, H. Riemenschneider, and L. Van Gool, “Creating summaries from user videos,” in Computer Vision – ECCV 2014 , D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds., vol. 8695, Series Title: Lecture Notes in Computer Science, Cham: Springer International...
2014 doi
-
[175]
VideoStory: A new multimedia embedding for few-example recognition and translation of events,
A. Habibian, T. Mensink, and C. G. Snoek, “VideoStory: A new multimedia embedding for few-example recognition and translation of events,” in Proceedings of the 22nd ACM international conference on Multimedia, Orlando Florida USA: ACM, 2014, pp. 17–26,ISBN : 978-1-4503-3063-3. ...
2014
-
[176]
Large-scale video classification with convolutional neural networks,
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA: IEEE, 2014, pp. 1725– 1732, ISBN : 978-1-...
2014
-
[177]
Free english and czech telephone speech corpus shared under the CC-BY-SA 3.0 license,
M. Korvas, O. Plátek, O. Dušek, L. Žilka, and F. Jurˇcíˇcek, “Free english and czech telephone speech corpus shared under the CC-BY-SA 3.0 license,” in Proceedings of the Ninth In- ternational Conference on Language Resources and Evaluation (LREC’14), N. Calzolari, K. Choukri,...
2014
-
[178]
The language of actions: Recovering the syntax and semantics of goal-directed human activities,
H. Kuehne, A. Arslan, and T. Serre, “The language of actions: Recovering the syntax and semantics of goal-directed human activities,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA: IEEE, 2014, pp. 780–787, ISBN : 978-1-4799- 5118-5. DOI:...
2014
-
[179]
StoryGraphs: Visualizing character interactions as a timeline,
M. Tapaswi, M. Bauml, and R. Stiefelhagen, “StoryGraphs: Visualizing character interactions as a timeline,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 827–834. [Online]. Available: https://openaccess.thecvf. com/content_cvpr_201...
2014
-
[180]
Extraction of relations between genes and diseases from text and large-scale data analysis: Implications for translational research,
A. Bravo, J. Pinero, N. Queralt-Rosinach, M. Rautschka, and L. I. Furlong, “Extraction of relations between genes and diseases from text and large-scale data analysis: Implications for translational research,” BMC Bioinformatics, vol. 16, no. 1, p. 55, 2015, ISSN : 1471-2105. ...
2015
-
[181]
Building large arabic multi-domain resources for sentiment analysis,
H. ElSahar and S. R. El-Beltagy, “Building large arabic multi-domain resources for sentiment analysis,” in Computational Linguistics and Intelligent Text Processing, A. Gelbukh, Ed., Cham: Springer International Publishing, 2015, pp. 23–34, ISBN : 978-3-319-18117-2. DOI: 10/g6k58r
2015
-
[182]
ActivityNet: A large-scale video benchmark for human activity understanding,
F. C. Heilbron, V . Escorcia, B. Ghanem, and J. C. Niebles, “ActivityNet: A large-scale video benchmark for human activity understanding,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA: IEEE, 2015, pp. 961–970, ISBN : 978- 1-4673-69...
2015
-
[183]
Librispeech: An ASR corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, Queensland, Australia: IEEE, 2015, pp. 5206–5210, IS...
2015
-
[184]
Compositional semantic parsing on semi-structured tables,
P. Pasupat and P. Liang, “Compositional semantic parsing on semi-structured tables,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C....
2015
-
[185]
Rohrbach, M
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele,A dataset for movie description, 2015. arXiv: 1501.02530[cs] . [Online]. Available: http://arxiv.org/abs/1501. 02530 (visited on 05/01/2024)
2015 arXiv
-
[186]
An open/free database and benchmark for uyghur speaker recognition,
A. Rozi, Dong Wang, Zhiyong Zhang, and T. F. Zheng, “An open/free database and benchmark for uyghur speaker recognition,” in 2015 International Conference Oriental COCOSDA held jointly with 2015 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE), Sh...
2015
-
[187]
A. M. Rush, S. Chopra, and J. Weston, A neural attention model for abstractive sentence summarization, 2015. arXiv: 1509.00685[cs]. [Online]. Available: http://arxiv. org/abs/1509.00685 (visited on 05/29/2024)
2015 arXiv
-
[188]
ImageNet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, et al., “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015, ISSN : 1573-
2015
-
[190]
Wang and X
D. Wang and X. Zhang, THCHS-30 : A free chinese speech corpus , 2015. arXiv: 1512. 01882[cs]. [Online]. Available: http://arxiv.org/abs/1512.01882 (visited on 05/01/2024)
2015 arXiv
-
[191]
Weston, A
J. Weston, A. Bordes, S. Chopra, et al., Towards AI-complete question answering: A set of prerequisite toy tasks , 2015. arXiv: 1502 . 05698[cs , stat]. [Online]. Available: http://arxiv.org/abs/1502.05698 (visited on 05/29/2024)
2015 arXiv
-
[192]
TVSum: Summarizing web videos using titles,
Yale Song, J. Vallmitjana, A. Stent, and A. Jaimes, “TVSum: Summarizing web videos using titles,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA: IEEE, 2015, pp. 5179–5187, ISBN : 978-1-4673-6964-0. DOI: 10/gfsj74. [Online]. Availabl...
2015
-
[193]
Abu-El-Haija, N
S. Abu-El-Haija, N. Kothari, J. Lee, et al., YouTube-8m: A large-scale video classification benchmark, 2016. arXiv: 1609 . 08675[cs]. [Online]. Available: http : / / arxiv . org/abs/1609.08675 (visited on 05/01/2024). 41 The Data Provenance Initiative, 2024
2016 arXiv
-
[194]
Alayrac, P
J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien,Unsuper- vised learning from narrated instruction videos, 2016. arXiv: 1506.09215[cs]. [Online]. Available: http://arxiv.org/abs/1506.09215 (visited on 05/01/2024)
2016 arXiv
-
[195]
BRAD 1.0: Book reviews in arabic dataset,
A. Elnagar and O. Einea, “BRAD 1.0: Book reviews in arabic dataset,” in2016 IEEE/ACS 13th International Conference of Computer Systems and Applications (AICCSA) , Agadir, Morocco: IEEE, 2016, pp. 1–8, ISBN : 978-1-5090-4320-0. DOI: 10/g6k6jm . [Online]. Available: http://ieeex...
2016
-
[196]
Ibrahim, S
M. Ibrahim, S. Muralidharan, Z. Deng, A. Vahdat, and G. Mori,A hierarchical deep temporal model for group activity recognition, 2016. arXiv: 1511.06040[cs]. [Online]. Available: http://arxiv.org/abs/1511.06040 (visited on 05/01/2024)
2016 arXiv
-
[197]
Jurczyk, M
T. Jurczyk, M. Zhai, and J. D. Choi, SelQA: A new benchmark for selection-based question answering, 2016. arXiv: 1606.08513[cs]. [Online]. Available:http://arxiv.org/ abs/1606.08513 (visited on 10/02/2024)
2016 arXiv
-
[198]
I. A. El-khair, 1.5 billion words arabic corpus, 2016. arXiv: 1611.04033[cs]. [Online]. Available: http://arxiv.org/abs/1611.04033 (visited on 10/02/2024)
2016 arXiv
-
[199]
Lebret, D
R. Lebret, D. Grangier, and M. Auli, Neural text generation from structured data with application to the biography domain, 2016. arXiv: 1603.07771[cs]. [Online]. Available: http://arxiv.org/abs/1603.07771 (visited on 05/29/2024)
2016 arXiv
-
[200]
Y . Li, Y . Song, L. Cao, et al. , TGIF: A new dataset and benchmark on animated GIF description, 2016. arXiv: 1604 . 02748[cs]. [Online]. Available: http : / / arxiv . org/abs/1604.02748 (visited on 05/01/2024)
2016 arXiv
-
[201]
R. Lowe, N. Pow, I. Serban, and J. Pineau, The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems, 2016. arXiv: 1506.08909[cs]. [Online]. Available: http://arxiv.org/abs/1506.08909 (visited on 10/02/2024)
2016 arXiv
-
[202]
Merity, C
S. Merity, C. Xiong, J. Bradbury, and R. Socher,Pointer sentinel mixture models, 2016. arXiv: 1609.07843[cs]. [Online]. Available: http://arxiv.org/abs/1609.07843 (visited on 05/29/2024)
2016 arXiv
-
[203]
A benchmark dataset and evaluation methodology for video object segmentation,
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), ISSN: 1063-6919, 2016, pp. 724–732...
2016
-
[204]
Rajpurkar, J
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang,SQuAD: 100,000+ questions for machine comprehension of text, 2016. arXiv: 1606.05250[cs]. [Online]. Available: http:// arxiv.org/abs/1606.05250 (visited on 05/29/2024)
2016 arXiv
-
[205]
Rohrbach, A
A. Rohrbach, A. Torabi, M. Rohrbach, et al. , Movie description , 2016. arXiv: 1605 . 03705[cs]. [Online]. Available: http : / / arxiv . org / abs / 1605 . 03705(vis- ited on 05/29/2024)
2016
-
[206]
Recognizing fine-grained and composite activities using hand-centric features and script data,
M. Rohrbach, A. Rohrbach, M. Regneri, et al., “Recognizing fine-grained and composite activities using hand-centric features and script data,” International Journal of Computer Vision, vol. 119, no. 3, pp. 346–373, 2016, ISSN : 0920-5691, 1573-1405. DOI: 10/f8w6kp. arXiv: 1502...
2016 arXiv
-
[207]
Shahroudy, J
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, NTU RGB+d: A large scale dataset for 3d human activity analysis , 2016. arXiv: 1604 . 02808[cs]. [Online]. Available: http : //arxiv.org/abs/1604.02808 (visited on 05/02/2024)
2016 arXiv
-
[208]
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, Hollywood in homes: Crowdsourcing data collection for activity understanding , 2016. arXiv: 1604. 01753[cs]. [Online]. Available: http://arxiv.org/abs/1604.01753 (visited on 05/01/2024)
2016 arXiv
-
[209]
Tapaswi, Y
M. Tapaswi, Y . Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler,MovieQA: Under- standing stories in movies through question-answering, 2016. arXiv: 1512.02902[cs]. [Online]. Available: http://arxiv.org/abs/1512.02902 (visited on 05/29/2024). 42 The Data Provenance...
2016 arXiv
-
[210]
MSR-VTT: A large video description dataset for bridging video and language,
J. Xu, T. Mei, T. Yao, and Y . Rui, “MSR-VTT: A large video description dataset for bridging video and language,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), ISSN: 1063-6919, 2016, pp. 5288–5296. DOI: 10/ggv9gj. [Online]. Available: https://ieeex...
2016
-
[211]
The value of semantic parse labeling for knowledge base question answering,
W.-t. Yih, M. Richardson, C. Meek, M.-W. Chang, and J. Suh, “The value of semantic parse labeling for knowledge base question answering,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), K. Erk and N. A. Smith...
2016
-
[212]
Zeng, T.-H
K.-H. Zeng, T.-H. Chen, J. C. Niebles, and M. Sun, Title generation for user generated videos, 2016. arXiv: 1608.07068[cs]. [Online]. Available: http://arxiv.org/ abs/1608.07068 (visited on 05/01/2024)
2016 arXiv
-
[213]
Zhang, J
X. Zhang, J. Zhao, and Y . LeCun,Character-level convolutional networks for text classifica- tion, 2016. arXiv: 1509.01626[cs]. [Online]. Available: http://arxiv.org/abs/ 1509.01626 (visited on 05/29/2024)
2016 arXiv
-
[214]
MARS: A video benchmark for large-scale person re- identification,
L. Zheng, Z. Bie, Y . Sun, et al., “MARS: A video benchmark for large-scale person re- identification,” in Computer Vision – ECCV 2016 , B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., Cham: Springer International Publishing, 2016, pp. 868–884, ISBN : 978-3- 319-46466-4. DO...
2016
-
[215]
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline , 2017. arXiv: 1709 . 05522[cs]. [Online]. Available: http://arxiv.org/abs/1709.05522 (visited on 05/01/2024)
2017 arXiv
-
[216]
Chouldechova, Fair prediction with disparate impact: A study of bias in recidivism pre- diction instruments, 2017
A. Chouldechova, Fair prediction with disparate impact: A study of bias in recidivism pre- diction instruments, 2017. arXiv: 1703.00056[cs,stat]. [Online]. Available: http: //arxiv.org/abs/1703.00056 (visited on 10/02/2024)
2017 arXiv
-
[217]
Frames: A corpus for adding memory to goal- oriented dialogue systems,
L. El Asri, H. Schulz, S. Sharma, et al., “Frames: A corpus for adding memory to goal- oriented dialogue systems,” in Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, K. Jokinen, M. Stede, D. DeVault, and A. Louis, Eds., Saarbrücken, Germany: Associati...
2017
-
[218]
Eric and C
M. Eric and C. D. Manning, Key-value retrieval networks for task-oriented dialogue, 2017. arXiv: 1705.05414[cs] . [Online]. Available: http://arxiv.org/abs/1705. 05414 (visited on 05/02/2024)
2017 arXiv
-
[219]
D. F. Fouhey, W.-c. Kuo, A. A. Efros, and J. Malik, From lifestyle vlogs to everyday inter- actions, 2017. arXiv: 1712.02310[cs]. [Online]. Available: http://arxiv.org/ abs/1712.02310 (visited on 05/01/2024)
2017 arXiv
-
[220]
something something
R. Goyal, S. E. Kahou, V . Michalski,et al., The "something something" video database for learning and evaluating visual common sense, 2017. arXiv: 1706.04261[cs]. [Online]. Available: http://arxiv.org/abs/1706.04261 (visited on 05/01/2024)
2017 arXiv
-
[221]
Ha and D
D. Ha and D. Eck, A neural representation of sketch drawings , 2017. arXiv: 1704 . 03477[cs,stat] . [Online]. Available: http://arxiv.org/abs/1704.03477 (visited on 10/02/2024)
2017 arXiv
-
[222]
The THUMOS challenge on action recognition for videos
H. Idrees, A. R. Zamir, Y .-G. Jiang, et al., “The THUMOS challenge on action recognition for videos "in the wild",” Computer Vision and Image Understanding, vol. 155, pp. 1–23, 2017, ISSN : 10773142. DOI: 10/f9rwnr. arXiv: 1604.06182[cs]. [Online]. Available: http://arxiv.org...
2017 arXiv
-
[223]
Ito and L
K. Ito and L. Johnson, The LJ Speech Dataset , 2017. [Online]. Available: https : / / keithito.com/LJ-Speech-Dataset (visited on 05/01/2024)
2017
-
[224]
S. Iyer, I. Konstas, A. Cheung, J. Krishnamurthy, and L. Zettlemoyer, Learning a neural semantic parser from user feedback, 2017. arXiv: 1704.08760[cs]. [Online]. Available: http://arxiv.org/abs/1704.08760 (visited on 10/02/2024). 43 The Data Provenance Initiative, 2024
2017 arXiv
-
[225]
Search-based neural structured learning for se- quential question answering,
M. Iyyer, W.-t. Yih, and M. -W. Chang, “Search-based neural structured learning for se- quential question answering,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), R. Barzilay and M.-Y . Kan, Eds., Vancouver...
2017
-
[226]
Joshi, E
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer, TriviaQA: A large scale distantly su- pervised challenge dataset for reading comprehension, 2017. arXiv: 1705.03551[cs]. [Online]. Available: http://arxiv.org/abs/1705.03551 (visited on 05/29/2024)
2017 arXiv
-
[227]
W. Kay, J. Carreira, K. Simonyan,et al., The kinetics human action video dataset, 2017. arXiv: 1705.06950[cs]. [Online]. Available: http://arxiv.org/abs/1705.06950 (visited on 05/01/2024)
2017 arXiv
-
[228]
Korzinek, K
D. Korzinek, K. Marasek, L. Brocki, and K. Wolk,Polish read speech corpus for speech tools and services, 2017. arXiv: 1706.00245[cs] . [Online]. Available: http://arxiv. org/abs/1706.00245 (visited on 05/29/2024)
2017 arXiv
-
[229]
G. Lai, Q. Xie, H. Liu, Y . Yang, and E. Hovy,RACE: Large-scale ReAding comprehension dataset from examinations, 2017. arXiv: 1704.04683[cs]. [Online]. Available: http: //arxiv.org/abs/1704.04683 (visited on 05/29/2024)
2017 arXiv
-
[230]
Lewis, D
M. Lewis, D. Yarats, Y . N. Dauphin, D. Parikh, and D. Batra,Deal or no deal? end-to-end learning for negotiation dialogues, 2017. arXiv: 1706.05125[cs]. [Online]. Available: http://arxiv.org/abs/1706.05125 (visited on 05/29/2024)
2017 arXiv
-
[231]
Free linguistic and speech resources for tibetan,
G. Li, H. Yu, T. F. Zheng, J. Yan, and S. Xu, “Free linguistic and speech resources for tibetan,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Kuala Lumpur: IEEE, 2017, pp. 733–736, ISBN : 978-1-5386- 1542-3. DOI...
2017
-
[232]
W. Ling, D. Yogatama, C. Dyer, and P. Blunsom,Program induction by rationale generation : Learning to solve and explain algebraic word problems, 2017. arXiv: 1705.04146[cs]. [Online]. Available: http://arxiv.org/abs/1705.04146 (visited on 05/29/2024)
2017 arXiv
-
[233]
C. Liu, Y . Hu, Y . Li, S. Song, and J. Liu,PKU-MMD: A large scale benchmark for continuous multi-modal human action understanding , 2017. arXiv: 1703 . 07475[cs]. [Online]. Available: http://arxiv.org/abs/1703.07475 (visited on 05/01/2024)
2017 arXiv
-
[234]
Mrksic, D
N. Mrksic, D. O. Seaghdha, T.-H. Wen, B. Thomson, and S. Young,Neural belief tracker: Data-driven dialogue state tracking, 2017. arXiv: 1606.03777[cs]. [Online]. Available: http://arxiv.org/abs/1606.03777 (visited on 05/01/2024)
2017 arXiv
-
[235]
Novikova, O
J. Novikova, O. Dušek, and V . Rieser, The e2e dataset: New challenges for end-to-end generation, 2017. arXiv: 1706 . 09254[cs]. [Online]. Available: http : / / arxiv . org/abs/1706.09254 (visited on 05/29/2024)
2017 arXiv
-
[236]
A. See, P. J. Liu, and C. D. Manning, Get to the point: Summarization with pointer-generator networks, 2017. arXiv: 1704.04368[cs]. [Online]. Available: http://arxiv.org/ abs/1704.04368 (visited on 05/29/2024)
2017 arXiv
-
[237]
Sharghi, J
A. Sharghi, J. S. Laurel, and B. Gong, Query-focused video summarization: Dataset, evalua- tion, and a memory network based approach, 2017. arXiv: 1707.04960[cs]. [Online]. Available: http://arxiv.org/abs/1707.04960 (visited on 05/01/2024)
2017 arXiv
-
[238]
A free kazakh speech database and a speech recognition baseline,
Y . Shi, A. Hamdullah, Z. Tang, D. Wang, and T. F. Zheng, “A free kazakh speech database and a speech recognition baseline,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , Kuala Lumpur: IEEE, 2017, pp. 745–748, IS...
2017
-
[239]
Welbl, N
J. Welbl, N. F. Liu, and M. Gardner, Crowdsourcing multiple choice science questions, 2017. arXiv: 1707.06209[cs, stat]. [Online]. Available: http://arxiv.org/abs/ 1707.06209 (visited on 05/29/2024). 44 The Data Provenance Initiative, 2024
2017 arXiv
-
[240]
Yeung, O
S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, and L. Fei-Fei, Every mo- ment counts: Dense detailed labeling of actions in complex videos , 2017. arXiv: 1507 . 05738[cs]. [Online]. Available: http://arxiv.org/abs/1507.05738 (visited on 05/02/2024)
2017 arXiv
-
[241]
Zhong, C
V . Zhong, C. Xiong, and R. Socher,Seq2sql: Generating structured queries from natural lan- guage using reinforcement learning, 2017. arXiv: 1709.00103[cs]. [Online]. Available: http://arxiv.org/abs/1709.00103 (visited on 05/01/2024)
2017 arXiv
-
[242]
L. Zhou, C. Xu, and J. J. Corso, Towards automatic learning of procedures from web instructional videos, 2017. arXiv: 1703.09788[cs]. [Online]. Available: https:// arxiv.org/abs/1703.09788v3 (visited on 05/02/2024)
2017 arXiv
-
[243]
Bajaj, D
P. Bajaj, D. Campos, N. Craswell, et al., MS MARCO: A human generated MAchine reading COmprehension dataset, 2018. arXiv: 1611 . 09268[cs]. [Online]. Available: http : //arxiv.org/abs/1611.09268 (visited on 05/29/2024)
2018 arXiv
-
[244]
J. A. Botha, M. Faruqui, J. Alex, J. Baldridge, and D. Das, Learning to split and rephrase from wikipedia edit history, 2018. arXiv: 1808.09468[cs]. [Online]. Available: http: //arxiv.org/abs/1808.09468 (visited on 10/02/2024)
2018 arXiv
-
[245]
Carreira, E
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman, A short note about kinetics-600, 2018. arXiv: 1808.01340[cs] . [Online]. Available: http://arxiv. org/abs/1808.01340 (visited on 05/01/2024)
2018 arXiv
-
[246]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, et al., Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018. arXiv: 1803.05457[cs]. [Online]. Available: http://arxiv.org/abs/1803.05457 (visited on 05/29/2024)
2018 arXiv
-
[247]
Coucke, A
A. Coucke, A. Saade, A. Ball, et al., Snips voice platform: An embedded spoken language un- derstanding system for private-by-design voice interfaces, 2018. arXiv: 1805.10190[cs]. [Online]. Available: http://arxiv.org/abs/1805.10190 (visited on 05/01/2024)
2018 arXiv
-
[248]
Damen, H
D. Damen, H. Doughty, G. M. Farinella, et al. , Scaling egocentric vision: The EPIC- KITCHENS dataset, 2018. arXiv: 1804 . 02748[cs]. [Online]. Available: http : / / arxiv.org/abs/1804.02748 (visited on 05/01/2024)
2018 arXiv
-
[249]
J. Du, X. Na, X. Liu, and H. Bu, AISHELL-2: Transforming mandarin ASR research into industrial scale, 2018. arXiv: 1808.10583[cs]. [Online]. Available: http://arxiv. org/abs/1808.10583 (visited on 05/31/2024)
2018 arXiv
-
[250]
Faruqui and D
M. Faruqui and D. Das, Identifying well-formed natural language questions, 2018. arXiv: 1808.09419[cs]. [Online]. Available: http://arxiv.org/abs/1808.09419 (visited on 05/29/2024)
2018 arXiv
-
[251]
Gorrell, K
G. Gorrell, K. Bontcheva, L. Derczynski, E. Kochkina, M. Liakata, and A. Zubiaga, Ru- mourEval 2019: Determining rumour veracity and support for rumours, 2018. arXiv: 1809. 06683[cs]. [Online]. Available: http://arxiv.org/abs/1809.06683 (visited on 05/29/2024)
2019 arXiv
-
[252]
C. Gu, C. Sun, D. A. Ross, et al., AVA: A video dataset of spatio-temporally localized atomic visual actions, 2018. arXiv: 1705.08421[cs]. [Online]. Available: http://arxiv. org/abs/1705.08421 (visited on 05/29/2024)
2018 arXiv
-
[253]
MMQA: A multi-domain multi- lingual question-answering framework for english and hindi,
D. Gupta, S. Kumari, A. Ekbal, and P. Bhattacharyya, “MMQA: A multi-domain multi- lingual question-answering framework for english and hindi,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), N. Calzolari, K. Choukri, C....
2018
-
[254]
Gupta, R
S. Gupta, R. Shah, M. Mohit, A. Kumar, and M. Lewis, Semantic parsing for task ori- ented dialog using hierarchical representations, 2018. arXiv: 1810.07942[cs]. [Online]. Available: http://arxiv.org/abs/1810.07942 (visited on 05/01/2024)
2018 arXiv
-
[255]
H. He, D. Chen, A. Balakrishnan, and P. Liang, Decoupling strategy and generation in negotiation dialogues, 2018. arXiv: 1808.09637[cs]. [Online]. Available: http:// arxiv.org/abs/1808.09637 (visited on 05/01/2024). 45 The Data Provenance Initiative, 2024
2018 arXiv
-
[256]
Localizing moments in video with temporal language,
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with temporal language,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, E...
2018
-
[257]
TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,
F. Hernandez, V . Nguyen, S. Ghannay, N. Tomashenko, and Y . Estève, “TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Lecture Notes in Computer Science , vol. 11096, Springer, Cham, 2018, pp. 198–208. DOI: 10 . 1007/978- 3- 319-...
2018 arXiv
-
[258]
SciTaiL: A textual entailment dataset from science question answering,
T. Khot, A. Sabharwal, and P. Clark, “SciTaiL: A textual entailment dataset from science question answering,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018, ISSN : 2374-3468, 2159-5399. DOI: 10/grm22d. [Online]. Available: https: / / ojs ....
2018
-
[259]
S. Kim, I. Kang, and N. Kwak, Semantic sentence matching with densely-connected recurrent and co-attentive information, 2018. arXiv: 1805.11360[cs]. [Online]. Available: http: //arxiv.org/abs/1805.11360 (visited on 10/02/2024)
2018 arXiv
-
[260]
Crowd-sourced speech corpora for javanese, sundanese, sinhala, nepali, and bangladeshi bengali,
O. Kjartansson, S. Sarin, K. Pipatsrisawat, M. Jansche, and L. Ha, “Crowd-sourced speech corpora for javanese, sundanese, sinhala, nepali, and bangladeshi bengali,” in6th Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU 2018), ISCA, 2018, pp. 52–55....
2018
-
[261]
B. M. Lake and M. Baroni, Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks, 2018. arXiv: 1711.00350[cs]. [Online]. Available: http://arxiv.org/abs/1711.00350 (visited on 05/29/2024)
2018 arXiv
-
[262]
Event representations for automated story generation with deep neural nets,
L. J. Martin, P. Ammanabrolu, X. Wang,et al., “Event representations for automated story generation with deep neural nets,” Proceedings of the AAAI Conference on Artificial In- telligence, vol. 32, no. 1, 2018, ISSN : 2374-3468, 2159-5399. DOI: 10/g6k72p . arXiv: 1706.01331[cs...
2018 arXiv
-
[263]
Mihaylov, P
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal,Can a suit of armor conduct electricity? a new dataset for open book question answering, 2018. arXiv: 1809.02789[cs]. [Online]. Available: http://arxiv.org/abs/1809.02789 (visited on 05/29/2024)
2018 arXiv
-
[264]
Moniz and L
N. Moniz and L. Torgo, Multi-source social feedback of online news feeds , 2018. arXiv: 1801.07055[cs]. [Online]. Available: http://arxiv.org/abs/1801.07055 (visited on 10/02/2024)
2018 arXiv
-
[265]
Mrkši ´c and I
N. Mrkši ´c and I. Vuli ´c, Fully statistical neural belief tracking , 2018. arXiv: 1805 . 11350[cs]. [Online]. Available: http : / / arxiv . org / abs / 1805 . 11350(vis- ited on 05/01/2024)
2018
-
[266]
Nagrani, J
A. Nagrani, J. S. Chung, and A. Zisserman, VoxCeleb: A large-scale speaker identifi- cation dataset , 2018. DOI: 10 . 21437 / Interspeech . 2017 - 950. arXiv: 1706 . 08612[cs]. [Online]. Available: http://arxiv.org/abs/1706.08612 (visited on 05/29/2024)
2018 arXiv
-
[267]
Narayan, S
S. Narayan, S. B. Cohen, and M. Lapata, Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization, 2018. arXiv: 1808. 08745[cs]. [Online]. Available: http://arxiv.org/abs/1808.08745 (visited on 05/29/2024)
2018 arXiv
-
[268]
Royer, K
A. Royer, K. Bousmalis, S. Gouws, et al., XGAN: Unsupervised image-to-image translation for many-to-many mappings, 2018. arXiv: 1711.05139[cs]. [Online]. Available: http: //arxiv.org/abs/1711.05139 (visited on 10/02/2024)
2018 arXiv
-
[269]
Saeidi, M
M. Saeidi, M. Bartolo, P. Lewis, et al., Interpretation of natural language rules in conver- sational machine reading, 2018. arXiv: 1809.01494[cs,stat] . [Online]. Available: http://arxiv.org/abs/1809.01494 (visited on 10/02/2024). 46 The Data Provenance Initiative, 2024
2018 arXiv
-
[270]
A. Saha, R. Aralikatte, M. M. Khapra, and K. Sankaranarayanan, DuoRC: Towards complex language understanding with paraphrased reading comprehension , 2018. arXiv: 1804 . 07927[cs]. [Online]. Available: http://arxiv.org/abs/1804.07927 (visited on 05/29/2024)
2018 arXiv
-
[271]
Sanabria, O
R. Sanabria, O. Caglayan, S. Palaskar, et al., How2: A large-scale dataset for multimodal language understanding, 2018. arXiv: 1811.00347[cs] . [Online]. Available: http: //arxiv.org/abs/1811.00347 (visited on 05/01/2024)
2018 arXiv
-
[272]
P. Shah, D. Hakkani-Tür, G. Tür, et al., Building a conversational agent overnight with dia- logue self-play, 2018. arXiv: 1801.04871[cs]. [Online]. Available: http://arxiv. org/abs/1801.04871 (visited on 05/01/2024)
2018 arXiv
-
[273]
Unsupervised abstractive meeting summarization with multi-sentence compression and budgeted submodular maximization,
G. Shang, W. Ding, Z. Zhang, et al., “Unsupervised abstractive meeting summarization with multi-sentence compression and budgeted submodular maximization,” in Proceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. ...
2018
-
[274]
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, Actor and observer: Joint modeling of first and third-person videos, 2018. arXiv: 1804.09627[cs]. [Online]. Available: http://arxiv.org/abs/1804.09627 (visited on 05/01/2024)
2018 arXiv
-
[275]
Tafjord, P
O. Tafjord, P. Clark, M. Gardner, W.-t. Yih, and A. Sabharwal, QuaRel: A dataset and models for answering questions about qualitative relationships, 2018. arXiv: 1811.08048[cs]. [Online]. Available: http://arxiv.org/abs/1811.08048 (visited on 10/02/2024)
2018 arXiv
-
[276]
The web as a knowledge-base for answering complex questions,
A. Talmor and J. Berant, “The web as a knowledge-base for answering complex questions,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), M. Walker, H. Ji, ...
2018
-
[277]
Thorne, A
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, FEVER: A large-scale dataset for fact extraction and VERification, 2018. arXiv: 1803.05355[cs]. [Online]. Available: http://arxiv.org/abs/1803.05355 (visited on 05/29/2024)
2018 arXiv
-
[278]
Vicol, M
P. Vicol, M. Tapaswi, L. Castrejon, and S. Fidler, MovieGraphs: Towards understanding human-centric situations from videos, 2018. arXiv: 1712.06761[cs]. [Online]. Available: http://arxiv.org/abs/1712.06761 (visited on 05/29/2024)
2018 arXiv
-
[279]
AirDialogue: An environment for goal-oriented dialogue research,
W. Wei, Q. Le, A. Dai, and J. Li, “AirDialogue: An environment for goal-oriented dialogue research,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds., Brussels, Belgium: As- soc...
2018
-
[280]
Welbl, P
J. Welbl, P. Stenetorp, and S. Riedel,Constructing datasets for multi-hop reading compre- hension across documents, 2018. arXiv: 1710.06481[cs]. [Online]. Available: http: //arxiv.org/abs/1710.06481 (visited on 10/02/2024)
2018 arXiv
-
[281]
Z. Wu, B. Ramsundar, E. N. Feinberg, et al., MoleculeNet: A benchmark for molecular machine learning, 2018. arXiv: 1703.00564[physics, stat]. [Online]. Available: http://arxiv.org/abs/1703.00564 (visited on 10/02/2024)
2018 arXiv
-
[282]
Z. Yang, P. Qi, S. Zhang, et al., HotpotQA: A dataset for diverse, explainable multi-hop question answering, 2018. arXiv: 1809 . 09600[cs]. [Online]. Available: http : / / arxiv.org/abs/1809.09600 (visited on 05/29/2024)
2018 arXiv
-
[283]
Amini, S
A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi, MathQA: Towards interpretable math word problem solving with operation-based formalisms, 2019. arXiv: 1905.13319[cs] . [Online]. Available: http://arxiv.org/abs/1905. 13319 (visited on 05/29/2024)...
2019 arXiv
-
[284]
Balakrishnan, J
A. Balakrishnan, J. Rao, K. Upasani, M. White, and R. Subba, Constrained decoding for neural NLG from compositional representations in task-oriented dialogue , 2019. arXiv: 1906.07220[cs]. [Online]. Available: http://arxiv.org/abs/1906.07220 (visited on 10/02/2024)
2019 arXiv
-
[285]
The spoken wikipedia corpus collection: Harvesting, alignment and an application to hyperlistening,
T. Baumann, A. Köhn, and F. Hennig, “The spoken wikipedia corpus collection: Harvesting, alignment and an application to hyperlistening,”Language Resources and Evaluation, vol. 53, no. 2, pp. 303–329, 2019, ISSN : 1574-0218. DOI: 10/gq5xdf. [Online]. Available: https: //doi.or...
2019 doi
-
[286]
Y . Bisk, R. Zellers, R. L. Bras, J. Gao, and Y . Choi, PIQA: Reasoning about physical commonsense in natural language, 2019. arXiv: 1911.11641[cs]. [Online]. Available: http://arxiv.org/abs/1911.11641 (visited on 05/29/2024)
2019 arXiv
-
[287]
Byrne, K
B. Byrne, K. Krishnamoorthi, C. Sankar, et al., Taskmaster-1: Toward a realistic and diverse dialog dataset, 2019. arXiv: 1909.05358[cs]. [Online]. Available: http://arxiv. org/abs/1909.05358 (visited on 05/01/2024)
2019 arXiv
-
[288]
Named entity dis- ambiguation using deep learning on graphs,
A. Cetoli, M. Akbari, S. Bragaglia, A. D. O’Harney, and M. Sloan, “Named entity dis- ambiguation using deep learning on graphs,” in vol. 11438, 2019, pp. 78–86. DOI: 10 . 1007/978- 3- 030- 15719- 7_10 . arXiv: 1810.09164[cs] . [Online]. Available: http://arxiv.org/abs/1810.091...
2019 arXiv
-
[290]
[Online]
arXiv: 1906.02059[cs] . [Online]. Available: http:// arxiv.org/abs/ 1906.02059 (visited on 10/02/2024)
1906 arXiv
-
[294]
Clark, K
C. Clark, K. Lee, M. -W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, BoolQ: Exploring the surprising difficulty of natural yes/no questions , 2019. arXiv: 1905 . 10044[cs]. [Online]. Available: http : / / arxiv . org / abs / 1905 . 10044(vis- ited on 05/29/2024)
2019
-
[297]
Dasigi, N
P. Dasigi, N. F. Liu, A. Marasovi´c, N. A. Smith, and M. Gardner, Quoref: A reading com- prehension dataset with questions requiring coreferential reasoning, 2019. arXiv: 1908. 05803[cs]. [Online]. Available: http://arxiv.org/abs/1908.05803 (visited on 05/29/2024). 48 The Data...
2019 arXiv
-
[298]
MuST-c: A multilingual speech translation corpus,
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-c: A multilingual speech translation corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (...
2019
-
[299]
Dinan, V
E. Dinan, V . Logacheva, V . Malykh,et al., The second conversational intelligence challenge (ConvAI2), 2019. arXiv: 1902.00098[cs]. [Online]. Available:http://arxiv.org/ abs/1902.00098 (visited on 05/01/2024)
2019 arXiv
-
[300]
Dinan, S
E. Dinan, S. Roller, K. Shuster, A. Fan, M. Auli, and J. Weston, Wizard of wikipedia: Knowledge-powered conversational agents, 2019. arXiv: 1811 . 01241[cs]. [Online]. Available: http://arxiv.org/abs/1811.01241 (visited on 05/01/2024)
2019 arXiv
-
[1405]
[Online]
DOI: 10/gcgk7w. [Online]. Available:https://doi.org/10.1007/s11263- 015-0816-y (visited on 05/29/2024)
2024 doi
-
[2019]
DOI: 10.1109/TETCI.2019.2892755
2019
-
[2020]
[Online]
arXiv: 2010.14701 [cs.LG] . [Online]. Available: https://arxiv.org/ abs/2010.14701
2010 arXiv
-
[2023]
arXiv: 2310.12941 [cs.LG]
-
[2024]
Available: https : / / www
[Online]. Available: https : / / www . 404media . co / nvidia - sued - for - scraping-youtube-after-404-media-investigation/
-
[4294]
Available: https://proceedings.neurips.cc/paper_files/ paper/2020/file/2cd4e8a2ce081c3d7c32c3cde4312ef7-Paper.pdf
[Online]. Available: https://proceedings.neurips.cc/paper_files/ paper/2020/file/2cd4e8a2ce081c3d7c32c3cde4312ef7-Paper.pdf
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.