Pith. sign in

REVIEW 3 major objections 4 minor 139 references

LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that LEAP, an LLM-powered library, turns natural-language social science queries over unstructured data into ML annotations and SQL analysis, achieving 100% pass@3 and 92% pass@1 on a new 120-query benchmark at an average…

desk verdict The engineering is real and the ablations are solid, but the headline performance numbers rest on a benchmark whose ground truth is never explained, and the paper's own cost arithmetic makes independent human annotation implausible. read the letter →

arxiv 2501.03892 v1 pith:6PAVMOPZ submitted 2025-01-07 cs.DB

classification cs.DB
keywords naturallanguagetoSQLunstructureddatamachinelearningannotationsocialsciencequeriesvaguequeryhandlingLLMfunctioncallingbenchmarkdatasetend-to-endcost
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Social scientists increasingly want to ask research questions directly over raw text, video, and PDFs—questions like which posts are persuasive or whether public mood predicts economic indicators—but the semantic information they need is not stored in the data and must be extracted by machine-learning models. The paper introduces QUIET-ML, a benchmark of 120 real social-science queries with ground-truth answers, and LEAP, an end-to-end library that takes a natural-language query and raw data as input and returns a result. LEAP first decides whether the query is vague, plans a chain of ML functions, annotates the data into a structured table, then generates and executes SQL-style code over that table. The paper reports 92% pass@1 and 100% pass@3 on QUIET-ML at an average end-to-end cost of $1.06 per query, of which code generation costs $0.02. If true, this would let domain experts get reliable answers to social-science questions on unstructured data without manually selecting models or writing analysis code.

What carries the argument

The load-bearing mechanism is a four-part LLM pipeline. The forward planning filter uses chain-of-thought prompting to decide whether a query is vague and to produce a planned ML function chain; it also returns reformulated queries for vague cases. The stage selector runs table generation, code generation, execution, and display as function calls, tracking progress. Within table generation, the supported ML function list is organized as a function tree so only one leaf node's functions are passed to the LLM at a time—this fits the token limit and cuts cost by 55%. Doubly linked lists connect functions with mutual output dependencies (for example, emotion classification output feeds emotion-trigger identification), raising accuracy on implicit-dependency queries from 20% to 87.7%, and alias check blocks reuse existing columns instead of rerunning expensive ML functions. The code generator is an NL2SQL step over the annotated table.

What would settle it

A direct check: take a sample of QUIET-ML queries, compute answers with independently trained models or human annotators, and compare those answers to LEAP's outputs; if agreement is substantially below the reported 92% pass@1, the benchmark would be measuring reuse of the library's own functions rather than correct social-science answers.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that the barrier to ML-based social science analysis is not the ML models themselves but the glue: choosing the right models, ordering dependent calls, noticing vague queries, and translating results into queries. LEAP packages that glue into an automatic pipeline. Given raw unstructured data and a natural-language query, a forward planning filter decides whether the query is deterministic and, if it is vague, stops and proposes specified alternatives; otherwise it plans a function chain. A stage selector then extends the data with ML annotations—selecting functions from a tree so the prompt fits the token limit, following doubly linked dependency edges for implicit calls like emotion classification before emotion-trigger extraction—and generates and executes the final query. On the 120-query QUIET-ML benchmark, LEAP achieves 92% pass@1, 100% pass@3 and pass@5, and 96.97% success on the 33 vague queries, compared with 41.21% for the best baseline.

Load-bearing premise

The load-bearing premise is that the QUIET-ML ground-truth answers are correct and were computed independently of the ML functions LEAP selects, so that matching them demonstrates semantic correctness rather than self-consistency with its own annotation functions.

Editorial extensions

If this is right

  • Social science researchers can issue queries in plain language over raw Tweets, PDFs, or videos and receive an answer plus a structured, annotated table, without writing ML code.
  • Queries that are too vague to be answered deterministically are rejected with suggested reformulations, so users learn what is missing instead of getting wrong answers.
  • For non-vague queries with unspecified numeric thresholds, LEAP warns and then picks a data-driven value, preserving the query's intent.
  • End-to-end cost per query is about $1.06, under 0.1% of the estimated $1,700–$2,300 traditional annotation cost, so exploratory analyses become affordable.
  • User-defined ML functions can be dropped into the same pipeline, allowing researchers to reuse their own models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the QUIET-ML benchmark's ground truth is described as 'ground-truth query results' but the paper does not say how it was produced; if the same internally supported ML functions generated those answers, then LEAP's high scores show it can reconstruct its own function outputs, not that the outputs are semantically correct for new data.
  • Editorial inference: the vague-query filter inherits the judgment of the underlying LLM, so the 96–98% component accuracies are likely to drift as models or prompts change; a deterministic check on data sufficiency and value ranges would make the guarantees more stable.
  • Editorial inference: the same architecture could extend beyond text to images, audio, or video as long as ML functions exist that can annotate those modalities; QUIET-ML already includes videos and PDFs, so the main requirement is function coverage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces QUIET-ML, a new benchmark of 120 natural-language social science queries over unstructured data (text, PDFs, videos), together with claimed ground-truth answers. It also proposes LEAP, an LLM-powered pipeline that first uses a forward planning filter to detect vague queries and suggest disambiguated alternatives, then selects and applies ML functions (from a function tree with doubly linked dependency lists) to extend the data into an annotated table, and finally generates and executes code (mostly SQL) to answer the query. The paper reports 92% pass@1, 100% pass@3, and a cost of $1.06 per query on QUIET-ML, with ablations showing the importance of each system component.

Significance. If the performance and cost claims are supported by a sound, independent benchmark, this is a significant step toward making ML-based analysis of unstructured data accessible to social scientists. The paper's strengths are its broad benchmark coverage (9 domains, 68 sources, 33 vague queries), its modular system design, the systematic evaluation with 5 runs per query and clear pass@k definitions, the thorough ablations (forward planning filter, doubly linked lists, function tree), and the public artifact. However, the benchmark's ground-truth construction is not documented, and the evaluation metrics for vague queries are defined in a way that risks circularity. These issues must be resolved before the headline numbers can be trusted.

major comments (3)
  1. [Section 2, Section 5.2] The paper states that QUIET-ML "includes the ground-truth query results on the provided unstructured data" (Section 2) but never documents how these reference answers were produced. Given the average data size of 22,323 records and an average of 2.0 annotations per record (Section 2), manual construction at the paper's own cost estimates of $1,736–$2,266 per query (Section 5.2) would be prohibitively expensive, implying that the gold answers were generated automatically, most plausibly by the same or similar ML models that constitute LEAP's "internally supported functions." If that is the case, the reported 92% pass@1 and 100% pass@3 measure whether LEAP can replay a reference pipeline (function chain, parameters, SQL) rather than whether the answers are semantically valid. The authors must report the ground-truth generation procedure, including the models, parameters, and any manual verification; without this, the central performance claim is not independently verifiable.
  2. [Section 5.2] For the 33 vague queries, a run is scored as successful if "LEAP correctly rejects the query and recommends alternatives, one of which yields a result that matches the ground truth." Since a vague query such as Q2 ("Is the public mood correlated with...") has no unique answer until disambiguated, the ground truth itself must encode a specific disambiguation. If that disambiguation was generated by one of LEAP's internally supported ML functions (e.g., the emotion distribution), then the evaluation only verifies that LEAP can propose the particular alternative that was arbitrarily selected as gold. The paper should specify how the ground-truth disambiguation for each vague query was chosen, report the set of acceptable alternative queries, and state whether success requires the first recommendation or allows any recommendation.
  3. [Section 4.2.1] The table generation prompt includes a worked example that states the "ground-truth function selections" for the query "I want to count the number of positive paragraphs in the PDF document," including the exact function chain (OCR, paragraph separator, sentiment analyzer, stopper) and parameter selections. The paper does not state whether this exemplar is one of the 120 QUIET-ML test queries or whether the test set contains queries with the same function chain. If the exemplar is not disjoint from the test set, the prompt leaks the correct answer for those queries and inflates pass@k. Please disclose the exemplar and confirm its disjointness from the test queries, or remove it from the evaluation.
minor comments (4)
  1. [Figure 6] The label "LogicBeam" in the horizontal bar chart should be "LogicalBeam" to match the text and reference [8].
  2. [Section 5.2] The sentence "LEAP achieves an accuracy of 93.3% on the 33 vague queries when also provided with unstructured data" is confusing because the immediately preceding comparison uses structured tables; please clarify which numbers correspond to which setting in the main results.
  3. [Section 4.1] The forward planning filter's third case mentions "generating an alternative query list Q based on F, D, and q," but the success criterion in Section 5.2 requires that one of the recommended alternatives yields a result that matches the ground truth; please state explicitly whether the alternative list is scored on the annotated table or on the original unstructured data.
  4. [Section 2] The statement that "over half (61) of them require executing two or more ML models" is technically 61/120 = 50.8%; consider writing "61 (50.8%)" for precision.

Circularity Check

1 steps flagged · score 4.0 of 10

QUIET-ML's ground-truth construction is undocumented; if the gold answers were produced by the same ML functions that LEAP internally supports, the 92% pass@1 / 100% pass@3 claims measure replay fidelity rather than independent answer correctness.

  1. self definitional [Section 2 (QUIET-ML) and Section 5.2 (success criteria)]
    "For evaluation, QUIET-ML includes the ground-truth query results on the provided unstructured data. ... For vague queries, a run is successful if LEAP correctly rejects the query and recommends alternatives, one of which yields a result that matches the ground truth."

    The paper never states how QUIET-ML's gold answers were generated. The dataset's scale (22,323 data points and 2.0 annotations per query on average) and the paper's own cost estimates ($1,736-$2,266 per query for human annotation) make independent human gold for all 120 queries implausible; the remaining feasible construction is to run the same published ML models that LEAP wraps as its 'internally supported' functions (e.g., f_emotion, f_pe). Under that construction, the gold query result is, by definition, the output of the very function chain LEAP is asked to select, so pass@k reduces to whether LEAP replays the reference chain and SQL.

full rationale

LEAP's internal components (forward planning filter, stage selector, function tree, doubly linked lists, alias checks) are described with their own ablations and are not derived from the benchmark labels, so their engineering logic is self-contained. There is no load-bearing self-citation or imported uniqueness theorem. The central circularity risk is external to the derivation chain: QUIET-ML's 'ground-truth query results' are asserted but their provenance is never documented, and the reported data size and human-annotation cost comparison imply automated gold generation via the same ML functions that constitute LEAP's supported function list. If that is what was done, the headline pass@k numbers are a self-consistency/replay test, not a validation of answer semantics; if instead the gold tables were independently human-verified or taken from the original papers' published results, the claim would be independent and the score would be near 0. Because the paper does not disclose this, the honest finding is a partial evaluation-circularity risk rather than a demonstrated derivation collapse.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No mathematical or physical free parameters are introduced; the system's behavior is governed by LLM prompts and pre-trained ML functions. The main unstated premises are about dataset validity and ML function correctness.

assumptions (4)
  • domain assumption Ground truth answers in QUIET-ML are correct and independent of LEAP's components.
    Section 2 states ground-truth query results are included but does not describe how they were produced; evaluation validity depends on this.
  • domain assumption The 120 collected queries are representative of real social science research questions.
    Section 2 says they cover SALT survey and CS 224C topics, but selection is not randomized and may emphasize queries solvable with available ML functions.
  • domain assumption The internally supported ML functions yield semantically valid annotations.
    LEAP's answers inherit the accuracy of the ML functions; the paper does not evaluate the accuracy of these functions on this data.
  • domain assumption gpt-4-0613 reliably performs function calling and code generation as prompted.
    All LLM steps use gpt-4-0613; if API behavior changes, results may not reproduce.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data." pith.science (2026). https://pith.science/paper/6PAVMOPZ

@misc{pith2026250103892,
  author       = {Pith},
  title        = {Pith review of: LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6PAVMOPZ}},
  note         = {Machine review of arXiv:2501.03892}
}
abstract

Social scientists are increasingly interested in analyzing the semantic information (e.g., emotion) of unstructured data (e.g., Tweets), where the semantic information is not natively present. Performing this analysis in a cost-efficient manner requires using machine learning (ML) models to extract the semantic information and subsequently analyze the now structured data. However, this process remains challenging for domain experts. To demonstrate the challenges in social science analytics, we collect a dataset, QUIET-ML, of 120 real-world social science queries in natural language and their ground truth answers. Existing systems struggle with these queries since (1) they require selecting and applying ML models, and (2) more than a quarter of these queries are vague, making standard tools like natural language to SQL systems unsuited. To address these issues, we develop LEAP, an end-to-end library that answers social science queries in natural language with ML. LEAP filters vague queries to ensure that the answers are deterministic and selects from internally supported and user-defined ML functions to extend the unstructured data to structured tables with necessary annotations. LEAP further generates and executes code to respond to these natural language queries. LEAP achieves a 100% pass @ 3 and 92% pass @ 1 on QUIET-ML, with a \$1.06 average end-to-end cost, of which code generation costs \$0.02.

Figures

Figures reproduced from arXiv: 2501.03892 by the authors.

Figure 1
Figure 1. Workflow of LEAP. With the recommended non-vague query, the library first loads the Tweets as a single-column table and appends an additional column containing emotion classes from applying the emotion classifier 𝑓emotion on the posts. The library then generates code similar to the following SQL code. SELECT emotion , COUNT ( emotion ) AS count FROM table GROUP BY emotion Finally, the library executes the code and d… view at source ↗
Figure 3
Figure 3. Structure of the supported function list [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. LEAP table generation stage’s display for 𝑄3: “For each (targeted) persona/in-group, I want to know the number of each type of dog whistle.” [75] LEAP displays the current stage during its execution to indicate the progress. During the table generation stage, LEAP dynamically updates the column mapping relation graph ( [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (2 more)
Figure 6
Figure 6. Figure 6: Pass @ 1 of LEAP and baselines over the 33 vague queries in QUIET-ML on structured tables [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Cost breakdown of LEAP and comparison with traditional social science research methods. extends beyond the table schema level and a limited number of text segments, where LEAP demonstrates a significant advantage. LEAP is cost-efficient. We compare the query cost of LE…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

139 extracted references · 20 canonical work pages

  1. [1]

    Faisal Alatawi, Paras Sheth, and Huan Liu. 2023. Quantifying the Echo Chamber Effect: An Embedding Distance-Based Approach. arXiv:2307.04668 [cs]

  2. [2]

    Gauss Algorithmic. 2023. https://huggingface.co/gaussalgo/T5-LM-Large- text2sql-spider Accessed: February 25, 2024

  3. [3]

    Tim Althoff, Cristian Danescu-Niculescu-Mizil, and Daniel Jurafsky. 2014. How to Ask for a Favor: A Case Study on the Success of Altruistic Requests. ArXiv abs/1405.3282 (2014). https://api.semanticscholar.org/CorpusID:8809599

  4. [4]

    Akari Asai, Sara Evensen, Behzad Golshan, Alon Halevy, Vivian Li, Andrei Lopatenko, Daniela Stepanov, Yoshihiko Suhara, Wang-Chiew Tan, and Yinzhan Xu. 2018. HappyDB: A Corpus of 100,000 Crowdsourced Happy Moments. arXiv:1801.07746 [cs.CL]

  5. [5]

    Mana Ashida and Mamoru Komachi. 2022. Towards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions. In Pro- ceedings of the Sixth Workshop on Online Abuse and Harms (WOAH) . Asso- ciation for Computational Linguistics, Seattle, Washington (Hybrid), 11–23. https://aclanthology.org/2022.woah-1.2

  6. [6]

    Ramy Baly, Giovanni Da San Martino, James Glass, and Preslav Nakov. 2020. We Can Detect Your Bias: Predicting the Political Ideology of News Articles. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Onlin...

  7. [7]

    David Bamman, Brendan O’Connor, and Noah A. Smith. 2013. Learning La- tent Personas of Film Characters. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Hinrich Schuetze, Pascale Fung, and Massimo Poesio (Eds.). Association for Computa- tional Linguistics, Sofia, Bulgaria, 352–361. https:...

  8. [8]

    Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe, and Sunita Sarawagi. 2023. Benchmarking and Improving Text-to-SQL Generation under Ambiguity. arXiv:2310.13659 [cs.CL]

Show all 139 references
  1. [9]

    Johan Bollen, Huina Mao, and Xiaojun Zeng. 2011. Twitter mood predicts the stock market. Journal of Computational Science 2, 1 (2011), 1–8. https: //doi.org/10.1016/j.jocs.2010.12.007

  2. [10]

    Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to Computer Programmer as Woman is to Home- maker? Debiasing Word Embeddings. In Advances in Neural Information Pro- cessing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and...

  3. [11]

    Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov. 2019. Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Intern...

  4. [12]

    Levinson

    Penelope Brown and Stephen C. Levinson. 1987. Politeness: Some universals in language usage. Vol. 4. Cambridge University Press

  5. [13]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  6. [14]

    Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, and Smaranda Muresan

  7. [15]

    Branden Chan, Timo Möller, Malte Pietsch, and Tanay Soni. 2023. https: //huggingface.co/deepset/roberta-base-squad2 Accessed: June 23, 2024

  8. [16]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...

  9. [17]

    Romero, and David Jurgens

    Minje Choi, Ceren Budak, Daniel M. Romero, and David Jurgens. 2021. More than Meets the Tie: Examining the Role of Interpersonal Relationships in Social Networks. arXiv:2105.06038 [cs.SI]

  10. [18]

    Eric Chu, Akanksha Baid, Ting Chen, AnHai Doan, and Jeffrey Naughton. 2007. A relational approach to incrementally extracting and querying structure in unstructured data. In Proceedings of the 33rd international conference on Very large data bases. 1045–1056

  11. [19]

    Eric Chu, Prashanth Vijayaraghavan, and Deb Roy. 2018. Learning Personas from Dialogue with Attentive Memory Networks. arXiv:1810.08717 [cs.CL]

  12. [20]

    Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini

  13. [21]

    Yi-Ling Chung, Serra Sinem Tekiroğlu, and Marco Guerini. 2021. Towards Knowledge-Grounded Counter Narrative Generation for Hate Speech. In Pro- ceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics. Association for Computational Linguistics, Online

  14. [22]

    Ann Copestake and Karen Sparck Jones. 1990. Natural language interfaces to databases. The Knowledge Engineering Review 5, 4 (1990), 225–249

  15. [23]

    Cristian Danescu-Niculescu-Mizil, Lillian Lee, Bo Pang, and Jon Kleinberg. 2012. Echoes of power: Language effects and power differences in social interaction. In Proceedings of WWW . 699–708

  16. [24]

    Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Dan Jurafsky, Jure Leskovec, and Christopher Potts. 2013. A computational approach to politeness with application to social factors. In Proceedings of ACL

  17. [25]

    Richard D De Veaux and Adam Eck. 2021. Machine Learning Methods for Computational Social Science.Handbook of Computational Social Science, Volume 2: Data Science, Statistical Modelling, and Machine Learning Methods (2021)

  18. [26]

    Clark, Vinodkumar Prab- hakaran, and Jacob Eisenstein

    Dorottya Demszky, Devyani Sharma, Jonathan H. Clark, Vinodkumar Prab- hakaran, and Jacob Eisenstein. 2021. Learning to Recognize Dialect Features. arXiv:2010.12707 [cs.CL]

  19. [27]

    Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson. 2021. Structure-Grounded Pretraining for Text-to-SQL. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics:...

  20. [28]

    Achim Edelmann, Tom Wolff, Danielle Montagne, and Christopher A. Bail

  21. [29]

    Julian Martin Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Under- standing tables with intermediate pre-training. arXiv:2010.00571 [cs.CL]

  22. [30]

    Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent Hatred: A Benchmark for Understanding Implicit Hate Speech. arXiv:2109.05322 [cs.CL]

  23. [31]

    Nora Freeman Engstrom, David Freeman Engstrom, Jonah Gelbach, Austin Peters, and Aaron Schaffer-Neitz. 2024. Secrecy by Stipulation. Duke Law Journal (May 2 2024). https://papers.ssrn.com/sol3/papers.cfm?abstract_id= 4811151

  24. [32]

    Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroğlu, and Marco Guerini

  25. [33]

    Emilio Ferrara and Zeyao Yang. 2015. Measuring Emotional Contagion in Social Media. PLOS ONE 10 (06 2015). https://doi.org/10.1371/journal.pone.0142390

  26. [34]

    Avrilia Floratou, Fotis Psallidas, Fuheng Zhao, and et al. [n.d.]. NL2SQL is a solved problem... Not! https://www.cidrdb.org/cidr2024/papers/p74-floratou. pdf

  27. [35]

    Han Fu, Chang Liu, Bin Wu, Feifei Li, Jian Tan, and Jianling Sun. 2023. CatSQL: Towards Real World Natural Language to SQL Applications. Proceedings of the VLDB Endowment 16, 6 (2023), 1534–1547

  28. [36]

    Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roes- ner, Eunsol Choi, and Yejin Choi. 2022. Misinfo Reaction Frames: Reason- ing about Readers’ Reactions to News Headlines. In Proceedings of the 60th Annual Meeting of the Association for Computational Li...

  29. [37]

    Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, and Yejin Choi. 2022. Misinfo Reaction Frames: Reasoning about Readers’ Reactions to News Headlines. ACL (2022)

  30. [38]

    Dawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun, Yichen Qian, Bolin Ding, and Jingren Zhou. 2023. Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation. CoRR abs/2308.15363 (2023)

  31. [39]

    Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences 115, 16 (April 2018). https://doi.org/10.1073/ pnas.1720347115

  32. [40]

    Google. 2023. https://huggingface.co/google/tapas-large-finetuned-wtq Ac- cessed: February 25, 2024

  33. [41]

    Roberts, and Brandon M

    Justin Grimmer, Margaret E. Roberts, and Brandon M. Stewart. 2021. Machine Learning for Social Science: An Agnostic Approach. Annual Review of Political Science 24, 1 (2021), 395–419. https://doi.org/10.1146/annurev-polisci-053119- 015921 arXiv:https://doi.org/10.1146/annurev-...

  34. [42]

    Matthew Groh, Ziv Epstein, Chaz Firestone, and Rosalind Picard. 2021. Deep- fake detection by human crowds, machines, and machine-informed crowds. Proceedings of the National Academy of Sciences 119, 1 (Dec. 2021). https: //doi.org/10.1073/pnas.2110013119

  35. [43]

    Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. 2019. Towards Complex Text-to-SQL in Cross-Domain Data- base with Intermediate Representation. arXiv:1905.08205 [cs.CL]

  36. [44]

    Hatfield, J.T

    E. Hatfield, J.T. Cacioppo, and R.L. Rapson. 1993. Emotional Contagion. Cam- bridge University Press. https://books.google.com/books?id=BbA-BAAAQBAJ

  37. [45]

    Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020. TAPAS: Weakly Supervised Table Parsing via Pre-training. arXiv:2004.02349 [cs.IR]

  38. [46]

    Christopher Hidey, Elena Musi, Alyssa Hwang, Smaranda Muresan, and Kathy McKeown. 2017. Analyzing the Semantic Types of Claims and Premises in an Online Persuasive Forum. In Proceedings of the 4th Workshop on Argument Mining, Ivan Habernal, Iryna Gurevych, Kevin Ashley, Claire...

  39. [47]

    Chuxuan Hu, Qinghai Zhou, and Hanghang Tong. 2024. Genius: Subteam Replacement with Clustering-based Graph Neural Networks. In Proceedings of the 2024 SIAM International Conference on Data Mining (SDM) . 10–18. https: //doi.org/10.1137/1.9781611978032.2

  40. [48]

    Pere-Lluís Huguet Cabot, Verna Dankers, David Abadi, Agneta Fischer, and Ekaterina Shutova. 2020. The Pragmatics behind Politics: Modelling Metaphor, Framing and Emotion in Political Discourse. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor C...

  41. [49]

    Ali Hürriyetoğlu, Hristo Tanev, Vanni Zavarella, Jakub Piskorski, Reyyan Yen- iterzi, Osman Mutlu, Deniz Yuret, and Aline Villavicencio. 2021. Challenges and Applications of Automated Extraction of Socio-political Events from Text (CASE 2021): Workshop and Shared Task Report. ...

  42. [50]

    Mohit Iyyer, Peter Enns, Jordan Boyd-Graber, and Philip Resnik. 2014. Political Ideology Detection Using Recursive Neural Networks. InProceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Kristina Toutanova and Hua Wu ...

  43. [51]

    Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based Neu- ral Structured Learning for Sequential Question Answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan...

  44. [52]

    Jacobs and Annette Kinder

    Arthur M. Jacobs and Annette Kinder. 2018. What Makes a Metaphor Literary? Answers From Two Computational Studies. Metaphor and Symbol 33, 2 (2018), 85–100. https://doi.org/10.1080/10926488.2018.1434943

  45. [53]

    Zubin Jelveh, Bruce Kogut, and Suresh Naidu. 2014. Detecting Latent Ideology in Expert Text: Evidence From Academic Papers in Economics. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Alessandro Moschitti, Bo Pang, and Walter...

  46. [54]

    Yohan Jo, Seojin Bang, Emaad Manzoor, Eduard Hovy, and Chris Reed. 2020. De- tecting Attackable Sentences in Arguments. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.)....

  47. [55]

    Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia

  48. [56]

    Bailis, Tatsunori Hashimoto, and Matei Zaharia

    Daniel Kang, John Guibas, Peter D. Bailis, Tatsunori Hashimoto, and Matei Zaharia. 2022. TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data. In Proceedings of the 2022 International Conference on Management of Data (Philadelphia, PA, USA) (SIGMOD...

  49. [57]

    George Katsogiannis-Meimarakis and Georgia Koutrika. 2023. A survey on deep learning approaches for text-to-SQL. The VLDB Journal 32, 4 (2023), 905–936. https://doi.org/10.1007/s00778-022-00776-8

  50. [58]

    Adam Kramer, Jamie Guillory, and Jeffrey Hancock. 2014. Experimental Evi- dence of Massive-Scale Emotional Contagion Through Social Networks. Pro- ceedings of the National Academy of Sciences of the United States of America 111 (06 2014). https://doi.org/10.1073/pnas.1320040111

  51. [59]

    Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy Liang. 2019. SPoC: Search-based Pseudocode to Code. arXiv:1906.04908 [cs.LG]

  52. [60]

    David Lazer, Alex Pentland, Lada Adamic, Sinan Aral, Albert-László Barabási, Devon Brewer, Nicholas Christakis, Noshir Contractor, James Fowler, Myron Gutmann, et al. 2009. Computational social science. Science 323, 5915 (2009), 721–723

  53. [61]

    David MJ Lazer, Alex Pentland, Duncan J Watts, Sinan Aral, Susan Athey, Noshir Contractor, Deen Freelon, Sandra Gonzalez-Bailon, Gary King, Helen Margetts, et al. 2020. Computational social science: Obstacles and opportunities. Science 369, 6507 (2020), 1060–1062

  54. [62]

    David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berin- sky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Ny- han, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jon...

  55. [63]

    Jens Lemmens, Ilia Markov, and Walter Daelemans. 2021. Improving Hate Speech Type and Target Detection with Hateful Metaphor Features. In Pro- ceedings of the Fourth Workshop on NLP for Internet Freedom: Censorship, Disin- formation, and Propaganda, Anna Feldman, Giovanni Da S...

  56. [64]

    Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al. 2024. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. Advances in Neural Information Processing Sys...

  57. [65]

    Mingyang Li, Louis Hickman, Louis Tay, Lyle Ungar, and Sharath Chandra Guntuku. 2020. Studying Politeness across Cultures using English Twitter and Mandarin Weibo. Proc. ACM Hum.-Comput. Interact. 4, CSCW2, Article 119 (oct 2020), 15 pages. https://doi.org/10.1145/3415190

  58. [66]

    Sha Li, Heng Ji, and Jiawei Han. 2021. Document-Level Event Argument Ex- traction by Conditional Generation. arXiv:2104.05919 [cs.CL]

  59. [67]

    Zihao Li, Dongqi Fu, and Jingrui He. 2023. Everything Evolves in Personalized PageRank. In Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023 , Ying Ding, Jie Tang, Juan F. Sequeda, Lora Aroyo, Carlos Castillo, and Geert-Jan Hoube...

  60. [68]

    Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. TAPEX: Table Pre-training via Learning a Neural SQL Executor. arXiv:2107.07653 [cs.CL] https://arxiv.org/abs/2107.07653

  61. [69]

    Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021. Towards Emotional Support Dialog Systems. arXiv:2106.01144 [cs.CL]

  62. [70]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov

  63. [71]

    Shuai Ma, Wenbin Jiang, Xiang Ao, Meng Tian, Xinwei Feng, Yajuan Lyu, Qiaoqiao She, and Qing He. 2023. Semantic-Driven Instance Generation for Table Question Answering. In International Conference on Database Systems for Advanced Applications. Springer, 3–18

  64. [72]

    Binny Mathew, Anurag Illendula, Punyajoy Saha, Soumya Sarkar, Pawan Goyal, and Animesh Mukherjee. 2020. Hate begets Hate: A Temporal Study of Hate Speech. arXiv:1909.10966 [cs.SI]

  65. [73]

    Daniel A McFarland, Kevin Lewis, and Amir Goldberg. 2016. Sociology in the era of big data: The ascent of forensic social science. The American Sociologist 47 (2016), 12–35

  66. [74]

    McRae and J

    K. McRae and J. J. Gross. 2020. Emotion regulation. Emotion 20, 1 (2020), 1–9

  67. [75]

    Julia Mendelsohn, Ronan Le Bras, Yejin Choi, and Maarten Sap. 2023. From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, J...

  68. [76]

    CoRR abs/1907.11692 (2019)

    RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907.11692

  69. [77]

    Vlad Niculae and Cristian Danescu-Niculescu-Mizil. 2014. Brighter than Gold: Figurative Language in User Generated Comparisons. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Alessandro Moschitti, Bo Pang, and Walter Daelema...

  70. [78]

    OpenAI. 2023. https://chat.openai.com/ Accessed: October 24, 2024

  71. [79]

    OpenAI. 2023. https://platform.openai.com/docs/models/gpt-4-and-gpt-4- turbo Accessed: October 24, 2024

  72. [80]

    OpenAI. 2023. https://platform.openai.com/docs/guides/function-calling/ supported-models Accessed: October 24, 2024

  73. [81]

    OpenAI. 2023. https://platform.openai.com/examples/default-sql-translate Accessed: October 24, 2024

  74. [82]

    Saif Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry. 2016. SemEval-2016 Task 6: Detecting Stance in Tweets. In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval- 2016), Steven Bethard, Marine Carpuat, Daniel Cer, Dav...

  75. [83]

    Panupong Pasupat and Percy Liang. 2015. Compositional Semantic Parsing on Semi-Structured Tables. CoRR abs/1508.00305 (2015). arXiv:1508.00305 http://arxiv.org/abs/1508.00305

  76. [84]

    Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, and Mihai Burzo

  77. [85]

    Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019. WiC: the Word- in-Context Dataset for Evaluating Context-Sensitive Meaning Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...

  78. [86]

    Mohammadreza Pourreza and Davood Rafiei. 2024. Din-sql: Decomposed in-context learning of text-to-sql with self-correction. Advances in Neural Information Processing Systems 36 (2024)

  79. [87]

    Marcelo O. R. Prates, Pedro H. C. Avelar, and Luis Lamb. 2019. Assessing Gender Bias in Machine Translation – A Case Study with Google Translate. arXiv:1809.02208 [cs.CY] https://arxiv.org/abs/1809.02208

  80. [88]

    OpenAI. 2024. https://openai.com/pricing

  81. [89]

    Bowen Qin, Binyuan Hui, Lihan Wang, Min Yang, Jinyang Li, Binhua Li, Ruiying Geng, Rongyu Cao, Jian Sun, Luo Si, et al. 2022. A survey on text-to-sql parsing: Concepts, methods, and future directions. arXiv preprint arXiv:2208.13629 (2022)

  82. [90]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 [cs.LG]

  83. [91]

    Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. arXiv:1806.03822 [cs.CL] https: //arxiv.org/abs/1806.03822

  84. [92]

    Manuel Romero. 2023. https://huggingface.co/mrm8488/t5-base-finetuned- wikiSQL Accessed: February 25, 2024

  85. [93]

    Barbara Rothbaum, Elizabeth Meadows, and Patricia Resick. 2012. Cognitive- Behavioral Therapy. Journal of Traumatic Stress 13 (10 2012)

  86. [94]

    Barbara O Rothbaum, Elizabeth A Meadows, Patricia Resick, and David W Foy

  87. [95]

    Daniel Preoţiuc-Pietro, Ye Liu, Daniel Hopkins, and Lyle Ungar. 2017. Beyond Binary Labels: Political Ideology Prediction of Twitter Users. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Regina Barzilay and M...

  88. [96]

    Smith, and Yejin Choi

    Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Impli- cations of Language. arXiv:1911.03891 [cs.CL]

  89. [97]

    Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In ACL

  90. [98]

    Smith, and James Pennebaker

    Maarten Sap, Eric Horvitz, Yejin Choi, Noah A. Smith, and James Pennebaker

  91. [99]

    Maarten Sap, Anna Jafarpour, Choi Yejin, Noah Smith, James Pennebaker, and Eric Horvitz. 2022. Quantifying the narrative flow of imagined ver- sus autobiographical stories. Proceedings of the National Academy of Sci- ences of the United States of America 119 (11 2022), e221171...

  92. [100]

    Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen

  93. [101]

    Scale. [n.d.]. https://scale.com/rapid Accessed: October 24, 2024

  94. [102]

    Harvard Law School. 2023. https://case.law/ Accessed: October 24, 2024

  95. [103]

    Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The Risk of Racial Bias in Hate Speech Detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korho- nen, David Traum, and Lluís Màrquez (Eds.)....

  96. [104]

    Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. 2020. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? arXiv preprint arXiv:2010.12725 (2020)

  97. [105]

    Rachele Sprugnoli and Sara Tonelli. 2019. Novel Event Detection and Classifica- tion for Historical Texts. Computational Linguistics 45, 2 (June 2019), 229–265. https://doi.org/10.1162/coli_a_00347

  98. [106]

    Kevin Stowe, Prasetya Utama, and Iryna Gurevych. 2022. IMPLI: Investigating NLI Models’ Performance on Figurative Language. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and ...

  99. [107]

    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)

    Recollection versus Imagination: Exploring Human Memory and Cogni- tion via Neural Language Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association f...

  100. [108]

    Chenhao Tan, Vlad Niculae, Cristian Danescu-Niculescu-Mizil, and Lillian Lee

  101. [109]

    Anja Thieme, Danielle Belgrave, and Gavin Doherty. 2020. Machine Learning in Mental Health: A Systematic Review of the HCI Literature to Support the Development of Effective and Implementable ML Systems.ACM Trans. Comput.- Hum. Interact. 27, 5, Article 34 (aug 2020), 53 pages....

  102. [110]

    Toma, Jeffrey T

    Catalina L. Toma, Jeffrey T. Hancock, and Nicole B. Ellison. 2008. Sepa- rating Fact From Fiction: An Examination of Deceptive Self-Presentation in Online Dating Profiles. Personality and Social Psychology Bul- letin 34, 8 (2008), 1023–1036. https://doi.org/10.1177/01461672083...

  103. [111]

    Bing Wang, Yan Gao, Zhoujun Li, and Jian-Guang Lou. 2023. Know What I don’t Know: Handling Ambiguous and Unanswerable Questions for Text-to- SQL. arXiv:2212.08902 [cs.CL]

  104. [112]

    Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2021. RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers. arXiv:1911.04942 [cs.CL]

  105. [113]

    Miner, David C

    Ashish Sharma, Adam S. Miner, David C. Atkins, and Tim Althoff. 2020. A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support. arXiv:2009.08441 [cs.CL]

  106. [114]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903 [cs.CL]

  107. [115]

    Orion Weller and Kevin Seppi. 2019. Humor Detection: A Transformer Gets the Last Laugh. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro I...

  108. [116]

    Diyi Yang, Jiaao Chen, Zichao Yang, Dan Jurafsky, and Eduard Hovy. 2019. Let’s Make Your Request More Persuasive: Modeling Persuasive Strategies via Semi-Supervised Neural Nets on Crowdfunding Platforms. In Proceedings of the 2019 Conference of the North American Chapter of th...

  109. [117]

    Alane Suhr, Ming-Wei Chang, Peter Shaw, and Kenton Lee. 2020. Exploring Unexplored Generalization Challenges for Cross-Database Semantic Parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schlu...

  110. [118]

    Tao Yu, Zifan Li, Zilin Zhang, Rui Zhang, and Dragomir Radev. 2018. TypeSQL: Knowledge-based Type-Aware Neural Text-to-SQL Generation. arXiv:1804.09769 [cs.CL]

  111. [119]

    Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In P...

  112. [120]

    Hongli Zhan, Tiberiu Sosea, Cornelia Caragea, and Junyi Jessy Li. 2022. Why Do You Feel This Way? Summarizing Triggers of Emotions in Social Media Posts. ArXiv abs/2210.12531 (2022). https://api.semanticscholar.org/CorpusID: 253098848

  113. [121]

    Zhang, Bryan Culbertson, and Praveen K

    Amy X. Zhang, Bryan Culbertson, and Praveen K. Paritosh. 2017. Char- acterizing Online Discussion Using Coarse Discourse Sequences. Proceed- ings of the International AAAI Conference on Web and Social Media (2017). https://api.semanticscholar.org/CorpusID:35696952

  114. [122]

    Justine Zhang, Jonathan Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Dario Taraborelli, and Nithum Thain. 2018. Conversations Gone Awry: Detecting Early Signs of Conversational Failure. In Proceedings of the 56th Annual Meeting of the Association for Compu...

  115. [123]

    Jun Zhang, Wei Wang, Feng Xia, Yu-Ru Lin, and Hanghang Tong. 2020. Data- Driven Computational Social Science: A Survey. Big Data Research 21 (2020), 100145. https://doi.org/10.1016/j.bdr.2020.100145

  116. [124]

    Chuangxian Wei, Bin Wu, Sheng Wang, Renjie Lou, Chaoqun Zhan, Feifei Li, and Yuanzhe Cai. 2020. AnalyticDB-V: a hybrid analytical engine towards query fusion for structured and unstructured data. Proc. VLDB Endow. 13, 12 (aug 2020), 3152–3165. https://doi.org/10.14778/3415478.3415541

  117. [125]

    Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.CoRR abs/1709.00103 (2017)

  118. [126]

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023. Can Large Language Models Transform Computational Social Science? arXiv:2305.03514 [cs.CL]

  119. [127]

    Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. Multi-VALUE: A Framework for Cross-Dialectal English NLP. arXiv:2212.08011 [cs.CL]

  120. [128]

    Diyi Yang, Kaitlyn Zhou, and Wenna Qin. 2023. CS 224C: NLP for Computational Social Science. Stanford University. https://web.stanford.edu/class/cs224c/

  121. [135]

    Yi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten, Georgia Koutrika, and Kurt Stockinger. 2023. ScienceBenchmark: A Com- plex Real-World Benchmark for Evaluating Natural Language to SQL Systems. arXiv:2306.04743 [cs.DB]

  122. [139]

    Caleb Ziems, Minzhi Li, Anthony Zhang, and Diyi Yang. 2022. Inducing Positive Perspectives with Text Reframing. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), Smaranda Muresan, Preslav Nakov, and Aline Vill...

  123. [2000]

    In Effective treatments for PTSD: Practice guidelines from the International Society for Traumatic Stress Studies , Edna B Foa, Terence M Keane, and Matthew J Friedman (Eds.)

    Cognitive-behavioral therapy. In Effective treatments for PTSD: Practice guidelines from the International Society for Traumatic Stress Studies , Edna B Foa, Terence M Keane, and Matthew J Friedman (Eds.). The Guilford Press, 320–325

  124. [2015]

    In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction(Seattle, Washington, USA) (ICMI ’15)

    Deception Detection using Real-life Trial Data. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction(Seattle, Washington, USA) (ICMI ’15). Association for Computing Machinery, New York, NY, USA, 59–66. https://doi.org/10.1145/2818346.2820758

  125. [2016]

    In Proceedings of the 25th International Confer- ence on World Wide Web (Montréal, Québec, Canada) (WWW ’16)

    Winning Arguments: Interaction Dynamics and Persuasion Strategies in Good-faith Online Discussions. In Proceedings of the 25th International Confer- ence on World Wide Web (Montréal, Québec, Canada) (WWW ’16). International World Wide Web Conferences Steering Committee, Republ...

  126. [2017]

    arXiv:1703.02529 [cs.DB]

    NoScope: Optimizing Neural Network Queries over Video at Scale. arXiv:1703.02529 [cs.DB]

  127. [2018]

    In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.)

    CARER: Contextualized Affect Representations for Emotion Recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics...

  128. [2019]

    In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

    CONAN - COunter NArratives through Nichesourcing: a Multilingual Dataset of Responses to Fight Online Hate Speech. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, Italy, 2819–2829...

  129. [2020]

    Annual Review of Sociol- ogy 46, 1 (2020), 61–81

    Computational Social Science and Sociology. Annual Review of Sociol- ogy 46, 1 (2020), 61–81. https://doi.org/10.1146/annurev-soc-121919-054621 arXiv:https://doi.org/10.1146/annurev-soc-121919-054621

  130. [2021]

    In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics

    Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate Speech. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics

  131. [2022]

    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.)

    FLUTE: Figurative Language Understanding through Textual Explanations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, Unite...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.