Pith. sign in

REVIEW 4 major objections 4 minor 299 references

FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read FindMyText claims that grouping position-offset-aligned fingerprint matches into chains detects text containment at ~0.99 AUC-ROC across Wikipedia, ArXiv, and a 50.7M-document web crawl, at ~450 ms per query.

desk verdict The chain-clustering score is a clean, working idea that beats its baselines on its own benchmark, but the benchmark never gives a positive a gap inside the copied fragment—so the robustness claims for real crawled corpora are overreaching. read the letter →

arxiv 2607.10020 v2 pith:M3S3TIBB submitted 2026-07-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords textcontainmentdocumentfingerprintingwinnowingnear-duplicatedetectioninvertedindexpositionalclusteringcopyrightweb-scalesearch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents FindMyText, a method to decide whether a query text appears—in whole or in part—inside a large text corpus, a task that matters for verifying whether copyrighted material made it into an LLM's training data. The central claim is that this 'text containment' can be detected far more reliably by looking for chains of matching fingerprints than by counting shared fingerprints or comparing embeddings. Each shared fingerprint is placed at a point (position in the query, offset relative to the corpus document); connected clusters of such points indicate a contiguous reused passage, and the size of the largest cluster is the containment score. On a synthetic benchmark built with a Smith-Waterman oracle, this chain score achieves AUC-ROC around 0.99 on Wikipedia, ArXiv, and a 50.7-million-document web crawl, while semantic-similarity baselines perform near chance. The method runs in about 450 ms per query on the largest corpus using a disk-based inverted index.

What carries the argument

The central mechanism is the fingerprint-chain score s_chain(q,d_i)=max_k |C_k|, built from winnowed, position-stamped hashes. Winnowing selects the minimum hash in each sliding window, guaranteeing that any shared substring of length k+w−1 yields a common fingerprint. The novelty is clustering shared fingerprints in (query position, offset) space with tolerance thresholds (τ_pos, τ_off), which converts a set of independent coincidental hashes into evidence of a contiguous, aligned copy.

What would settle it

Create a synthetic or real example where a query shares two consecutive copied blocks with an inserted 100-character paragraph between them (so Smith-Waterman still scores above the positive threshold but no single fingerprint chain connects the blocks). If FindMyText's chain score for this pair falls below its positive threshold, the method's robustness claim is contradicted.

Watch

Extended reading notes

Core claim

The core discovery is that an explicit, geometrically coherent sequence of shared fingerprints—rather than a raw count—separates true near-verbatim reuse from mere topical similarity. For a query q and a corpus document d_i, FindMyText winnows both texts into k-gram hashes, records character-level positions, and for each shared hash h computes the offset δ(h) = p_{d_i}(h) − p_q(h). Shared hashes are plotted as points (p_q(h), δ(h)); a reused passage appears as a nearly horizontal run of points. Two points are connected if their query positions differ by at most τ_pos = 30 characters and their offsets differ by at most τ_off = 10 characters; clusters of at least κ = 5 fingerprints are kept, a

Load-bearing premise

The reported accuracy transfers to real corpora only if the synthetic benchmark's edit menu and Smith-Waterman thresholds faithfully represent the noise actually introduced by web crawling and cleaning.

Editorial extensions

If this is right

  • If the chain score is right, near-verbatim reuse can be detected at web scale with a disk-based inverted index: ~450 ms per query on 50.7M documents.
  • The method distinguishes containment from similarity: BM25 and dense retrieval score near chance on the benchmark, while the chain method stays above 0.98 AUC-ROC.
  • Robustness covers common cleaning edits—case changes, hyphenation, ligatures, boilerplate insertion—but not paraphrase or large reorderings.
  • The approach yields localizable evidence (which passages match), which is useful for copyright-compliance audits of training corpora.
  • The synthetic benchmark itself is a contribution: a reproducible recipe for evaluating containment methods with Smith-Waterman as oracle.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The (position, offset) clustering is essentially a Hough-transform–style vote for a linear alignment; a natural extension is to return multiple disjoint chains and their boundaries, giving a full containment map rather than a single max score.
  • The tight tolerances (τ_pos=30, τ_off=10 characters) suggest the method is tuned for clean text; OCR-laden or heavily segmented web content with larger local insertions might need adaptive thresholds or multi-scale winnowing—a testable hypothesis.
  • Because negatives in the benchmark are constructed adversarially (permute, paraphrase, boilerplate), the reported AUC may understate real-world ease if actual non-contained documents are less similar; conversely, if real copies undergo heavier edits, performance could degrade—this asymmetry is worth probing.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper presents FindMyText, an open-source Python toolkit for detecting whether a query text appears wholly or partially in a corpus. The method extracts winnowed character-level fingerprints, indexes them in a disk-based inverted index, and scores each candidate document by the size of the largest position-coherent 'chain' of shared fingerprints (Eq. 2). This chain score is compared with shared-fingerprint counts, BM25, and dense retrieval on a synthetic benchmark built from Wikipedia, ArXiv, and the HPLT web corpus. The central claim is that the chain-based method discriminates positive from negative instances with AUC-ROC around 0.99 and Precision@Recall=90% above 0.98 on all three corpora, while baselines fail (Table 1). The authors also provide a web interface and release the toolkit under an MIT license.

Significance. If the reported results transfer to real-world settings, FindMyText would be a practically valuable tool for copyright-compliance screening and training-data provenance analysis, because it scales to tens of millions of documents with sub-second queries and explicitly identifies contiguous reused passages rather than merely scoring overall similarity. The paper ships an open-source implementation and a reproducible synthetic benchmark, and the core chain-clustering idea is a natural, well-motivated extension of winnowing. The main limitation is that the evaluation's ground truth is generated by a synthetic procedure whose positive instances are, by construction, a single lightly edited contiguous fragment. Therefore the headline numbers demonstrate that the method detects exactly the kind of near-verbatim segment the benchmark instantiates, but they do not yet establish robustness to the full range of preprocessing transformations that the paper motivates in Sec. 1 (OCR noise, boilerplate insertion/removal, segmentation, and reordering).

major comments (4)
  1. [§5.1, Eq. (2)] The benchmark's positive-instance construction is the load-bearing element for the headline AUC≈0.99, and it excludes exactly the chain-breaking cases that matter for real corpora. A positive is created by taking one consecutive fragment, applying small character-level edits until the Smith-Waterman score exceeds T_pos=1000, and concatenating unrelated filler around it. Because the chain score connects fingerprints only when gaps are ≤ τ_pos=30 characters and the offset drift is ≤ τ_off=10 characters, a single inserted paragraph, a deleted sentence block, or an OCR-induced gap inside the copied segment will split the chain and can push the largest cluster below κ=5. The benchmark contains no positive instance with such discontinuities, so the reported P@R=90% of 0.998–1.00 does not demonstrate robustness to the very transformations (boilerplate removal, document segmentation, cleaning) d
  2. [Table 1, Dense Retrieval row] An AUC-ROC of exactly 0.000 for dense retrieval on Wikipedia is a degenerate result that cannot support the paper's conclusion that the baselines 'are unable to disentangle' positives and negatives. AUC=0 means every positive is ranked below every negative, which is not a typical failure mode of a document-embedding retriever; it indicates that the baseline is being applied in a way that is systematically inverted for this benchmark (e.g., using whole-document similarity when the benchmark negatives are topic-matched near-paraphrases and the positives are a short copied span embedded in unrelated filler). Since the dense-retriever results are unavailable for ArXiv and HPLT, the comparison is also incomplete. The authors should either configure the dense baseline to operate on passage-level or windowed embeddings with a proper thresholding/ranking protocol, or explicitly state that the ba
  3. [§5.2, §3.2, Table 1] The chain parameters τ_pos=30, τ_off=10, and κ=5 are calibrated on a validation set generated by the same synthetic procedure used for the test set, and no sensitivity analysis is reported. Because the positive examples are constructed to be nearly contiguous, the chain score is essentially measuring whether the input has a long unbroken run of shared fingerprints; the benchmark does not provide evidence about how the method degrades when τ_pos or κ are mismatched to the actual noise level. In particular, the paper does not report a null model or empirical distribution of chain scores for unrelated documents, so it is unclear how the method avoids false positives in a corpus with, say, reused boilerplate or common long n-grams. I ask the authors to add (i) a sweep over τ_pos, τ_off, and κ on the validation set, (ii) a false-positive analysis on random non-overlapping documents from the s
  4. [§4, §5.4] The scalability claim is supported only by query latency (≈450 ms for 50.7M documents). Index construction time, disk footprint, memory-mapped index size, and query throughput under concurrency are not reported. Since the paper's stated contribution includes 'scaling to large web-crawled datasets,' the lack of index-construction measurements makes it difficult for a reader to judge whether the approach is practical at the intended scale. Please report wall-clock index building time for each corpus, index size on disk, and, if possible, the peak memory usage.
minor comments (4)
  1. [§1, Abstract] Several typos and spacing issues: 'FindMyTextextends' (missing space), 'thiry-fourth' in the Broder reference, and inconsistent use of 'd′i' vs 'd_i'. A light proofreading pass is recommended.
  2. [Table 1, note] The dense retrieval baseline is listed as '/' for ArXiv and HPLT. If the authors retain the baseline, they should either compute embeddings for those corpora or explicitly mark the baseline as not evaluated, rather than leaving an ambiguous slash.
  3. [Appendix C] The benchmark examples are useful, but the 'positive' example appears to contain a very long insertion of garden-blog content before the copied Simpsons paragraph. This is actually a good stress case, yet it is not representative of the automated generation described in §5.1. Clarify whether such long prepended unrelated spans are part of the standard positive generation or a hand-constructed illustrative example, and whether the chain method still detects the copied span when the unrelated content is longer than the copied span.
  4. [Appendix A] Embedding a live password ('EMNLP2026') in the paper is not standard practice; if the web interface is meant to be publicly accessible for review, provide an anonymous/unauthenticated link or a clearly temporary password. Also, the GitHub URL is given but no repository DOI or version number is provided; a versioned release would improve reproducibility.

Circularity Check

1 steps flagged · score 6.0 of 10

Benchmark positives are generated as consecutive, lightly edited fragments and negatives by operations that break chains; the chain score (Eq. 2) is defined to detect exactly that structure, so the 0.99 AUC is largely a construction match rather than an independent validation.

  1. fitted input called prediction [Sec. 3.2 (Eq. 2), Sec. 5.1 (Positive/Negative instances), Sec. 5.2 (Hyper-parameters), Sec. 5.4 (Table 1)]
    "we first sample a document d_i from C and then select a consecutive fragment d′_i ⊂ d_i ... Those edits are repeated until we obtain a fragment d′_i that is right above a given threshold T_pos. Finally, the fragment d′_i is concatenated with some other text content ... The pair (d′′_i, d_i) is a positive example as the two documents share by construction a common text region according to the oracle ... we consider two shared fingerprints h_i and h_j as connected if |p_q(h_i)−p_q(h_j)| ≤ τ_pos and |δ(h_i)−δ(h_j)| ≤ τ_off ... Hyper-parameters – such as the τ_pos and τ_off thresholds of the chain"

    The benchmark's positive instances are by construction a single consecutive fragment of d_i, edited only by normalizations and then concatenated with filler; the chain score s_chain(q,d_i)=max_k|C_k| connects shared fingerprints with query-position gaps ≤ τ_pos=30 and offset drift ≤ τ_off=10, which is exactly the structure such consecutive, lightly edited fragments produce. Negatives are generated by permutations, paraphrase, and boilerplate insertion — operations designed to destroy long consecutive shared spans, which is precisely what the chain criterion measures. Thus the near-perfect AUC-ROC/P@R in Table 1 is a match between Eq. (2) and the label-generation process, not an independent test of whether chain score recovers Smith-Waterman-defined containment. Since τ_pos and τ_off were t

full rationale

No self-citation chain, imported uniqueness, or ansatz-via-citation is present: the winnowing backbone is grounded in Schleimer et al. (2003), and the method is self-contained. However, the central experimental claim — that the chain-based method 'discriminates much more effectively between positive and negative examples, achieving an AUC-ROC around 0.99' — is supported only by a benchmark whose labels were generated to have exactly the contiguous, small-offset shared-span property that Eq. (2) detects. Positives are constructed by selecting a consecutive fragment and editing it only until the Smith-Waterman oracle barely exceeds T_pos=1000, so the shared region is guaranteed to form a long fingerprint chain; negatives are constructed by operations that break such chains. The thresholds τ_pos/τ_off are then tuned on a validation set from the same generator. The performance numbers therefore reduce partly to the benchmark construction: they demonstrate that the implementation finds fingerprint chains when the benchmark has placed a chain in every positive and removed chains from every negative, rather than independently establishing robustness to arbitrary SW-positive containment. This is a partial, evaluation-level circularity — the derived score and the label definition are not identical, but the generative process makes them nearly equivalent by construction. The scaling, indexing, and runtime results are independent and not affected by this circularity.

Assumptions & free parameters 8 free parameters · 5 assumptions · 0 invented entities

The paper's own contribution rests on winnowing's exact-match guarantee from Schleimer et al. 2003 (adopted, not re-derived) plus seven hand-set or empirically calibrated parameters (k, w, τ_pos, τ_off, κ, T_pos, T_neg). There is no null model for random fingerprint coincidences: κ=5 and the tolerance windows are justified only by validation-set calibration. No invented entities are introduced — chains are derived objects. The free-parameter count is the honest measure of what the benchmark evaluation required the authors to tune by hand.

free parameters (8)
  • k (shingle length) = 4 words
    Hand-chosen in Sec. 4; sets the minimum granularity of detectable shared fragments; not swept.
  • w (winnowing window) = 6 hashes
    Hand-chosen in Sec. 4; controls fingerprint density together with k.
  • τ_pos (maximum gap between connected fingerprints) = 30 characters
    Calibrated empirically on a validation set synthesized per Sec. 5.1 (Sec. 5.2); directly controls chain connectivity.
  • τ_off (maximum offset drift) = 10 characters
    Calibrated empirically on the same validation set; bounds robustness to local insertions/deletions.
  • κ (minimum cluster size) = 5 fingerprints
    Hand-set in Sec. 3.2; filters spurious chains; no null model justifies the value.
  • T_pos (positive-instance SW-score threshold) = 1000
    Arbitrary benchmark threshold (Sec. 5.1); the 'half a page' equivalence is asserted, not measured.
  • T_neg (negative-instance SW-score threshold) = 100
    Arbitrary benchmark threshold (Sec. 5.1); defines absence of any shared segment longer than 1-2 sentences.
  • LLM paraphrase settings for negative generation = Qwen2-7B-Instruct, ≈400 char chunks, temperature 0.8
    Generation-time choice (Sec. 5.1) that controls how hard the negatives are; stochastic.
assumptions (5)
  • standard math Winnowing guarantee: any shared substring of length ≥ k+w−1 yields at least one common fingerprint (Schleimer et al., 2003)
    Adopted from prior literature in Sec. 3.1; not re-derived; the whole method's recall floor depends on it.
  • domain assumption Word-level k-gram tokenization that ignores punctuation preserves shingle identity under the benchmark's edit operations
    Sec. 4 sets k=4 'ignoring punctuation'; the ligature and de-hyphenation edits in Appendix C change characters inside words, which only survives because hashing is on word shingles.
  • domain assumption Smith-Waterman local alignment score with thresholds T_pos=1000 / T_neg=100 is a valid oracle for 'shared text segment' containment
    Sec. 5.1 defines benchmark labels via this oracle; no calibration or independent validation of the 'half a page' equivalence is given.
  • standard math Fingerprint hash collisions are negligible at the scales used
    Unstated in Sec. 3.1; standard for 64-bit hashing but relied on for correctness of the inverted index.
  • ad hoc to paper A connected cluster of ≥κ=5 fingerprints within τ_pos=30 / τ_off=10 is evidence of reuse, with no null model for chance chains
    Sec. 3.2 sets these thresholds without a random-coincidence model, so false-positive behaviour on varied corpora is unquantified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora." pith.science (2026). https://pith.science/paper/M3S3TIBB

@misc{pith2026260710020,
  author       = {Pith},
  title        = {Pith review of: FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M3S3TIBB}},
  note         = {Machine review of arXiv:2607.10020}
}
read the original abstract

We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints. By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities. This makes FindMyText particularly suited for verifying the presence of copyrighted material in a corpus. Leveraging a distributed, disk-based indexing framework, the system scales to large web-crawled datasets. Using a new benchmark for evaluating text containment methods, we show that FindMyText outperforms alternative approaches across three datasets (ArXiv papers, Wikipedia, and generic web content).

Figures

Figures reproduced from arXiv: 2607.10020 by the authors.

Figure 1
Figure 1. General sketch of the FindMyText approach. The reliability of those attacks have, however, been called into question (Meeus et al., 2024; Zhang et al., 2024a; Liu et al., 2025), particularly for pro￾duction LLMs that are carefully designed to avoid generating copyrighted content. Transparency obligations introduced by recent regulations such as the EU AI Act (European Parlia￾ment and Council, 2024) may in the future… view at source ↗
Figure 2
Figure 2. Illustration of the method for identifying fin [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. FindMyText web interface. 9 [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

299 extracted references · 2 linked inside Pith

  1. [1]

    Why keep hospital clinical records?

    Nicol, A and Sheppard, J. Why keep hospital clinical records?. British medical journal (Clinical research ed.). doi:10.1136/bmj.290.6464.263

  2. [2]

    We hacked a robot vacuum -- and could watch live through its camera

    Fell, Julian. We hacked a robot vacuum -- and could watch live through its camera. ABC News

  3. [3]

    Robot roaming some federal office buildings raises privacy concerns

    van Rooy, Natalie. Robot roaming some federal office buildings raises privacy concerns

  4. [4]

    N kan KI-generert tekst vannmerkes

    Lison, Pierre. N kan KI-generert tekst vannmerkes

  5. [5]

    Kan kunstig intelligens forst spr k?

    Lison, Pierre and Falkum, Ingrid Lossius. Kan kunstig intelligens forst spr k?

  6. [6]

    Privacy and confidentiality issues in historical health sciences collections 2012 law & informatics issue

    Gilliland, Anne T and Wiener, Judith A. Privacy and confidentiality issues in historical health sciences collections 2012 law & informatics issue. Northern Kentucky law review

  7. [7]

    Unsupervised Text Deidentification

    Morris, John and Chiu, Justin and Zabih, Ramin and Rush, Alexander. Unsupervised Text Deidentification. Findings of the Association for Computational Linguistics: EMNLP 2022. doi:10.18653/v1/2022.findings-emnlp.352

  8. [8]

    skweak: Weak Supervision Made Easy for NLP

    Lison, Pierre and Barnes, Jeremy and Hubin, Aliaksandr. skweak: Weak Supervision Made Easy for NLP. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations. doi:10.18653/v1/2021.acl-demo.40

Show all 299 references
  1. [9]

    Replication in visual diffusion models: A survey and outlook

    Wang, Wenhao and Sun, Yifan and Yang, Zongxin and Hu, Zhengdong and Tan, Zhentao and Yang, Yi. Replication in visual diffusion models: A survey and outlook. arXiv [cs.CV]

  2. [10]

    AIStorySimilarity: Quantifying story similarity using narrative for search, IP infringement, and guided creativity

    Chun, Jon. AIStorySimilarity: Quantifying story similarity using narrative for search, IP infringement, and guided creativity. Proceedings of the 28th Conference on Computational Natural Language Learning. doi:10.18653/v1/2024.conll-1.13

  3. [11]

    CopyBench: Measuring literal and non-literal reproduction of copyright-protected text in language model generation

    Chen, Tong and Asai, Akari and Mireshghallah, Niloofar and Min, Sewon and Grimmelmann, James and Choi, Yejin and Hajishirzi, Hannaneh and Zettlemoyer, Luke and Koh, Pang Wei. CopyBench: Measuring literal and non-literal reproduction of copyright-protected text in language mode...

  4. [12]

    Alignment for Honesty

    Yang, Yuqing and Chern, Ethan and Qiu, Xipeng and Neubig, Graham and Liu, Pengfei. Alignment for Honesty. The Thirty-eighth Annual Conference on Neural Information Processing Systems

  5. [13]

    Scaling up membership inference: When and how attacks succeed on large language models

    Puerto, Haritz and Gubri, Martin and Yun, Sangdoo and Oh, Seong Joon. Scaling up membership inference: When and how attacks succeed on large language models. arXiv [cs.CL]

  6. [14]

    Do Membership Inference Attacks Work on Large Language Models?

    Duan, Michael and Suri, Anshuman and Mireshghallah, Niloofar and Min, Sewon and Shi, Weijia and Zettlemoyer, Luke and Tsvetkov, Yulia and Choi, Yejin and Evans, David and Hajishirzi, Hannaneh. Do Membership Inference Attacks Work on Large Language Models?. First Conference on ...

  7. [15]

    Membership inference attacks cannot prove that a model was trained on your data

    Zhang, Jie and Das, Debeshee and Kamath, Gautam and Tram\` e r, Florian. Membership inference attacks cannot prove that a model was trained on your data. arXiv [cs.LG]

  8. [16]

    Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration

    Fu, Wenjie and Wang, Huandong and Gao, Chen and Liu, Guanghua and Li, Yong and Jiang, Tao. Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration. The Thirty-eighth Annual Conference on Neural Information Processing Systems

  9. [17]

    Noisy Neighbors: Efficient membership inference attacks against LLMs

    Galli, Filippo and Melis, Luca and Cucinotta, Tommaso. Noisy Neighbors: Efficient membership inference attacks against LLMs. Proceedings of the Fifth Workshop on Privacy in Natural Language Processing

  10. [18]

    DF-MIA: A distribution-free membership Inference Attack on fine-tuned Large Language Models

    Huang, Zhiheng and Liu, Yannan and He, Daojing and Li, Yu. DF-MIA: A distribution-free membership Inference Attack on fine-tuned Large Language Models. Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence. doi:10.1609/aa...

  11. [19]

    Membership inference attacks against large vision-language models

    Li, Zhan and Wu, Yongtao and Chen, Yihang and Tonin, Francesco and Abad-Rocamora, El\' as and Cevher, V. Membership inference attacks against large vision-language models. Neural Information Processing Systems. doi:10.48550/arXiv.2411.02902

  12. [20]

    Not all tokens are equal: Membership inference attacks against fine-tuned language models

    Song, Changtian and Zhao, Dongdong and Xiang, Jianwen. Not all tokens are equal: Membership inference attacks against fine-tuned language models. 2024 Annual Computer Security Applications Conference (ACSAC). doi:10.1109/acsac63791.2024.00020

  13. [21]

    Data Attribution for Text-to-Image Models by Unlearning Synthesized Images

    Wang, Sheng-Yu and Hertzmann, Aaron and Efros, Alexei A and Zhu, Jun-Yan and Zhang, Richard. Data Attribution for Text-to-Image Models by Unlearning Synthesized Images. Advances in Neural Information Processing Systems

  14. [22]

    Source-Aware Training Enables Knowledge Attribution in Language Models

    Khalifa, Muhammad and Wadden, David and Strubell, Emma and Lee, Honglak and Wang, Lu and Beltagy, Iz and Peng, Hao. Source-Aware Training Enables Knowledge Attribution in Language Models. First Conference on Language Modeling

  15. [23]

    Large language models for automated open-domain scientific hypotheses discovery

    Yang, Zonglin and Du, Xinya and Li, Junxian and Zheng, Jie and Poria, Soujanya and Cambria, Erik. Large language models for automated open-domain scientific hypotheses discovery. Findings of the Association for Computational Linguistics ACL 2024. doi:10.18653/v1/2024.findings-acl.804

  16. [24]

    HoneyBee: Progressive instruction finetuning of large language models for materials science

    Song, Yu and Miret, Santiago and Zhang, Huan and Liu, Bang. HoneyBee: Progressive instruction finetuning of large language models for materials science. Findings of the Association for Computational Linguistics: EMNLP 2023. doi:10.18653/v1/2023.findings-emnlp.380

  17. [25]

    A large-scale audit of dataset licensing and attribution in AI

    Longpre, Shayne and Mahari, Robert and Chen, Anthony and Obeng-Marnu, Naana and Sileo, Damien and Brannon, William and Muennighoff, Niklas and Khazam, Nathan and Kabbara, Jad and Perisetla, Kartik and Wu, Xinyi and Shippole, Enrico and Bollacker, Kurt and Wu, Tongshuang and Vi...

  18. [26]

    Inherently privacy-preserving vision for trustworthy autonomous systems: Needs and solutions

    Taras, Adam K and S \" u nderhauf, Niko and Corke, Peter and Dansereau, Donald G. Inherently privacy-preserving vision for trustworthy autonomous systems: Needs and solutions. Journal of Responsible Technology. doi:10.1016/j.jrt.2024.100079

  19. [27]

    Unsupervised utility evaluation of text anonymization methods via neural language models

    Manzanares-Salor, Benet and S\' a nchez, David and Lison, Pierre. Unsupervised utility evaluation of text anonymization methods via neural language models. Neural Networks: The Official Journal of the International Neural Network Society. doi:10.1016/j.neunet.2026.109079

  20. [28]

    Thermal imaging in robotics as a privacy-enhancing or privacy-invasive measure? Misconceptions of privacy when using thermal cameras in robots

    Lintvedt, Mona Naomi. Thermal imaging in robotics as a privacy-enhancing or privacy-invasive measure? Misconceptions of privacy when using thermal cameras in robots. Social Science Research Network

  21. [29]

    Privacy-preserving robot vision with anonymized faces by extreme low resolution

    Kim, Myeung Un and Lee, Harim and Yang, Hyun Jong and Ryoo, Michael S. Privacy-preserving robot vision with anonymized faces by extreme low resolution. 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). doi:10.1109/iros40897.2019.8967681

  22. [30]

    MicPro: Microphone-based voice privacy protection

    Xiao, Shilin and Ji, Xiaoyu and Yan, Chen and Zheng, Zhicong and Xu, Wenyuan. MicPro: Microphone-based voice privacy protection. Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. doi:10.1145/3576915.3616616

  23. [31]

    Investigating privacy in the context of office delivery robots

    Grasso, Maria Antonietta and Willamowski, Jutta and Park, Jisun and Bak, Sure. Investigating privacy in the context of office delivery robots. 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN). doi:10.1109/ro-man60168.2024.10731254

  24. [32]

    Robots as AI Double Agents: Privacy in Motion Planning

    Shome, Rahul and Kingston, Zachary and Kavraki, Lydia E. Robots as AI Double Agents: Privacy in Motion Planning. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). doi:10.1109/iros55552.2023.10341460

  25. [33]

    Themes and research directions in privacy-sensitive robotics

    Rueben, Matthew and Aroyo, Alexander Mois and Lutz, Christoph and Schmolz, Johannes and Van Cleynenbreugel, Pieter and Corti, Andrea and Agrawal, Siddharth and Smart, William D. Themes and research directions in privacy-sensitive robotics. 2018 IEEE Workshop on Advanced Roboti...

  26. [34]

    Privacy beyond data: Assessment and mitigation of privacy risks in robotic technology for elderly care

    Grabler, Reinhard and Koeszegi, Sabine Theresia. Privacy beyond data: Assessment and mitigation of privacy risks in robotic technology for elderly care. ACM transactions on human-robot interaction. doi:10.1145/3689216

  27. [35]

    Understanding users' perception of privacy in human-robot interaction

    Lee, Min Kyung and Tang, Karen P and Forlizzi, Jodi and Kiesler, Sara. Understanding users' perception of privacy in human-robot interaction. Proceedings of the 6th international conference on Human-robot interaction. doi:10.1145/1957656.1957721

  28. [36]

    Privacy considerations for socially assistive robots

    Fosch-Villaronga, Eduard and Drukarch, Hadassah and Custers, Bart. Privacy considerations for socially assistive robots. The Cambridge Handbook of the Law, Policy, and Regulation for Human--Robot Interaction. doi:10.1017/9781009386708.037

  29. [37]

    Surveying adult perceptions of privacy and attitudes towards social robots in the home

    Levinson, Leigh and Barrett, Tyler and Gomez, Randy and Sabanovi\' c , Selma. Surveying adult perceptions of privacy and attitudes towards social robots in the home. 2024 ACM/IEEE International Conference on Human-Robot Interaction. doi:10.1145/3610978.3640719

  30. [38]

    Remix: Making art commerce thrive hybrid economy

    Lessig, L. Remix: Making art commerce thrive hybrid economy

  31. [39]

    Conceptions Code: How metaphors explain legal challenges digital times

    Larsson, S. Conceptions Code: How metaphors explain legal challenges digital times

  32. [40]

    DE-COP: Detecting Copyrighted Content in Language Models Training Data

    Duarte, Andr\' e V and Zhao, Xuandong and Oliveira, Arlindo L and Li, Lei. DE-COP: Detecting Copyrighted Content in Language Models Training Data. arXiv [cs.CL]

  33. [41]

    HoneyComb: A Flexible LLM-Based Agent System for Materials Science

    Zhang, Huan and Song, Yu and Hou, Ziyu and Miret, Santiago and Liu, Bang. HoneyComb: A Flexible LLM-Based Agent System for Materials Science. Findings of the Association for Computational Linguistics: EMNLP 2024

  34. [42]

    Materials science in the era of large language models: a perspective

    Lei, Ge and Docherty, Ronan and Cooper, Samuel J. Materials science in the era of large language models: a perspective. Digital Discovery

  35. [43]

    MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration

    Ni, Ziqi and Li, Yahao and Hu, Kaijia and Han, Kunyuan and Xu, Ming and Chen, Xingyu and Liu, Fengqi and Ye, Yicong and Bai, Shuxin. MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration. arXiv preprint arXiv:2411. 08063

  36. [44]

    MaScQA: investigating materials science knowledge of large language models

    Zaki, Mohd and Jayadeva and Mausam and Krishnan, N M Anoop. MaScQA: investigating materials science knowledge of large language models. Digit. Discov

  37. [45]

    Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents

    Kumbhar, Shrinidhi and Mishra, Venkatesh and Coutinho, Kevin and Handa, Divij and Iquebal, Ashif and Baral, Chitta. Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents. arXiv preprint arXiv:2501. 13299

  38. [46]

    Crystal structure generation with autoregressive large language modeling

    Antunes, Luis M and Butler, Keith T and Grau-Crespo, Ricardo. Crystal structure generation with autoregressive large language modeling. Nature Communications

  39. [47]

    Leveraging large language models for predictive chemistry

    Jablonka, K and Schwaller, P and Ortega-Guerrero, Andres and Smit, Berend. Leveraging large language models for predictive chemistry. Nat. Mac. Intell. doi:10.1038/s42256-023-00788-1

  40. [48]

    From text to insight: large language models for materials science data extraction

    Schilling-Wilhelmi, Mara and R\' os-Garc\' a, Marti\ n o and Shabih, Sherjeel and Gil, Mar\' a Victoria and Miret, Santiago and Koch, Christoph T and M\' a rquez, Jos\' e A and Jablonka, Kevin Maik. From text to insight: large language models for materials science data extract...

  41. [49]

    LLMatDesign: Autonomous Materials Discovery with Large Language Models

    Jia, Shuyi and Zhang, Chao and Fung, Victor. LLMatDesign: Autonomous Materials Discovery with Large Language Models. arXiv preprint arXiv:2406. 13163

  42. [50]

    MatText: Do Language Models Need More than Text & Scale for Materials Modeling?

    Alampara, Nawaf and Miret, Santiago and Jablonka, Kevin Maik. MatText: Do Language Models Need More than Text & Scale for Materials Modeling?. AI for Accelerated Materials Design - Vienna 2024

  43. [51]

    MechGPT, a Language-Based Strategy for Mechanics and Materials Modeling That Connects Knowledge Across Scales, Disciplines, and Modalities

    Buehler, Markus J. MechGPT, a Language-Based Strategy for Mechanics and Materials Modeling That Connects Knowledge Across Scales, Disciplines, and Modalities. Applied Mechanics Reviews

  44. [52]

    Self-Alignment with Instruction Backtranslation

    Li, Xian and Yu, Ping and Zhou, Chunting and Schick, Timo and Zettlemoyer, Luke and Levy, Omer and Weston, Jason and Lewis, Mike. Self-Alignment with Instruction Backtranslation. arXiv [cs.CL]

  45. [53]

    Effects of social behaviors of robots in privacy-sensitive situations

    Yang, Daseul and Chae, Yu-Jung and Kim, Doogon and Lim, Yoonseob and Kim, Dong Hwan and Kim, Changhwan and Park, Sung-Kee and Nam, Changjoo. Effects of social behaviors of robots in privacy-sensitive situations. International journal of social robotics. doi:10.1007/s12369-021-00809-2

  46. [54]

    A Watermark for Large Language Models

    Kirchenbauer, John and Geiping, Jonas and Wen, Yuxin and Katz, Jonathan and Miers, Ian and Goldstein, Tom. A Watermark for Large Language Models. International Conference on Machine Learning

  47. [55]

    Privacy and Transparency in Human--Robot Interaction

    Ho, Chih-Hsing. Privacy and Transparency in Human--Robot Interaction. The Cambridge Handbook of the Law, Policy, and Regulation for Human--Robot Interaction

  48. [56]

    Consent in Crisis: The Rapid Decline of the AI Data Commons

    Longpre, Shayne and Mahari, Robert and Lee, Ariel and Lund, Campbell and Oderinwale, Hamidah and Brannon, William and Saxena, Nayan and Obeng-Marnu, Naana and South, Tobin and Hunter, Cole and Klyman, Kevin and Klamm, Christopher and Schoelkopf, Hailey and Singh, Nikhil and Ch...

  49. [57]

    Retrieval-augmented generation for natural language processing: A survey

    Wu, Shangyu and Xiong, Ying and Cui, Yufei and Wu, Haolun and Chen, Can and Yuan, Ye and Huang, Lianming and Liu, Xue and Kuo, Tei-Wei and Guan, Nan and Xue, Chun Jason. Retrieval-augmented generation for natural language processing: A survey. arXiv [cs.CL]

  50. [58]

    What ' s in the Box? An Analysis of Undesirable Content in the C ommon C rawl Corpus

    Luccioni, Alexandra and Viviano, Joseph. What ' s in the Box? An Analysis of Undesirable Content in the C ommon C rawl Corpus. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Languag...

  51. [59]

    GRUEN for Evaluating Linguistic Quality of Generated Text

    Zhu, Wanzheng and Bhat, Suma. GRUEN for Evaluating Linguistic Quality of Generated Text. Findings of the Association for Computational Linguistics: EMNLP 2020. doi:10.18653/v1/2020.findings-emnlp.9

  52. [60]

    Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity

    Wang, Cunxiang and Liu, Xiaoze and Yue, Yuanhao and Tang, Xiangru and Zhang, Tianhang and Jiayang, Cheng and Yao, Yunzhi and Gao, Wenyang and Hu, Xuming and Qi, Zehan and Wang, Yidong and Yang, Linyi and Wang, Jindong and Xie, Xing and Zhang, Zheng and Zhang, Yue. Survey on Fa...

  53. [61]

    PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change

    Valmeekam, Karthik and Marquez, Matthew and Olmo, Alberto and Sreedharan, Sarath and Kambhampati, Subbarao. PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change. arXiv [cs.CL]

  54. [62]

    Copyright and Generative AI: Opinion of the European Copyright Society

    Dusollier, Severine and Kretschmer, Martin and Margoni, Thomas and Mezei, P\' e ter and Quintais, Jo\ a o Pedro and Rognstad, Ole Andreas. Copyright and Generative AI: Opinion of the European Copyright Society. Social Science Research Network

  55. [63]

    Generative AI meets copyright

    Samuelson, Pamela. Generative AI meets copyright. Science (New York, N.Y.). doi:10.1126/science.adi0656

  56. [64]

    In search of balance: text data mining and copyright in the digital single market directive from a fundamental rights persperctive

    Manteghi, M. In search of balance: text data mining and copyright in the digital single market directive from a fundamental rights persperctive. European law review

  57. [65]

    Copyright Traps for Large Language Models

    Meeus, Matthieu and Shilov, Igor and Faysse, Manuel and De Montjoye, Yves-Alexandre. Copyright Traps for Large Language Models. International Conference on Machine Learning

  58. [66]

    The Ai-Copyright Trap

    Craig, Carys J. The Ai-Copyright Trap. Social Science Research Network. doi:10.2139/ssrn.4905118

  59. [67]

    Consent and compensation: Resolving generative AI's copyright crisis essay

    Pasquale, Frank and Sun, Haochen. Consent and compensation: Resolving generative AI's copyright crisis essay. Virginia Law Review Online

  60. [68]

    We Need Smart Intellectual Property Laws for Artificial Intelligence

    Love, James. We Need Smart Intellectual Property Laws for Artificial Intelligence. Scientific American

  61. [69]

    The Data That Powers A.I

    Roose, Kevin. The Data That Powers A.I. Is Disappearing Fast. The New York Times

  62. [70]

    The State and Fate of Linguistic Diversity and Inclusion in the NLP World

    Joshi, Pratik and Santy, Sebastin and Budhiraja, Amar and Bali, Kalika and Choudhury, Monojit. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. doi:10.18653/v1/20...

  63. [71]

    A survey of fake news: Fundamental theories, detection methods, and opportunities

    Zhou, Xinyi. A survey of fake news: Fundamental theories, detection methods, and opportunities. ACM Computing Surveys (CSUR)

  64. [72]

    On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?

    Bender, Emily M and Gebru, Timnit and McMillan-Major, Angelina and Shmitchell, Shmargaret. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. doi:10.1145/3442188.3445922

  65. [73]

    The work of copyright law in the age of generative AI

    Crawford, Kate and Schultz, Jason. The work of copyright law in the age of generative AI. Grey room. doi:10.1162/grey\_a\_00389

  66. [74]

    Copyright Violations and Large Language Models

    Karamolegkou, Antonia and Li, Jiaang and Zhou, Li and S gaard, Anders. Copyright Violations and Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/2023.emnlp-main.458

  67. [76]

    Detecting Pretraining Data from Large Language Models

    Shi, Weijia and Ajith, Anirudh and Xia, Mengzhou and Huang, Yangsibo and Liu, Daogao and Blevins, Terra and Chen, Danqi and Zettlemoyer, Luke. Detecting Pretraining Data from Large Language Models. arXiv [cs.CL]

  68. [77]

    A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

    Huang, Lei and Yu, Weijiang and Ma, Weitao and Zhong, Weihong and Feng, Zhangyin and Wang, Haotian and Chen, Qianglong and Peng, Weihua and Feng, Xiaocheng and Qin, Bing and Liu, Ting. A survey on hallucination in large language models: Principles, taxonomy, challenges, and op...

  69. [78]

    Generative AI and Author Remuneration

    Senftleben, Martin. Generative AI and Author Remuneration. IIC - International Review of Intellectual Property and Competition Law. doi:10.1007/s40319-023-01399-4

  70. [79]

    From text to talk: H arnessing conversational corpora for humane and diversity-aware language technology

    Dingemanse, Mark and Liesenfeld, Andreas. From text to talk: H arnessing conversational corpora for humane and diversity-aware language technology. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). doi:10.18653/v1/...

  71. [80]

    Self-Rewarding Language Models

    Yuan, Weizhe and Pang, Richard Yuanzhe and Cho, Kyunghyun and Sukhbaatar, Sainbayar and Xu, Jing and Weston, Jason. Self-Rewarding Language Models. arXiv [cs.CL]

  72. [81]

    Protecting User Data Through Privacy-Sensitive Robot Design

    Sullivan, Dakota and Mutlu, Bilge. Protecting User Data Through Privacy-Sensitive Robot Design. Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. doi:10.5555/3721488.3721804

  73. [82]

    Privacy concerns in the smart home context

    Guhr, Nadine and Werth, Oliver and Blacha, Philip Peter Hermann and Breitner, Michael H. Privacy concerns in the smart home context. SN applied sciences. doi:10.1007/s42452-020-2025-8

  74. [83]

    Evaluating current automatic de-identification methods with Veteran's health administration clinical documents

    Ferr\' a ndez, O and South, B R and Shen, S and Friedlin, F J and Samore, M H and Meystre, S M. Evaluating current automatic de-identification methods with Veteran's health administration clinical documents. BMC medical research methodology

  75. [84]

    Personalised Reranking of Paper Recommendations Using Paper Content and User Behavior

    Li, Xinyi and Chen, Yifan and Pettit, Benjamin and Rijke, Maarten De. Personalised Reranking of Paper Recommendations Using Paper Content and User Behavior. ACM Transactions on Information and System Security. doi:10.1145/3312528

  76. [85]

    Data valuation in machine learning:``ingredients'', strategies, and open challenges

    Sim, Rachael Hwee Ling and Xu, Xinyi and Low, Bryan Kian Hsiang. Data valuation in machine learning:``ingredients'', strategies, and open challenges. Proc. IJCAI

  77. [86]

    Publicly detectable watermarking for language models

    Fairoze, Jaiden and Garg, Sanjam and Jha, Somesh and Mahloujifar, Saeed and Mahmoody, Mohammad and Wang, Mingyuan. Publicly detectable watermarking for language models. arXiv [cs.LG]

  78. [87]

    SMET: Semantic Mapping of CVE to ATT&CK and Its Application to Cybersecurity

    Abdeen, Basel and Al-Shaer, Ehab and Singhal, Anoop and Khan, Latifur and Hamlen, Kevin. SMET: Semantic Mapping of CVE to ATT&CK and Its Application to Cybersecurity. Data and Applications Security and Privacy XXXVII. doi:10.1007/978-3-031-37586-6\_15

  79. [88]

    A hybrid approach to dialogue management based on probabilistic rules

    Lison, Pierre. A hybrid approach to dialogue management based on probabilistic rules. Computer Speech & Language. doi:10.1016/j.csl.2015.01.001

  80. [89]

    Situated dialogue processing for human-robot interaction

    Kruijff, Geert-Jan M and Lison, Pierre and Benjamin, Trevor and Jacobsson, Henrik and Zender, Hendrik and Kruijff-Korbayov\' a , Ivana and Hawes, Nick. Situated dialogue processing for human-robot interaction. Cognitive Systems Monographs

  81. [90]

    Self-understanding and self-extension: A systems and representational approach

    Wyatt, Jeremy L and Aydemir, Alper and Brenner, Michael and Hanheide, Marc and Hawes, Nick and Jensfelt, Patric and Kristan, Matej and Kruijff, Geert-Jan M and Lison, Pierre and Pronobis, Andrzej and Sjoo, Kristoffer and Vrecko, Alen and Zender, Hendrik and Zillich, Michael an...

  82. [91]

    Spoken dialogue systems: the new frontier in human-computer interaction

    Lison, Pierre and Meena, Raveesh. Spoken dialogue systems: the new frontier in human-computer interaction. XRDS Crossroads The ACM Magazine for Students. doi:10.1145/2659891

  83. [92]

    Automatic turn segmentation for Movie & TV subtitles

    Lison, Pierre and Meena, Raveesh. Automatic turn segmentation for Movie & TV subtitles. 2016 IEEE Spoken Language Technology Workshop (SLT). doi:10.1109/slt.2016.7846272

  84. [93]

    PyOpenDial: A python-based domain-independent toolkit for developing spoken dialogue systems with probabilistic rules

    Jang, Youngsoo and Lee, Jongmin and Park, Jaeyoung and Lee, Kyeng-Hun and Lison, Pierre and Kim, Kee-Eung. PyOpenDial: A python-based domain-independent toolkit for developing spoken dialogue systems with probabilistic rules. Proceedings of the 2019 Conference on Empirical Met...

  85. [94]

    A salience-driven approach to speech recognition for human-robot interaction

    Lison, Pierre. A salience-driven approach to speech recognition for human-robot interaction. Interfaces: Explorations in Logic, Language and Computation. doi:10.1007/978-3-642-14729-6\_8

  86. [95]

    Towards online planning for dialogue management with rich domain knowledge

    Lison, Pierre. Towards online planning for dialogue management with rich domain knowledge. Natural Interaction with Robots, Knowbots and Smartphones. doi:10.1007/978-1-4614-8280-2\_11

  87. [96]

    Should we use movie subtitles to study linguistic patterns of conversational speech? A study based on French, English and Taiwan Mandarin

    Pr\' e vot, Laurent and Magistry, Pierre and Lison, Pierre. Should we use movie subtitles to study linguistic patterns of conversational speech? A study based on French, English and Taiwan Mandarin. Third International Symposium on Linguitic Patters of Spontaneous Speech

  88. [97]

    Neural Text Sanitization with Privacy Risk Indicators: An Empirical Analysis

    Papadopoulou, Anthi and Lison, Pierre and Anderson, Mark and vrelid, Lilja and Pil\' a n, Ildik\' o. Neural Text Sanitization with Privacy Risk Indicators: An Empirical Analysis. arXiv [cs.CL]. doi:10.48550/arXiv.2310.14312

  89. [98]

    Identifying Token-Level Dialectal Features in Social Media

    Barnes, Jeremy and Touileb, Samia and M hlum, Petter and Lison, Pierre. Identifying Token-Level Dialectal Features in Social Media. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)

  90. [99]

    Who's in charge? Roles and responsibilities of decision-making components in conversational robots

    Lison, Pierre and Kennington, Casey. Who's in charge? Roles and responsibilities of decision-making components in conversational robots. HRI 2023 Workshop on Human-Robot Conversational Interaction

  91. [100]

    PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models

    Li, Peixuan and Cheng, Pengzhou and Li, Fangqi and Du, Wei and Zhao, Haodong and Liu, Gongshen. PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models. Proceedings of the AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v37i12.26750

  92. [101]

    Three Bricks to Consolidate Watermarks for Large Language Models

    Fernandez, Pierre and Chaffin, Antoine and Tit, Karim and Chappelier, Vivien and Furon, Teddy. Three Bricks to Consolidate Watermarks for Large Language Models. 2023 IEEE International Workshop on Information Forensics and Security (WIFS). doi:10.1109/WIFS58808.2023.10374576

  93. [102]

    Towards Efficient Data Valuation Based on the Shapley Value

    Jia, Ruoxi and Dao, David and Wang, Boxin and Hubis, Frances Ann and Hynes, Nick and G \" u rel, Nezihe Merve and Li, Bo and Zhang, Ce and Song, Dawn and Spanos, Costas J. Towards Efficient Data Valuation Based on the Shapley Value. Proceedings of the Twenty-Second Internation...

  94. [103]

    Data Shapley: Equitable Valuation of Data for Machine Learning

    Ghorbani, Amirata and Zou, James. Data Shapley: Equitable Valuation of Data for Machine Learning. Proceedings of the 36th International Conference on Machine Learning

  95. [104]

    The Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence

    Crawford, Kate. The Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence

  96. [105]

    Dis, c'est quoi l'intelligence artificielle ?

    Lison, Pierre and Julia, Luc. Dis, c'est quoi l'intelligence artificielle ?

  97. [106]

    Gemini: a family of highly capable multimodal models

    Google. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv

  98. [107]

    GLM: General Language Model Pretraining with Autoregressive Blank Infilling

    Du, Zhengxiao and Qian, Yujie and Liu, Xiao and Ding, Ming and Qiu, Jiezhong and Yang, Zhilin and Tang, Jie. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. arXiv [cs.CL]

  99. [108]

    Retrieval Augmented Language Model Pre-Training

    Guu, Kelvin and Lee, Kenton and Tung, Zora and Pasupat, Panupong and Chang, Mingwei. Retrieval Augmented Language Model Pre-Training. Proceedings of the 37th International Conference on Machine Learning

  100. [109]

    In-Context Retrieval-Augmented Language Models

    Ram, Ori and Levine, Yoav and Dalmedigos, Itay and Muhlgay, Dor and Shashua, Amnon and Leyton-Brown, Kevin and Shoham, Yoav. In-Context Retrieval-Augmented Language Models. arXiv [cs.CL]

  101. [110]

    RETA-LLM: A Retrieval-Augmented Large Language Model Toolkit

    Liu, Jiongnan and Jin, Jiajie and Wang, Zihan and Cheng, Jiehan and Dou, Zhicheng and Wen, Ji-Rong. RETA-LLM: A Retrieval-Augmented Large Language Model Toolkit. arXiv [cs.IR]

  102. [111]

    Improving Language Models by Retrieving from Trillions of Tokens

    Borgeaud, Sebastian and Mensch, Arthur and Hoffmann, Jordan and Cai, Trevor and Rutherford, Eliza and Millican, Katie and Van Den Driessche, George Bm and Lespiau, Jean-Baptiste and Damoc, Bogdan and Clark, Aidan and De Las Casas, Diego and Guy, Aurelia and Menick, Jacob and R...

  103. [112]

    Dense Passage Retrieval for Open-Domain Question Answering

    Karpukhin, Vladimir and O g uz, Barlas and Min, Sewon and Lewis, Patrick and Wu, Ledell and Edunov, Sergey and Chen, Danqi and Yih, Wen-Tau. Dense Passage Retrieval for Open-Domain Question Answering. arXiv [cs.CL]

  104. [113]

    Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

    Xiong, Lee and Xiong, Chenyan and Li, Ye and Tang, Kwok-Fung and Liu, Jialin and Bennett, Paul and Ahmed, Junaid and Overwijk, Arnold. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. arXiv [cs.IR]

  105. [114]

    WebGPT: Browser-assisted question-answering with human feedback

    Nakano, Reiichiro and Hilton, Jacob and Balaji, Suchir and Wu, Jeff and Ouyang, Long and Kim, Christina and Hesse, Christopher and Jain, Shantanu and Kosaraju, Vineet and Saunders, William and Jiang, Xu and Cobbe, Karl and Eloundou, Tyna and Krueger, Gretchen and Button, Kevin...

  106. [115]

    Fine-Tuning Language Models from Human Preferences

    Ziegler, Daniel M and Stiennon, Nisan and Wu, Jeffrey and Brown, Tom B and Radford, Alec and Amodei, Dario and Christiano, Paul and Irving, Geoffrey. Fine-Tuning Language Models from Human Preferences. arXiv [cs.CL]

  107. [116]

    A survey of human-in-the-loop for machine learning

    Wu, Xingjiao and Xiao, Luwei and Sun, Yixuan and Zhang, Junhang and Ma, Tianlong and He, Liang. A survey of human-in-the-loop for machine learning. Future generations computer systems: FGCS. doi:10.1016/j.future.2022.05.014

  108. [117]

    Training language models to follow instructions with human feedback

    Ouyang, Long and Wu, Jeff and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll L and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and Schulman, John and Hilton, Jacob and Kelton, Fraser and Miller, Luke and Simens, Maddie and Ask...

  109. [118]

    RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    Zitkovich, Brianna and Yu, Tianhe and Xu, Sichun and Xu, Peng and Xiao, Ted and Xia, Fei and Wu, Jialin and Wohlhart, Paul and Welker, Stefan and Wahid, Ayzaan and Vuong, Quan and Vanhoucke, Vincent and Tran, Huong and Soricut, Radu and Singh, Anikait and Singh, Jaspiar and Se...

  110. [119]

    Large Language Models for Robotics: A Survey

    Zeng, Fanlong and Gan, Wensheng and Wang, Yongheng and Liu, Ning and Yu, Philip S. Large Language Models for Robotics: A Survey. arXiv [cs.RO]

  111. [120]

    Code as Policies: Language Model Programs for Embodied Control

    Liang, Jacky and Huang, Wenlong and Xia, Fei and Xu, Peng and Hausman, Karol and Ichter, Brian and Florence, Pete and Zeng, Andy. Code as Policies: Language Model Programs for Embodied Control. 2023 IEEE International Conference on Robotics and Automation (ICRA). doi:10.1109/I...

  112. [121]

    Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

    Ichter, Brian and Brohan, Anthony and Chebotar, Yevgen and Finn, Chelsea and Hausman, Karol and Herzog, Alexander and Ho, Daniel and Ibarz, Julian and Irpan, Alex and Jang, Eric and Julian, Ryan and Kalashnikov, Dmitry and Levine, Sergey and Lu, Yao and Parada, Carolina and Ra...

  113. [122]

    doi:10.7907/q75sz-e1e79

    CMS/ACM 117: Probability Theory & Computational Mathematics. doi:10.7907/q75sz-e1e79

  114. [123]

    Efficiently Modeling Long Sequences with Structured State Spaces

    Gu, Albert and Goel, Karan and R\' e , Christopher. Efficiently Modeling Long Sequences with Structured State Spaces. arXiv [cs.LG]

  115. [124]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces

    Gu, Albert and Dao, Tri. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv [cs.LG]

  116. [125]

    DataPerf: Benchmarks for Data-Centric AI Development

    Mazumder, Mark and Banbury, Colby and Yao, Xiaozhe and Karla s , Bojan and Rojas, William Gaviria and Diamos, Sudnya and Diamos, Greg and He, Lynn and Parrish, Alicia and Kirk, Hannah Rose and Quaye, Jessica and Rastogi, Charvi and Kiela, Douwe and Jurado, David and Kanter, Da...

  117. [126]

    Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

    Khattab, Omar and Santhanam, Keshav and Li, Xiang Lisa and Hall, David and Liang, Percy and Potts, Christopher and Zaharia, Matei. Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP. arXiv [cs.CL]

  118. [127]

    First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models

    Saphra, Naomi and Fleisig, Eve and Cho, Kyunghyun and Lopez, Adam. First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models. arXiv [cs.CL]

  119. [128]

    QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

    Xu, Yuhui and Xie, Lingxi and Gu, Xiaotao and Chen, Xin and Chang, Heng and Zhang, Hengheng and Chen, Zhengsu and Zhang, Xiaopeng and Tian, Qi. QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models. arXiv [cs.LG]

  120. [129]

    Privacy side channels in machine learning systems

    Debenedetti, Edoardo and Severi, Giorgio and Carlini, Nicholas and Choquette-Choo, Christopher A and Jagielski, Matthew and Nasr, Milad and Wallace, Eric and Tram\` e r, Florian. Privacy side channels in machine learning systems. arXiv [cs.CR]

  121. [130]

    Addressing the Blind Spots in Spoken Language Processing

    Moryossef, Amit. Addressing the Blind Spots in Spoken Language Processing. arXiv [eess.AS]

  122. [131]

    Probabilistic Machine Learning: Advanced Topics (Adaptive Computation and Machine Learning series)

    Murphy, Kevin P. Probabilistic Machine Learning: Advanced Topics (Adaptive Computation and Machine Learning series)

  123. [132]

    Uncertainty in Natural Language Generation: From Theory to Applications

    Baan, Joris and Daheim, Nico and Ilia, Evgenia and Ulmer, Dennis and Li, Haau-Sing and Fern\' a ndez, Raquel and Plank, Barbara and Sennrich, Rico and Zerva, Chrysoula and Aziz, Wilker. Uncertainty in Natural Language Generation: From Theory to Applications. arXiv [cs.CL]

  124. [133]

    Challenges and Applications of Large Language Models

    Kaddour, Jean and Harris, Joshua and Mozes, Maximilian and Bradley, Herbie and Raileanu, Roberta and McHardy, Robert. Challenges and Applications of Large Language Models. arXiv [cs.CL]

  125. [134]

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Ermon, Stefano and Manning, Christopher D and Finn, Chelsea. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv [cs.LG]

  126. [135]

    Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve

    Thomas McCoy, R and Yao, Shunyu and Friedman, Dan and Hardy, Matthew and Griffiths, Thomas L. Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve. arXiv [cs.CL]

  127. [136]

    Graph Neural Prompting with Large Language Models

    Tian, Yijun and Song, Huan and Wang, Zichen and Wang, Haozhu and Hu, Ziqing and Wang, Fang and Chawla, Nitesh V and Xu, Panpan. Graph Neural Prompting with Large Language Models. arXiv [cs.CL]

  128. [137]

    Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs

    Zhang, Qingru and Singh, Chandan and Liu, Liyuan and Liu, Xiaodong and Yu, Bin and Gao, Jianfeng and Zhao, Tuo. Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs. arXiv [cs.CL]

  129. [138]

    Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    Fernando, Chrisantha and Banarse, Dylan and Michalewski, Henryk and Osindero, Simon and Rockt \" a schel, Tim. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv [cs.CL]

  130. [139]

    a rli, Nathanael and Chowdhery, Aakanksha and Mansfield, Philip and Demner-Fushman, Dina and Ag \

    Singhal, Karan and Azizi, Shekoofeh and Tu, Tao and Mahdavi, S Sara and Wei, Jason and Chung, Hyung Won and Scales, Nathan and Tanwani, Ajay and Cole-Lewis, Heather and Pfohl, Stephen and Payne, Perry and Seneviratne, Martin and Gamble, Paul and Kelly, Chris and Babiker, Abuba...

  131. [140]

    Deepproblog: Neural probabilistic logic programming

    Manhaeve, R and Dumancic, S and Kimmig, A and others. Deepproblog: Neural probabilistic logic programming. Advances in neural information processing systems

  132. [141]

    Neurosymbolic AI: the 3rd wave

    Garcez, Artur D'avila and Lamb, Lu\' s C. Neurosymbolic AI: the 3rd wave. Artificial Intelligence Review. doi:10.1007/s10462-023-10448-w

  133. [142]

    Named Entity Recognition without Labelled Data: A Weak Supervision Approach

    Lison, Pierre and Barnes, Jeremy and Hubin, Aliaksandr and Touileb, Samia. Named Entity Recognition without Labelled Data: A Weak Supervision Approach. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. doi:10.18653/v1/2020.acl-main.139

  134. [143]

    Do as i can, not as i say: Grounding language in robotic affordances

    Ahn, Michael and Brohan, Anthony and Brown, Noah and Chebotar, Yevgen and Cortes, Omar and David, Byron and Finn, Chelsea and Fu, Chuyuan and Gopalakrishnan, Keerthana and Hausman, Karol and Others. Do as i can, not as i say: Grounding language in robotic affordances. arXiv pr...

  135. [144]

    On the bottleneck of graph neural networks and its practical implications

    Alon, Uri and Yahav, Eran. On the bottleneck of graph neural networks and its practical implications. arXiv preprint arXiv:2006. 05205

  136. [145]

    DIALOGPT: Large-Scale Generative Pre-training for Conversational Response Generation

    Zhang, Yizhe and Sun, Siqi and Galley, Michel and Chen, Yen-Chun and Brockett, Chris and Gao, Xiang and Gao, Jianfeng and Liu, Jingjing and Dolan, William B. DIALOGPT: Large-Scale Generative Pre-training for Conversational Response Generation. Proceedings of the 58th Annual Me...

  137. [146]

    Recipes for Building an Open-Domain Chatbot

    Roller, Stephen and Dinan, Emily and Goyal, Naman and Ju, Da and Williamson, Mary and Liu, Yinhan and Xu, Jing and Ott, Myle and Smith, Eric Michael and Boureau, Y-Lan and Others. Recipes for Building an Open-Domain Chatbot. Proceedings of the 16th Conference of the European C...

  138. [147]

    AARGH! End-to-end Retrieval-Generation for Task-Oriented Dialog

    Nekvinda, Tom\' a s and Du s ek, Ond r ej. AARGH! End-to-end Retrieval-Generation for Task-Oriented Dialog. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue

  139. [148]

    Few-shot Natural Language Generation for Task-Oriented Dialog

    Peng, Baolin and Zhu, Chenguang and Li, Chunyuan and Li, Xiujun and Li, Jinchao and Zeng, Michael and Gao, Jianfeng. Few-shot Natural Language Generation for Task-Oriented Dialog. Findings of the Association for Computational Linguistics: EMNLP 2020

  140. [149]

    A simple language model for task-oriented dialogue

    Hosseini-Asl, Ehsan and McCann, Bryan and Wu, Chien-Sheng and Yavuz, Semih and Socher, Richard. A simple language model for task-oriented dialogue. Advances in neural information processing systems

  141. [150]

    Flamingo: a visual language model for few-shot learning

    Alayrac, Jean-Baptiste and Donahue, Jeff and Luc, Pauline and Miech, Antoine and Barr, Iain and Hasson, Yana and Lenc, Karel and Mensch, Arthur and Millican, Katherine and Reynolds, Malcolm and Others. Flamingo: a visual language model for few-shot learning. Advances in neural...

  142. [151]

    Task-oriented dialogue as dataflow synthesis

    Andreas, Jacob and Bufe, John and Burkett, David and Chen, Charles and Clausman, Josh and Crawford, Jean and Crim, Kate and DeLoach, Jordan and Dorner, Leah and Eisner, Jason and Others. Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computat...

  143. [152]

    Abstract meaning representation for sembanking

    Banarescu, Laura and Bonial, Claire and Cai, Shu and Georgescu, Madalina and Griffitt, Kira and Hermjakob, Ulf and Knight, Kevin and Koehn, Philipp and Palmer, Martha and Schneider, Nathan. Abstract meaning representation for sembanking. Proceedings of the 7th linguistic annot...

  144. [153]

    Logical reasoning for task oriented dialogue systems

    Beygi, Sajjad and Fazel-Zarandi, Maryam and Cervone, Alessandra and Krishnan, Prakash and Jonnalagadda, Siddhartha Reddy. Logical reasoning for task oriented dialogue systems. arXiv preprint arXiv:2202. 04161

  145. [154]

    Language models for human-robot interaction

    Billing, Erik and Ros\' e n, Julia and Lamb, Maurice. Language models for human-robot interaction. ACM/IEEE International Conference on Human-Robot Interaction, March 13--16, 2023, Stockholm, Sweden

  146. [155]

    Dialogue-amr: abstract meaning representation for dialogue

    Bonial, Claire and Donatelli, Lucia and Abrams, Mitchell and Lukin, Stephanie and Tratz, Stephen and Marge, Matthew and Artstein, Ron and Traum, David and Voss, Clare. Dialogue-amr: abstract meaning representation for dialogue. Proceedings of the 12th Language Resources and Ev...

  147. [156]

    Knowledge-and ambiguity-aware robot learning from corrective and evaluative feedback

    Celemin, Carlos and Kober, Jens. Knowledge-and ambiguity-aware robot learning from corrective and evaluative feedback. Neural computing & applications

  148. [157]

    Large language models are few (1)-shot table reasoners

    Chen, Wenhu. Large language models are few (1)-shot table reasoners. arXiv preprint arXiv:2210. 06710

  149. [158]

    Explainable Conversational Question Answering over Heterogeneous Sources via Iterative Graph Neural Networks

    Christmann, Philipp and Saha Roy, Rishiraj and Weikum, Gerhard. Explainable Conversational Question Answering over Heterogeneous Sources via Iterative Graph Neural Networks. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

  150. [159]

    Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding

    Dziri, Nouha and Madotto, Andrea and Zaiane, Osmar R and Bose, Avishek Joey. Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

  151. [160]

    Space efficient context encoding for non-task-oriented dialogue generation with graph attention transformer

    Galetzka, Fabian and Rose, Jewgeni and Schlangen, David and Lehmann, Jens. Space efficient context encoding for non-task-oriented dialogue generation with graph attention transformer. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and t...

  152. [161]

    Generalization and representational limits of graph neural networks

    Garg, Vikas and Jegelka, Stefanie and Jaakkola, Tommi. Generalization and representational limits of graph neural networks. International Conference on Machine Learning

  153. [162]

    Human--robot interaction: a survey

    Goodrich, Michael A and Schultz, Alan C and Others. Human--robot interaction: a survey. Foundations and Trends in Human--Computer Interaction

  154. [163]

    Conversation graph: Data augmentation, training, and evaluation for non-deterministic dialogue management

    Gritta, Milan and Lampouras, Gerasimos and Iacobacci, Ignacio. Conversation graph: Data augmentation, training, and evaluation for non-deterministic dialogue management. Transactions of the Association for Computational Linguistics

  155. [164]

    Visual programming: Compositional visual reasoning without training

    Gupta, Tanmay and Kembhavi, Aniruddha. Visual programming: Compositional visual reasoning without training. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  156. [165]

    Inductive representation learning on large graphs

    Hamilton, Will and Ying, Zhitao and Leskovec, Jure. Inductive representation learning on large graphs. Advances in neural information processing systems

  157. [166]

    It's not what you do, it's how you do it: Grounding uncertainty for a simple robot

    Hough, Julian and Schlangen, David. It's not what you do, it's how you do it: Grounding uncertainty for a simple robot. Proceedings of the 2017 ACM/IEEE International Conference on Human-Robot Interaction

  158. [167]

    Inner Monologue: Embodied Reasoning through Planning with Language Models

    Huang, Wenlong and Xia, Fei and Xiao, Ted and Chan, Harris and Liang, Jacky and Florence, Pete and Zeng, Andy and Tompson, Jonathan and Mordatch, Igor and Chebotar, Yevgen and Others. Inner Monologue: Embodied Reasoning through Planning with Language Models. 6th Annual Confere...

  159. [168]

    Are LLMs All You Need for Task-Oriented Dialogue?

    Hude c ek, Vojt e ch and Du s ek, Ond r ej. Are LLMs All You Need for Task-Oriented Dialogue?. arXiv preprint arXiv:2304. 06556

  160. [169]

    A data source for reasoning embodied agents

    Lanchantin, Jack and Sukhbaatar, Sainbayar and Synnaeve, Gabriel and Sun, Yuxuan and Srinet, Kavya and Szlam, Arthur. A data source for reasoning embodied agents. Proceedings of the AAAI Conference on Artificial Intelligence

  161. [170]

    Embodied semantic scene graph generation

    Li, Xinghang and Guo, Di and Liu, Huaping and Sun, Fuchun. Embodied semantic scene graph generation. Conference on Robot Learning

  162. [171]

    CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers

    Li, Shiyang and Yavuz, Semih and Hashimoto, Kazuma and Li, Jia and Niu, Tong and Rajani, Nazneen and Yan, Xifeng and Zhou, Yingbo and Xiong, Caiming. CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers. International Conference on Learning Representations

  163. [172]

    Fusing topology contexts and logical rules in language models for knowledge graph completion

    Lin, Qika and Mao, Rui and Liu, Jun and Xu, Fangzhi and Cambria, Erik. Fusing topology contexts and logical rules in language models for knowledge graph completion. An international journal on information fusion

  164. [173]

    Interactive language: Talking to robots in real time

    Lynch, Corey and Wahid, Ayzaan and Tompson, Jonathan and Ding, Tianli and Betker, James and Baruch, Robert and Armstrong, Travis and Florence, Pete. Interactive language: Talking to robots in real time. IEEE Robotics and Automation Letters

  165. [174]

    Augmented Language Models: a Survey

    Mialon, Gr\' e goire and Dessi, Roberto and Lomeli, Maria and Nalmpantis, Christoforos and Pasunuru, Ramakanth and Raileanu, Roberta and Roziere, Baptiste and Schick, Timo and Dwivedi-Yu, Jane and Celikyilmaz, Asli and Grave, Edouard and LeCun, Yann and Scialom, Thomas. Augmen...

  166. [175]

    Embodiment, situatedness, and morphology for humanoid robots interacting with people

    Miller, Blanca and Feil-Seifer, David. Embodiment, situatedness, and morphology for humanoid robots interacting with people. Humanoid Robotics: A Reference

  167. [176]

    Simple Open-Vocabulary Object Detection with Vision Transformers

    Minderer, Matthias and Gritsenko, Alexey and Stone, Austin and Neumann, Maxim and Weissenborn, Dirk and Dosovitskiy, Alexey and Mahendran, Aravindh and Arnab, Anurag and Dehghani, Mostafa and Shen, Zhuoran and Wang, Xiao and Zhai, Xiaohua and Kipf, Thomas and Houlsby, Neil. Si...

  168. [177]

    GPT-4 Technical Report

    OpenAI. GPT-4 Technical Report. arXiv [cs.CL]

  169. [178]

    Optuna: A Next-generation Hyperparameter Optimization Framework

    Akiba, Takuya and Sano, Shotaro and Yanase, Toshihiko and Ohta, Takeru and Koyama, Masanori. Optuna: A Next-generation Hyperparameter Optimization Framework. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

  170. [179]

    Towards meaningful, grounded conversations with intelligent agents

    Papangelis, Alexandros and Ultes, Stefan. Towards meaningful, grounded conversations with intelligent agents. arXiv preprint arXiv:2006. 15768

  171. [180]

    Prediction and embodiment in dialogue

    Pickering, Martin J and Garrod, Simon. Prediction and embodiment in dialogue. European journal of social psychology

  172. [181]

    Bayesian networks for spoken dialogue management in multimodal systems of tour-guide robots

    Prodanov, Plamen and Drygajlo, Andrzej. Bayesian networks for spoken dialogue management in multimodal systems of tour-guide robots. Proceedings of the 8th European Conference on Speech Communication and Technology (Eurospeech)

  173. [182]

    Robotic information gathering using semantic language instructions

    Rankin, Ian C and McCammon, Seth and Hollinger, Geoffrey A. Robotic information gathering using semantic language instructions. 2021 IEEE International Conference on Robotics and Automation (ICRA)

  174. [183]

    Parallel Context Windows for Large Language Models

    Ratner, Nir and Levine, Yoav and Belinkov, Yonatan and Ram, Ori and Magar, Inbal and Abend, Omri and Karpas, Ehud and Shashua, Amnon and Leyton-Brown, Kevin and Shoham, Yoav. Parallel Context Windows for Large Language Models. Proceedings of the 61st Annual Meeting of the Asso...

  175. [184]

    Exploring the Relationship between LLM Hallucinations and Prompt Linguistic Nuances: Readability, Formality, and Concreteness

    Rawte, Vipula and Priya, Prachi and Tonmoy, S M and Zaman, S M and Sheth, Amit and Das, Amitava. Exploring the Relationship between LLM Hallucinations and Prompt Linguistic Nuances: Readability, Formality, and Concreteness. arXiv preprint arXiv:2309. 11064

  176. [185]

    z ej and Levine, Sergey and Others

    Shah, Dhruv and Osi\' n ski, B a\. z ej and Levine, Sergey and Others. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. Conference on Robot Learning

  177. [186]

    Large language models can be easily distracted by irrelevant context

    Shi, Freda and Chen, Xinyun and Misra, Kanishka and Scales, Nathan and Dohan, David and Chi, Ed H and Sch \" a rli, Nathanael and Zhou, Denny. Large language models can be easily distracted by irrelevant context. International Conference on Machine Learning

  178. [187]

    Llm-planner: Few-shot grounded planning for embodied agents with large language models

    Song, Chan Hee and Wu, Jiaman and Washington, Clayton and Sadler, Brian M and Chao, Wei-Lun and Su, Yu. Llm-planner: Few-shot grounded planning for embodied agents with large language models. arXiv preprint arXiv:2212. 04088

  179. [188]

    Link-Context Learning for Multimodal LLMs

    Tai, Yan and Fan, Weichen and Zhang, Zhao and Zhu, Feng and Zhao, Rui and Liu, Ziwei. Link-Context Learning for Multimodal LLMs. arXiv preprint arXiv:2308. 07891

  180. [189]

    Robots that use language

    Tellex, Stefanie and Gopalan, Nakul and Kress-Gazit, Hadas and Matuszek, Cynthia. Robots that use language. Annual Review of Control, Robotics, and Autonomous Systems

  181. [190]

    Bayesian update of dialogue state for robust dialogue systems

    Thomson, Blaise and Schatzmann, Jost and Young, Steve. Bayesian update of dialogue state for robust dialogue systems. 2008 IEEE International Conference on Acoustics, Speech and Signal Processing

  182. [191]

    Emergent abilities of large language models

    Wei, Jason and Tay, Yi and Bommasani, Rishi and Raffel, Colin and Zoph, Barret and Borgeaud, Sebastian and Yogatama, Dani and Bosma, Maarten and Zhou, Denny and Metzler, Donald and Others. Emergent abilities of large language models. arXiv preprint arXiv:2206. 07682

  183. [192]

    Towards Large-Scale Interpretable Knowledge Graph Reasoning for Dialogue Systems

    Tuan, Yi-Lin and Beygi, Sajjad and Fazel-Zarandi, Maryam and Gao, Qiaozi and Cervone, Alessandra and Wang, William Yang. Towards Large-Scale Interpretable Knowledge Graph Reasoning for Dialogue Systems. Findings of the Association for Computational Linguistics: ACL 2022. doi:1...

  184. [193]

    o rn and Steidl, Stefan and Batliner, Anton and Burkhardt, Felix and Devillers, Laurence and M \

    Schuller, Bj \" o rn and Steidl, Stefan and Batliner, Anton and Burkhardt, Felix and Devillers, Laurence and M \" u Ller, Christian and Narayanan, Shrikanth. Paralinguistics in speech and language--state-of-the-art and the challenge. Computer speech & language

  185. [194]

    A survey on retrieval-augmented text generation

    Li, Huayang and Su, Yixuan and Cai, Deng and Wang, Yan and Liu, Lemao. A survey on retrieval-augmented text generation. arXiv preprint arXiv:2202. 01110

  186. [195]

    Algorithms for hyper-parameter optimization

    Bergstra, James and Bardenet, R\' e mi and Bengio, Yoshua and K\' e gl, Bal\' a zs. Algorithms for hyper-parameter optimization. Advances in neural information processing systems

  187. [196]

    LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

    Dettmers, Tim and Lewis, Mike and Belkada, Younes and Zettlemoyer, Luke. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. arXiv preprint arXiv:2208. 07339

  188. [197]

    8-bit Optimizers via Block-wise Quantization

    Dettmers, Tim and Lewis, Mike and Shleifer, Sam and Zettlemoyer, Luke. 8-bit Optimizers via Block-wise Quantization. 9th International Conference on Learning Representations, ICLR

  189. [198]

    Llama: Open and efficient foundation language models

    Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth\' e e and Rozi\` e re, Baptiste and Goyal, Naman and Hambro, Eric and Azhar, Faisal and Others. Llama: Open and efficient foundation language models. arXiv p...

  190. [199]

    Large Language Models Still Can't Plan (A Benchmark for LLMs on Planning and Reasoning about Change)

    Valmeekam, Karthik and Olmo, Alberto and Sreedharan, Sarath and Kambhampati, Subbarao. Large Language Models Still Can't Plan (A Benchmark for LLMs on Planning and Reasoning about Change). arXiv preprint arXiv:2206. 10498

  191. [200]

    Attention is all you need

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia. Attention is all you need. Advances in neural information processing systems

  192. [201]

    The role of physical embodiment in human-robot interaction

    Wainer, Joshua and Feil-Seifer, David J and Shell, Dylan A and Mataric, Maja J. The role of physical embodiment in human-robot interaction. ROMAN 2006-The 15th IEEE International Symposium on Robot and Human Interactive Communication

  193. [202]

    Dialogue Management as Graph Transformations

    Walker, Nicholas Thomas and Dahl, Torbj rn and Lison, Pierre. Dialogue Management as Graph Transformations. Conversational AI for Natural Human-Centric Interaction. doi:10.1007/978-981-19-5538-9\_15

  194. [203]

    Algorithms for the reduction of the number of points required to represent a digitized line or its caricature

    Douglas, David H and Peucker, Thomas K. Algorithms for the reduction of the number of points required to represent a digitized line or its caricature. Cartographica: the international journal for geographic information and geovisualization

  195. [204]

    The dialog state tracking challenge series: A review

    Williams, Jason D and Raux, Antoine and Henderson, Matthew. The dialog state tracking challenge series: A review. Dialogue & Discourse

  196. [205]

    Towards Universal Dialogue State Tracking

    Ren, Liliang and Xie, Kaige and Chen, Lu and Yu, Kai. Towards Universal Dialogue State Tracking. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing

  197. [206]

    Improving dialogue response generation via knowledge graph filter

    Wang, Yanmeng and Wang, Ye and Lou, Xingyu and Rong, Wenge and Hao, Zhenghong and Wang, Shaojun. Improving dialogue response generation via knowledge graph filter. ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

  198. [207]

    Conversational AI and knowledge graphs for social robot interaction

    Wilcock, Graham and Jokinen, Kristiina. Conversational AI and knowledge graphs for social robot interaction. 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI)

  199. [208]

    Individual Comparisons by Ranking Methods

    Wilcoxon, Frank. Individual Comparisons by Ranking Methods. Biometrics Bulletin

  200. [209]

    Relational Temporal Graph Reasoning for Dual-task Dialogue Language Understanding

    Xing, Bowen and Tsang, Ivor W. Relational Temporal Graph Reasoning for Dual-task Dialogue Language Understanding. IEEE transactions on pattern analysis and machine intelligence

  201. [210]

    Sql-to-text generation with graph-to-sequence model

    Xu, Kun and Wu, Lingfei and Wang, Zhiguo and Feng, Yansong and Sheinin, Vadim. Sql-to-text generation with graph-to-sequence model. arXiv preprint arXiv:1809. 05255

  202. [211]

    A survey of scene graph: Generation and application

    Xu, Pengfei and Chang, Xiaojun and Guo, Ling and Huang, Po-Yao and Chen, Xiaojiang and Hauptmann, Alexander G. A survey of scene graph: Generation and application. IEEE Trans. Neural Netw. Learn. Syst

  203. [212]

    Evaluating human-robot interaction

    Young, James E and Sung, Jayoung and Voida, Amy and Sharlin, Ehud and Igarashi, Takeo and Christensen, Henrik I and Grinter, Rebecca E. Evaluating human-robot interaction. International Journal of Social Robotics

  204. [213]

    Program enhanced fact verification with verbalization and graph attention network

    Yang, Xiaoyu and Nie, Feng and Feng, Yufei and Liu, Quan and Chen, Zhigang and Zhu, Xiaodan. Program enhanced fact verification with verbalization and graph attention network. arXiv preprint arXiv:2010. 03084

  205. [214]

    Pomdp-based statistical spoken dialog systems: A review

    Young, Steve and Ga s i\' c , Milica and Thomson, Blaise and Williams, Jason D. Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE

  206. [215]

    Socratic models: Composing zero-shot multimodal reasoning with language

    Zeng, Andy and Attarian, Maria and Ichter, Brian and Choromanski, Krzysztof and Wong, Adrian and Welker, Stefan and Tombari, Federico and Purohit, Aveek and Ryoo, Michael and Sindhwani, Vikas and Others. Socratic models: Composing zero-shot multimodal reasoning with language. ...

  207. [216]

    Chat with the environment: Interactive multimodal perception using large language models

    Zhao, Xufeng and Li, Mengdi and Weber, Cornelius and Hafez, Muhammad Burhan and Wermter, Stefan. Chat with the environment: Interactive multimodal perception using large language models. arXiv preprint arXiv:2303. 08268

  208. [217]

    Memory-Augmented Dialogue Management for Task-Oriented Dialogue Systems

    Zhang, Zheng and Huang, Minlie and Zhao, Zhongzhou and Ji, Feng and Chen, Haiqing and Zhu, Xiaoyan. Memory-Augmented Dialogue Management for Task-Oriented Dialogue Systems. ACM Transactions on Information Systems

  209. [218]

    Topic scene graphs for image captioning

    Zhang, Min and Chen, Jingxiang and Li, Pengfei and Jiang, Ming and Zhou, Zhe. Topic scene graphs for image captioning. IET Computer Vision

  210. [219]

    Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

    Zhang, Yue and Li, Yafu and Cui, Leyang and Cai, Deng and Liu, Lemao and Fu, Tingchen and Huang, Xinting and Zhao, Enbo and Zhang, Yu and Chen, Yulong and Others. Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv preprint arXiv:2309. 01219

  211. [220]

    Neural Text Sanitization with Explicit Measures of Privacy Risk

    Papadopoulou, Anthi and Yu, Yunhao and Lison, Pierre and vrelid, Lilja. Neural Text Sanitization with Explicit Measures of Privacy Risk. Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Jo...

  212. [221]

    ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

    Clark, Kevin and Luong, Minh-Thang and Le, Quoc V and Manning, Christopher D. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020

  213. [222]

    A unified framework for evaluating the risk of re-identification of text de-identification tools

    Scaiano, Martin and Middleton, Grant and Arbuckle, Luk and Kolhatkar, Varada and Peyton, Liam and Dowling, Moira and Gipson, Debbie S and El Emam, Khaled. A unified framework for evaluating the risk of re-identification of text de-identification tools. Journal of biomedical in...

  214. [223]

    Salience and attention in surprisal-based accounts of language processing

    Zarcone, Alessandra and Van Schijndel, Marten and Vogels, Jorrig and Demberg, Vera. Salience and attention in surprisal-based accounts of language processing. Frontiers in psychology

  215. [224]

    Automatic Evaluation of Disclosure Risks of Text Anonymization Methods

    Manzanares-Salor, Benet and S\' a nchez, David and Lison, Pierre. Automatic Evaluation of Disclosure Risks of Text Anonymization Methods. Privacy in Statistical Databases: International Conference, PSD 2022, Paris, France, September 21--23, 2022, Proceedings. doi:10.1007/978-3...

  216. [225]

    Survey and evaluation of web search engine hit counts as research tools in computational linguistics

    S\' a nchez, David and Mart\' nez-Sanahuja, Laura and Batet, Montserrat. Survey and evaluation of web search engine hit counts as research tools in computational linguistics. Information systems

  217. [226]

    AutoML : A survey of the state-of-the-art

    He, Xin and Zhao, Kaiyong and Chu, Xiaowen. AutoML : A survey of the state-of-the-art. Knowledge-Based Systems. doi:10.1016/j.knosys.2020.106622

  218. [227]

    Data minimization for GDPR compliance in machine learning models

    Goldsteen, Abigail and Ezov, Gilad and Shmelkin, Ron and Moffie, Micha and Farkash, Ariel. Data minimization for GDPR compliance in machine learning models. AI and Ethics. doi:10.1007/s43681-021-00095-8

  219. [228]

    Generation of Replacement Options in Text Sanitization

    Olstad, Annika Willoch and Papadopoulou, Anthi and Lison, Pierre. Generation of Replacement Options in Text Sanitization. Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa)

  220. [229]

    Wikidata: A large-scale collaborative ontological medical database

    Turki, Houcemeddine and Shafee, Thomas and Hadj Taieb, Mohamed Ali and Ben Aouicha, Mohamed and Vrande c i\' c , Denny and Das, Diptanshu and Hamdi, Helmi. Wikidata: A large-scale collaborative ontological medical database. Journal of biomedical informatics. doi:10.1016/j.jbi....

  221. [230]

    Neural Text Generation from Structured Data with Application to the Biography Domain

    Lebret, R\' e mi and Grangier, David and Auli, Michael. Neural Text Generation from Structured Data with Application to the Biography Domain. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/D16-1128

  222. [231]

    Privacy- and Utility-Preserving NLP with Anonymized data: A case study of Pseudonymization

    Yermilov, Oleksandr and Raheja, Vipul and Chernodub, Artem. Privacy- and Utility-Preserving NLP with Anonymized data: A case study of Pseudonymization. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023)

  223. [232]

    AGORA: An intelligent system for the anonymization, information extraction and automatic mapping of sensitive documents

    Juez-Hernandez, Rodrigo and Quijano-S\' a nchez, Lara and Liberatore, Federico and G\' o mez, Jes\' u s. AGORA: An intelligent system for the anonymization, information extraction and automatic mapping of sensitive documents. Applied soft computing. doi:10.1016/j.asoc.2023.110540

  224. [233]

    INCOGNITUS : A Toolbox for Automated Clinical Notes Anonymization

    Ribeiro, Bruno and Rolla, Vitor and Santos, Ricardo. INCOGNITUS : A Toolbox for Automated Clinical Notes Anonymization. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations

  225. [234]

    Anonymisation Models for Text Data: State of the art, Challenges and Future Directions

    Lison, Pierre and Pil\' a n, Ildik\' o and S\' a nchez, David and Batet, Montserrat and vrelid, Lilja. Anonymisation Models for Text Data: State of the art, Challenges and Future Directions. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistic...

  226. [235]

    DeID-GPT : Zero-shot medical text de-identification by GPT -4

    Liu, Zhengliang and Yu, Xiaowei and Zhang, Lu and Wu, Zihao and Cao, Chao and Dai, Haixing and Zhao, Lin and Liu, Wei and Shen, Dinggang and Li, Quanzheng and Others. DeID-GPT : Zero-shot medical text de-identification by GPT -4. arXiv preprint arXiv:2303. 11032

  227. [236]

    The S tanford C ore NLP Natural Language Processing Toolkit

    Manning, Christopher and Surdeanu, Mihai and Bauer, John and Finkel, Jenny and Bethard, Steven and McClosky, David. The S tanford C ore NLP Natural Language Processing Toolkit. Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstr...

  228. [237]

    Introduction to the C o NLL -2003 Shared Task: Language-Independent Named Entity Recognition

    Tjong Kim Sang, Erik F and De Meulder, Fien. Introduction to the C o NLL -2003 Shared Task: Language-Independent Named Entity Recognition. Proceedings of the Seventh Conference on Natural Language Learning at HLT - NAACL 2003

  229. [238]

    Visualizing and Understanding Neural Models in NLP

    Li, Jiwei and Chen, Xinlei and Hovy, Eduard and Jurafsky, Dan. Visualizing and Understanding Neural Models in NLP. Proceedings of the 2016 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies. doi:10.18653/v1/N16-1082

  230. [239]

    Understanding Neural Networks through Representation Erasure

    Li, Jiwei and Monroe, Will and Jurafsky, Dan. Understanding Neural Networks through Representation Erasure. arXiv preprint arXiv:1612. 08220

  231. [240]

    Saliency-driven Word Alignment Interpretation for Neural Machine Translation

    Ding, Shuoyang and Xu, Hainan and Koehn, Philipp. Saliency-driven Word Alignment Interpretation for Neural Machine Translation. Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers). doi:10.18653/v1/W19-5201

  232. [241]

    u tt, Kristof T and D \

    Kindermans, Pieter-Jan and Hooker, Sara and Adebayo, Julius and Alber, Maximilian and Sch \" u tt, Kristof T and D \" a hne, Sven and Erhan, Dumitru and Kim, Been. The (Un)reliability of Saliency Methods. Explainable AI: Interpreting, Explaining and Visualizing Deep Learning. ...

  233. [242]

    Is Attention Interpretable?

    Serrano, Sofia and Smith, Noah A. Is Attention Interpretable?. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. doi:10.18653/v1/P19-1282

  234. [243]

    A Machine Learning Based System for Semi-Automatically Redacting Documents

    Cumby, Chad and Ghani, Rayid. A Machine Learning Based System for Semi-Automatically Redacting Documents. Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence. doi:10.1609/aaai.v25i2.18851

  235. [244]

    Differentially-Private Text Generation via Text Preprocessing to Reduce Utility Loss

    Sasada, Taisho and Kawai, Masataka and Taenaka, Yuzo and Fall, Doudou and Kadobayashi, Youki. Differentially-Private Text Generation via Text Preprocessing to Reduce Utility Loss. 2021 International Conference on Artificial Intelligence in Information and Communication (ICAIIC...

  236. [245]

    DP-BART for Privatized Text Rewriting under Local Differential Privacy

    Igamberdiev, Timour and Habernal, Ivan. DP-BART for Privatized Text Rewriting under Local Differential Privacy. Findings of the Association for Computational Linguistics: ACL 2023

  237. [246]

    Large Language Models Can Be Easily Distracted by Irrelevant Context

    Shi, Freda and Chen, Xinyun and Misra, Kanishka and Scales, Nathan and Dohan, David and Chi, Ed and Sch \" a rli, Nathanael and Zhou, Denny. Large Language Models Can Be Easily Distracted by Irrelevant Context. arXiv [cs.CL]

  238. [247]

    Leveraging synonymy and polysemy to improve semantic similarity assessments based on intrinsic information content

    Batet, Montserrat and S\' a nchez, David. Leveraging synonymy and polysemy to improve semantic similarity assessments based on intrinsic information content. Artificial Intelligence Review

  239. [248]

    Anaphora and coreference resolution: A review

    Sukthanker, Rhea and Poria, Soujanya and Cambria, Erik and Thirunavukarasu, Ramkumar. Anaphora and coreference resolution: A review. An international journal on information fusion

  240. [249]

    Using information content to evaluate semantic similarity in a taxonomy

    Resnik, Philip. Using information content to evaluate semantic similarity in a taxonomy. Proceedings of the 14th international joint conference on Artificial intelligence (IJCAI'95)

  241. [250]

    Web-Based Inference Detection

    Staddon, Jessica and Golle, Philippe and Zimny, Bryce. Web-Based Inference Detection. USENIX Security Symposium

  242. [251]

    The effect of the general data protection regulation on medical research

    Rumbold, John Mark Michael and Pierscionek, Barbara. The effect of the general data protection regulation on medical research. Journal of medical Internet research

  243. [252]

    Disruptive and avoidable: GDPR challenges to secondary research uses of data

    Peloquin, David and DiMaio, Michael and Bierer, Barbara and Barnes, Mark. Disruptive and avoidable: GDPR challenges to secondary research uses of data. European journal of human genetics: EJHG

  244. [253]

    The value of protecting privacy

    Santanen, Eric. The value of protecting privacy. Business horizons. doi:10.1016/j.bushor.2018.04.004

  245. [254]

    Custom NLP Approaches to Data Anonymization

    Mendels, Omri. Custom NLP Approaches to Data Anonymization. Towards data science

  246. [255]

    Approaches of anonymisation of an SMS corpus

    Patel, Namrata and Accorsi, Pierre and Inkpen, Diana and Lopez, C\' e dric and Roche, Mathieu. Approaches of anonymisation of an SMS corpus. Proceedings of the International Conference on Intelligent Text Processing and Computational Linguistics

  247. [256]

    Comment on 'Unique in the shopping mall: On the reidentifiability of credit card metadata'

    S\' a nchez, David and Mart\' nez, Sergio and Domingo-Ferrer, Josep. Comment on 'Unique in the shopping mall: On the reidentifiability of credit card metadata'. Science

  248. [257]

    Minimizing the disclosure risk of semantic correlations in document sanitization

    S\' a nchez, David and Batet, Montserrat and Viejo, Alexandre. Minimizing the disclosure risk of semantic correlations in document sanitization. Information sciences

  249. [258]

    TM-Score: A Misuseability Weight Measure for Textual Content

    Vartanian, Arik and Shabtai, Asaf. TM-Score: A Misuseability Weight Measure for Textual Content. IEEE Transactions on Information Forensics and Security

  250. [259]

    Text Classification for Data Loss Prevention

    Hart, Michael and Manadhata, Pratyusa and Johnson, Rob. Text Classification for Data Loss Prevention. Proceedings of the 11th Privacy Enhancing Technologies Symposium (PETS)

  251. [260]

    Document Sensitivity Classification for Data Leakage Prevention with Twitter-Based Document Embedding and Query Expansion

    Trieu, Lap Q and Tran, Trung-Nguyen and Tran, Mai-Khiem and Tran, Minh-Triet. Document Sensitivity Classification for Data Leakage Prevention with Twitter-Based Document Embedding and Query Expansion. Proceedings of the 13th International Conference on Computational Intelligen...

  252. [261]

    The Theory of Parsing, Translation and Compiling

    Aho, Alfred V and Ullman, Jeffrey D. The Theory of Parsing, Translation and Compiling

  253. [262]

    Uniqueness of simple demographics in the US population

    Sweeney, Latanya. Uniqueness of simple demographics in the US population

  254. [263]

    You Only Train Once: Loss-Conditional Training of Deep Networks

    Dosovitskiy, Alexey and Djolonga, Josip. You Only Train Once: Loss-Conditional Training of Deep Networks. 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020

  255. [264]

    Felix: Flexible text editing through tagging and insertion

    Mallinson, Jonathan and Severyn, Aliaksei and Malmi, Eric and Garrido, Guillermo. Felix: Flexible text editing through tagging and insertion. arXiv preprint arXiv:2003. 10687

  256. [265]

    On Syntactic Anonymity and Differential Privacy

    Clifton, Chris and Tassa, Tamir. On Syntactic Anonymity and Differential Privacy. Transactions on Data Privacy. Foundations and Technologies

  257. [266]

    Viewing the GDPR through a de-identification lens: a tool for compliance, clarification, and consistency

    Hintze, Mike. Viewing the GDPR through a de-identification lens: a tool for compliance, clarification, and consistency. International Data Privacy Law

  258. [267]

    Learning Differentially Private Recurrent Language Models

    McMahan, H Brendan and Ramage, Daniel and Talwar, Kunal and Zhang, Li. Learning Differentially Private Recurrent Language Models. arXiv:1710. 06963 [cs]

  259. [268]

    Using Clinical Natural Language Processing for Health Outcomes Research: Overview and Actionable Suggestions for Future Advances

    Velupillai, Sumithra and Suominen, Hanna and Liakata, Maria and Roberts, Angus and Shah, Anoop D and Morley, Katherine Irene and Osborn, David and Hayes, Joseph F and Stewart, Robert James and Downs, Johnny and Chapman, Wendy and Dutta, Rina. Using Clinical Natural Language Pr...

  260. [269]

    Generating Sentences by Editing Prototypes

    Guu, Kelvin and Hashimoto, Tatsunori B and Oren, Yonatan and Liang, Percy. Generating Sentences by Editing Prototypes. Transactions of the Association for Computational Linguistics

  261. [270]

    Natural language generation for electronic health records

    Lee, Scott H. Natural language generation for electronic health records. npj Digital Medicine. doi:10.1038/s41746-018-0070-0

  262. [271]

    Privacy, Anonymity, and Big Data in the Social Sciences

    Daries, Jon P and Reich, Justin and Waldo, Jim and Young, Elise M and Whittinghill, Jonathan and Ho, Andrew Dean and Seaton, Daniel Thomas and Chuang, Isaac. Privacy, Anonymity, and Big Data in the Social Sciences. Communications of the ACM. doi:10.1145/2643132

  263. [272]

    DBpedia - A Large-scale, Multilingual Knowledge Base Extracted from Wikipedia

    Lehmann, Jens and Isele, Robert and Jakob, Max and Jentzsch, Anja and Kontokostas, Dimitris and Mendes, Pablo N and Hellmann, Sebastian and Morsey, Mohamed and van Kleef, Patrick and Auer, S \" o ren and Bizer, Christian. DBpedia - A Large-scale, Multilingual Knowledge Base Ex...

  264. [273]

    Privacy and Freedom

    Westin, Alan F. Privacy and Freedom

  265. [274]

    Membership inference attacks against machine learning models

    Shokri, Reza and Stronati, Marco and Song, Congzheng and Shmatikov, Vitaly. Membership inference attacks against machine learning models. 2017 IEEE Symposium on Security and Privacy (SP)

  266. [275]

    When differential privacy meets NLP: The devil is in the detail

    Habernal, Ivan. When differential privacy meets NLP: The devil is in the detail. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing

  267. [276]

    What privacy is for

    Cohen, Julie E. What privacy is for. Harvard law review

  268. [277]

    Database Anonymization: Privacy Models, Data Utility, and Microaggregation-based Inter-model Connections

    Domingo-Ferrer, Josep and S\' a nchez, David and Soria-Comas, Jordi. Database Anonymization: Privacy Models, Data Utility, and Microaggregation-based Inter-model Connections

  269. [278]

    Seven types of privacy

    Finn, Rachel L and Wright, David and Friedewald, Michael. Seven types of privacy. European data protection: coming of age

  270. [279]

    The Rules of Redaction: Identify, Protect, Review (and Repeat)

    Bier, Eric A and Chow, Richard and Golle, Philippe and King, Tracy H and Staddon, J. The Rules of Redaction: Identify, Protect, Review (and Repeat). IEEE Security and Privacy Magazine

  271. [280]

    Protecting respondents' identities in microdata release

    Samarati, Pierangela. Protecting respondents' identities in microdata release. IEEE transactions on knowledge and data engineering

  272. [281]

    Recognizing textual entailment: Models and applications

    Dagan, Ido and Roth, Dan and Sammons, Mark and Zanzotto, Fabio Massimo. Recognizing textual entailment: Models and applications. Synthesis Lectures on Human Language Technologies

  273. [282]

    Adversarial Removal of Demographic Attributes from Text Data

    Elazar, Yanai and Goldberg, Yoav. Adversarial Removal of Demographic Attributes from Text Data. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/D18-1002

  274. [283]

    T ext H ide: Tackling Data Privacy in Language Understanding Tasks

    Huang, Yangsibo and Song, Zhao and Chen, Danqi and Li, Kai and Arora, Sanjeev. T ext H ide: Tackling Data Privacy in Language Understanding Tasks. Findings of the Association for Computational Linguistics: EMNLP 2020. doi:10.18653/v1/2020.findings-emnlp.123

  275. [284]

    Deep Reinforcement Learning-based Text Anonymization against Private-Attribute Inference

    Mosallanezhad, Ahmadreza and Beigi, Ghazaleh and Liu, Huan. Deep Reinforcement Learning-based Text Anonymization against Private-Attribute Inference. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conferen...

  276. [285]

    Utility-preserving privacy protection of nominal data sets via semantic rank swapping

    Rodr\' guez-Garc\' a, Mercedes and Batet, Montserrat and S\' a nchez, David. Utility-preserving privacy protection of nominal data sets via semantic rank swapping. An international journal on information fusion

  277. [286]

    Efficient Techniques for Document Sanitization

    Chakaravarthy, Venkatesan T and Gupta, Himanshu and Roy, Prasan and Mohania, Mukesh K. Efficient Techniques for Document Sanitization. Proceedings of the 17th ACM Conference on Information and Knowledge Management, CIKM 2008

  278. [287]

    Significance of Term Relationships on Anonymization

    Anandan, Balamurugan and Clifton, Chris. Significance of Term Relationships on Anonymization. Proceedings of the 2011 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology - Workshops, WI-IAT 2011

  279. [288]

    Detecting Privacy Leaks Using Corpus-Based Association Rules

    Chow, Richard and Golle, Philippe and Staddon, Jessica. Detecting Privacy Leaks Using Corpus-Based Association Rules. Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. doi:10.1145/1401890.1401997

  280. [289]

    The European Court of Human Rights and the protection of civil liberties: An overview

    Gearty, Conor A. The European Court of Human Rights and the protection of civil liberties: An overview. The Cambridge law journal

  281. [290]

    Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth Shared task Track 1

    Stubbs, Amber and Kotfila, Christopher and Uzuner, \" O zlem. Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth Shared task Track 1. Journal of biomedical informatics

  282. [291]

    De-identification of psychiatric intake records: Overview of 2016 CEGS N-GRID Shared Tasks Track 1

    Stubbs, Amber and Filannino, Michele and Uzuner, \" O zlem. De-identification of psychiatric intake records: Overview of 2016 CEGS N-GRID Shared Tasks Track 1. Journal of biomedical informatics

  283. [292]

    The UAB Informatics Institute and 2016 CEGS N-GRID de-identification shared task challenge

    Bui, Duy Duc An and Wyatt, Mathew and Cimino, James J. The UAB Informatics Institute and 2016 CEGS N-GRID de-identification shared task challenge. Journal of biomedical informatics

  284. [293]

    Customization scenarios for de-identification of clinical notes

    Hartman, Tzvika and Howell, Michael D and Dean, Jeff and Hoory, Shlomo and Slyper, Ronit and Laish, Itay and Gilon, Oren and Vainstein, Danny and Corrado, Greg and Chou, Katherine and Others. Customization scenarios for de-identification of clinical notes. BMC medical informat...

  285. [294]

    Privacy as a social good

    Kasper, Debbie V S. Privacy as a social good. Social thought & research

  286. [295]

    Obfuscating Gender in Social Media Writing

    Reddy, Sravana and Knight, Kevin. Obfuscating Gender in Social Media Writing. Proceedings of the First Workshop on NLP and Computational Social Science

  287. [296]

    Demographic Dialectal Variation in Social Media: A Case Study of A frican- A merican E nglish

    Blodgett, Su Lin and Green, Lisa and O'Connor, Brendan. Demographic Dialectal Variation in Social Media: A Case Study of A frican- A merican E nglish. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/D16-1120

  288. [297]

    Generation of Surrogates for De-Identification of Electronic Health Records

    Chen, Aipeng and Jonnagaddala, Jitendra and Nekkantti, Chandini and Liaw, Siaw-Teng. Generation of Surrogates for De-Identification of Electronic Health Records. MEDINFO 2019: Health and Wellbeing e-Networks for All - Proceedings of the 17th World Congress on Medical and Healt...

  289. [298]

    Can physicians recognize their own patients in de-identified notes?

    Meystre, St\' e phane and Shen, Shuying and Hofmann, Deborah and Gundlapalli, Adi. Can physicians recognize their own patients in de-identified notes?. Studies in health technology and informatics

  290. [299]

    Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text

    Carrell, David and Malin, Bradley and Aberdeen, John and Bayer, Samuel and Clark, Cheryl and Wellner, Ben and Hirschman, Lynette. Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text. Journal of the American Med...

  291. [300]

    A study of deep learning methods for de-identification of clinical notes in cross-institute settings

    Yang, Xi and Lyu, Tianchen and Li, Qian and Lee, Chih-Yin and Bian, Jiang and Hogan, William R and Wu, Yonghui. A study of deep learning methods for de-identification of clinical notes in cross-institute settings. BMC medical informatics and decision making

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.