Pith. sign in

REVIEW 4 major objections 6 minor 71 references

Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems

T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read ITSmark hides and fully recovers multi-bit watermarks in AI-generated traffic text, and only key-holders can extract them.

desk verdict ITSmark is a genuinely new multi-bit watermarking construction, but its central "entire extraction" guarantee rests on an unproven progress assumption (p=0 stalls) and no released code; worth refereeing for the core idea, not for the claims as stated. read the letter →

arxiv 2502.05425 v1 pith:MBM6SQNC submitted 2025-02-08 cs.CR cs.CL

classification cs.CRcs.CL
keywords multi-bitwatermarkinglargelanguagemodelsintelligenttransportationsystemscopyrightprotectiontamperlocalizationpermissionverificationtextextractionaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes ITSmark, a watermarking scheme that hides a custom multi-bit message, such as copyright information or a timestamp, inside text generated by a large language model, with the goal of protecting intelligent transportation data. The work tries to establish three capabilities that prior text watermarks lack: complete extraction of the embedded message, permission-gated extraction so that only authorized verifiers can recover it, and localization of tampered positions without access to the original document. If the scheme works as claimed, ITS departments and data-sharing platforms could verify the origin and integrity of AI-generated traffic reports and identify exactly which tokens were altered or forged. The authors report that ITSmark beats existing zero-bit and multi-bit baselines in data quality, extraction accuracy, and resistance to automatic detection, and that extraction succeeds completely in their experiments.

What carries the argument

The load-bearing object is the multi-bit space B, the set of all binary strings of length ε, divided into consecutive bit segments S. Each candidate token is assigned the segment whose length is proportional to its logit raised to a weight λ, and the token whose segment contains the current watermark chunk is chosen as the next token; at extraction, the same logits and the chosen token locate the segment, and the longest common prefix of the segment's lower and upper endpoints gives the embedded bits. This segment-prefix mechanism is what converts the one-to-many mapping from token back to watermark into a one-to-one reversible mapping, and it is what permits both complete extraction and tamper localization by rank.

What would settle it

Generate a large batch of watermarked texts with ITSmark on contexts where the model is highly confident (low-entropy next-token distributions) and compare the extracted watermark bit-by-bit with the embedded message; if any text shows an embedding step where the segment endpoints share no common prefix, the complete-extraction claim fails for that text.

Watch

Extended reading notes

Core claim

The central claim is that a multi-bit watermark can be embedded reversibly by making the next token's choice depend on the watermark: the copyright string is converted to bits, read a few bits at a time, and each chunk is located inside the space of all binary strings of that length, which is partitioned into contiguous segments weighted by the model's logits. The token whose segment contains the current chunk is generated, so the produced text literally contains the message. Because extraction replays the same logits and segment partition, the token reveals its segment, and the common prefix of that segment's endpoints recovers exactly the bits that were embedded at that step; the authors argue this inverse mapping makes embedding and extraction independent of the hyperparameters and gives complete extraction accuracy on the TV, BDD, AlpacaFarm, and FinQA datasets with LLaMA2-7B and ChatGLM3-6B. Encryption of the prompt and parameters into cipher data, plus a private key requirement, makes extraction fail for unauthorized users, and tokens whose likelihood rank falls outside the trusted top-k range are flagged as tampered.

Load-bearing premise

The scheme assumes that at every step of generation, the next piece of the watermark can be matched to the beginning of at least one allowed token's range, so no watermark bit ever has to be skipped.

Editorial extensions

If this is right

  • If the segment-prefix mechanism is correct, any LLM can carry ITSmark without retraining: the scheme needs only the base model's logits at generation and extraction time.
  • Authorized recipients can authenticate the origin of a traffic report and recover the full copyright string without needing the original unwatermarked text.
  • A failed full extraction after successful permission verification signals tampering, and the rank-based rule localizes the edited tokens within seconds.
  • Users can tune embedding load: smaller embedding ratios and larger λ preserve text quality, while smaller λ and larger ε increase payload, so the scheme adapts to short versus long copyright messages.
  • Because extraction failure is all-or-nothing without the key, ITSmark also acts as an access control layer on top of integrity checking.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the logit-based partition guarantees that every watermark chunk shares at least one leading bit with its assigned segment; if p=0 occurs for a required chunk, the chunk would be skipped and the advertised complete-extraction guarantee would fail, so a direct audit of p=0 frequency on low-entropy contexts would settle the matter.
  • The unforgeability result is measured as indistinguishability from unwatermarked text by a RoBERTa classifier, not as resistance to an adversary who knows the scheme; a stronger test would be adaptive forgery attempts with known parameters.
  • The tamper-tracing rule is a simple hand-set ranking heuristic; it could be turned into a calibrated probabilistic model by fitting the distribution of token ranks under substitution and rewrite attacks.
  • The same reversible multi-bit space idea transfers beyond ITS to any domain with LLM-generated text, such as legal or medical reporting, wherever provenance and edit localization matter.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents ITSmark, a multi-bit watermarking scheme for LLM-generated text in intelligent transportation systems. Copyright information is converted to a binary string and read in ε-bit windows; each window is located in a logit-weighted partition of the 2^ε bit space and the containing segment determines the next generated token. Extraction reverses this mapping: each observed token identifies its segment, and the common prefix of that segment's endpoints is recovered as the embedded bits. Extraction is gated by cipher data and a private key, and tamper localization is performed by a token-rank heuristic. Experiments against KGW and CTWL report improved perplexity, BERTScore, and ROUGE, 100% extraction success, unforgeability, permission-failure behavior, and ablations over η, λ, and ε.

Significance. The application is timely for T-ITS, and the construction has attractive properties: deterministic embedding/extraction if the partition and logits are fixed, no parameters fitted to produce the headline success-rate numbers, a built-in access-control layer, and a concrete tamper-tracing mechanism. The paper also scopes out strong robustness clearly. However, the distinguishing claim of entire extraction is not established: the scheme explicitly allows zero-bit progress, and no proof or stress test shows that the read pointer eventually reaches the end of the watermark. The empirical 100% extraction rate cannot be independently checked because no code or data are released. These gaps do not disprove the idea, but they leave the main advertised guarantee unsupported.

major comments (4)
  1. [§III.C, Eq. (14)–(15)] The load-bearing progress assumption is unproved. Equations (14)–(15) explicitly allow p=0, and the text states that in this case the match is an empty string and the next read is m_{1:ε}, so the watermark pointer does not advance. The paper never proves that a finite generated text embeds all a watermark bits; the Section V discussion of segment conflict only shows that the pointer does not jump backward, not that it eventually reaches the end. Because the abstract and Section IV.C advertise entire extraction, this is a central gap. A concrete remedy would be a construction or proof that every relevant ε-bit window lies in a segment with p ≥ 1, or a bound on the number of tokens needed to embed a bits. For reference, I do not think the variance-derivative objection to Eq. (13) lands: p_i and ln P_i are similarly ordered, so the derivative is indeed non-negative.
  2. [§IV.C, Table V] The claim that complete extraction is independent of λ, ε, and η is asserted without support and is inconsistent with the mechanism as written. The value of p is determined by segment endpoints, which depend on λ and the logit distribution; ε sets the size of the read window; and partial embedding with ratio η decides which sentences carry any bits at all. The 100.00% success rates in Table V are not accompanied by code, data, or an analysis of how often p=0 occurs, so the empirical result cannot be separated from the missing progress guarantee.
  3. [§III.D, Algorithm 1] The extraction algorithm lacks a termination condition and a rule for handling the final partial window. It outputs p bits per token, but it never specifies how the extractor knows the watermark length a, when to stop reading tokens, or what to do when p exceeds the number of remaining watermark bits. Without these details, "entire extraction" is not a well-defined procedure, especially if p=0 stalls can occur.
  4. [§III.B and §IV.F] There is a mismatch between the full-vocabulary description and the Top-k description. Algorithm 1 and Section III.B describe dividing the bit space according to the full vocabulary V, while Section IV.F states that the candidate pool is Top-k with k=40 and that the selected token must lie in that pool. If embedding and extraction use different token sets, the inverse property fails. The paper should state exactly which token set is used at both ends and prove that every token selected during embedding is also in the extraction set.
minor comments (6)
  1. [Table V caption] The caption contains a typo: "ESM ARK" should be "ITSmark."
  2. [Eq. (2)] The notation p(l, j, i) for segment length is easily confused with the common-prefix length p introduced later in Section III.C; consider renaming one of them.
  3. [Eqs. (7)–(13)] The symbol "In" should be "ln", and the phrase "this deviation is non-negative" should be "this derivative is non-negative."
  4. [Fig. 3] The labels "The beginning and end" are vague; annotating the segment endpoints β and β′ directly would make the figure self-contained.
  5. [§IV.A] Table II's presentation of dataset statistics is hard to read; the units (average tokens per text versus total tokens) should be stated explicitly in the table.
  6. [§IV.D] The unforgeability discussion would be clearer if actual ROC-AUC values were reported numerically rather than only through curves, since "smaller AUC" is otherwise ambiguous.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the multi-bit extraction guarantee is an exact-inverse construction property, and no headline claim reduces to a fitted parameter or to a self-citation chain.

full rationale

The paper's central claim is that an authorized user can entirely extract the embedded watermark. This is supported by the embedder/extractor design, not by a fitted parameter or by an imported self-cited theorem. In Section III.C, the embedder reads an epsilon-bit window, finds the segment containing it, and embeds only the common prefix p of the segment endpoints; Section III.D then has the extractor recompute the same logits and partition, locate the same segment from the token, and output the same prefix p. This is a deterministic reversible mapping, so the 100% success rate in Table V is a construction property of that inverse mapping rather than an empirically fitted prediction. The comparison metrics (perplexity, BERTScore, ROUGE) and ablations are measured externally against baselines or against the paper's own generated data; no parameter is fitted to the headline extraction-result and then renamed as a prediction. The paper's self-citations are background references and are not load-bearing for the watermark claim. The only substantive weakness is that Section III.C explicitly allows p=0 ('When p = 0, their match is an empty string'), and the paper does not prove that the reading position advances far enough to guarantee the full watermark is embedded before the text ends. That is an unproven completeness condition, not a circular reduction, and it does not affect the circularity verdict.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

Everything the central claims depend on is either standard in LLM watermarking or assumed without proof. The load-bearing assumptions are exact logit reproducibility at extraction, non-empty watermark progress at every token, and top-k containment. The ledger contains no invented physical entities. The main free parameters are hyperparameters lambda, epsilon, eta, and top-k, plus the hand-set tamper thresholds of Eq. 17.

free parameters (5)
  • lambda (logit weight) = 1.0 in main comparison; 0.4 to 1.5 in ablation
    Controls how much of the bit space is given to high-probability tokens; the quality and payload tradeoff depends on it.
  • epsilon (bits read per step) = 16 in main comparison; 8 to 20 in ablation
    Chunk size of watermark bits; the bit space has size 2^epsilon.
  • embedding ratio eta = 1.0 for full embedding; 0.1 to 0.9 for partial
    Fraction of sentences re-generated with the watermark; lower eta improves quality and reduces payload.
  • top-k candidate pool size = 40
    Defines the candidate token set and is used as the reference for tamper detection by probability rank.
  • tamper probability thresholds = 1, 0.75, 0.5, 0.3, 0 over rank ranges
    Hand-set scoring rule in Eq. 17 for traceability; not learned or independently justified.
assumptions (4)
  • domain assumption Same LLM and tokenizer reproduce identical logits at embedding and extraction.
    Algorithm 1 requires recalculating P(v_i|x_{1:t-1}) with the same system; no robustness to model updates or nondeterminism is shown.
  • ad hoc to paper The logit-based partition assigns a non-empty bit segment to every token that can be selected, and every epsilon-bit string lies in a segment with common prefix p greater than or equal to 1.
    Sections III.B and III.C define the partition but do not prove that embedding always makes progress; p=0 is explicitly allowed.
  • domain assumption Watermark-determined tokens always fall inside the top-k set used at extraction.
    Section IV.F bases tamper tracing on this property; if a watermarked token falls outside top-k, extraction of that token fails.
  • ad hoc to paper The derivative of variance in Eq. 13 is non-negative for all logit distributions when lambda is positive.
    Used to justify the lambda tradeoff discussion; the asserted non-negativity is not generally true and the sign can depend on the logit distribution.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems." pith.science (2026). https://pith.science/paper/MBM6SQNC

@misc{pith2026250205425,
  author       = {Pith},
  title        = {Pith review of: Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MBM6SQNC}},
  note         = {Machine review of arXiv:2502.05425}
}
read the original abstract

Intelligent transportation systems (ITS) use advanced technologies such as artificial intelligence to significantly improve traffic flow management efficiency, and promote the intelligent development of the transportation industry. However, if the data in ITS is attacked, such as tampering or forgery, it will endanger public safety and cause social losses. Therefore, this paper proposes a watermarking that can verify the integrity of copyright in response to the needs of ITS, termed ITSmark. ITSmark focuses on functions such as extracting watermarks, verifying permission, and tracing tampered locations. The scheme uses the copyright information to build the multi-bit space and divides this space into multiple segments. These segments will be assigned to tokens. Thus, the next token is determined by its segment which contains the copyright. In this way, the obtained data contains the custom watermark. To ensure the authorization, key parameters are encrypted during copyright embedding to obtain cipher data. Only by possessing the correct cipher data and private key, can the user entirely extract the watermark. Experiments show that ITSmark surpasses baseline performances in data quality, extraction accuracy, and unforgeability. It also shows unique capabilities of permission verification and tampered location tracing, which ensures the security of extraction and the reliability of copyright verification. Furthermore, ITSmark can also customize the watermark embedding position and proportion according to user needs, making embedding more flexible.

Figures

Figures reproduced from arXiv: 2502.05425 by the authors.

Figure 1
Figure 1. The performance of some SOTA schemes and the proposed ITSmark. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The embedding and extraction processes of ITSmark. “Pubkey” and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The working principle of the “Embedder” and “Extractor” in ITSmark. Here only a certain moment’s embedding and extraction processes are shown. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Comparison of ITSmark (Ours) and baselines regarding Perplexity. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Comparison of ITSmark (Ours) and baselines regarding BERTScore [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Comparison of ITSmark and baselines regarding the unforgeability in [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: The extraction results when the permission verification fails. The [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Fineness of tampered location tracing. The value range is [0%, 100%]. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Comparison of Perplexity under different embedding ratios [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 11
Figure 11. Figure 11: Comparison of data quality under different [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 60 canonical work pages

  1. [1]

    Transportation Internet: A Sustainable Solution for Intelligent Trans- portation Systems,

    Hui Li, Yongquan Chen, Keqiang Li, Chong Wang, and Bokui Chen, “Transportation Internet: A Sustainable Solution for Intelligent Trans- portation Systems,” IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 12, pp. 15818–15829, 2023

  2. [2]

    Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions,

    Khan Muhammad, Amin Ullah, Jaime Lloret, Javier Del Ser, and Victor Hugo C. de Albuquerque, “Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions,” IEEE Transactions on Intelligent Transportation Systems , vol. 22, no. 7, pp. 4316–4334, 2021

  3. [3]

    Transportation 5.0: The DAO to Safe, Secure, and Sustainable Intelligent Transportation Systems,

    Feiyue Wang, Yilun Lin, Petros A. Ioannou, Ljubo Vlacic, Xiaom- ing Liu, and Azim Eskandarian, “Transportation 5.0: The DAO to Safe, Secure, and Sustainable Intelligent Transportation Systems,” IEEE Transactions on Intelligent Transportation Systems , vol. 24, no. 10, pp. 10262–10278, 2023

  4. [4]

    Cooperative Incident Management in Mixed Traffic of CA Vs and Human-Driven Vehicles,

    Wenwei Yue, Changle Li, Shangbo Wang, Nan Xue, and Jiaming Wu, “Cooperative Incident Management in Mixed Traffic of CA Vs and Human-Driven Vehicles,” IEEE Transactions on Intelligent Transporta- tion Systems, vol. 24, no. 11, pp. 12462–12476, 2023

  5. [5]

    Advancing traffic safety through the safe system approach: A systematic review,

    Md Nasim Khan and Subasish Das, “Advancing traffic safety through the safe system approach: A systematic review,” Accident Analysis & Prevention, vol. 199, no. 107518, 2024

  6. [6]

    GPT-4 Tech- nical Report,

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, et al, “GPT-4 Tech- nical Report,” arXiv preprint arXiv:2303.08774 , 2023

  7. [7]

    Introducing Meta Llama 3: The most capable openly available LLM to date,

    Meta, “Introducing Meta Llama 3: The most capable openly available LLM to date,” available: https://ai.meta.com/blog/meta-llama-3/, 2024

  8. [8]

    Language Models are Unsupervised Multitask Learners,

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI blog, vol. 1, no. 8, pp. 9, 2019

Show all 71 references
  1. [9]

    Gemini: A Family of Highly Capable Multimodal Models,

    Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, et al, “Gemini: A Family of Highly Capable Multimodal Models,” arXiv preprint arXiv:2312.11805 , 2023

  2. [10]

    Advancing Gener- alizations of Multi-Scale GAN via Adversarial Perturbation Augmenta- tions,

    Jing Tang, Zeyu Gong, Bo Tao, and Zhouping Yin, “Advancing Gener- alizations of Multi-Scale GAN via Adversarial Perturbation Augmenta- tions,” Knowledge-Based Systems, V ol. 284, no. 111260, 2024

  3. [11]

    Multimodal Feature-Guided Pretraining for RGB-T Perception,

    Junlin Ouyang, Pengcheng Jin, and Qingwang Wang, “Multimodal Feature-Guided Pretraining for RGB-T Perception,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 17, pp. 16041–16050, 2024

  4. [12]

    Vision-Based Semantic Segmentation in Scene Understanding for Autonomous Driving: Recent Achievements, Challenges, and Outlooks,

    Khan Muhammad, Tanveer Hussain, Hayat Ullah, Javier Del Ser, Mahdi Rezaei, and Neeraj Kumar, “Vision-Based Semantic Segmentation in Scene Understanding for Autonomous Driving: Recent Achievements, Challenges, and Outlooks,” IEEE Transactions on Intelligent Trans- portation Sys...

  5. [13]

    An efficient quantum proactive incremental learning algorithm,

    Lingxiao Li, Jing Li, Yanqi Song, Sujuan Qin, Qiaoyan Wen, and Fei Gao, “An efficient quantum proactive incremental learning algorithm,” Science China Physics, Mechanics & Astronomy , vol. 68, no. 210313, 2025

  6. [14]

    SingleS2R: Single Sample Driven Sim-to-Real Transfer for Multi-Source Visual-Tactile Information Understanding using Multi-Scale Vision Transformers,

    Jing Tang, Zeyu Gong, Bo Tao, and Zhouping Yin, “SingleS2R: Single Sample Driven Sim-to-Real Transfer for Multi-Source Visual-Tactile Information Understanding using Multi-Scale Vision Transformers,” Information Fusion, V ol. 108, no. 102390, 2024

  7. [15]

    LLsM: Generative Linguistic Steganography with Large Language Model,

    Yihao Wang, Ruiqi Song, Ru Zhang, Jianyi Liu, and Lingxiao Li, “LLsM: Generative Linguistic Steganography with Large Language Model,” arXiv preprint arXiv:2401.15656 , 2024

  8. [16]

    Security of target recognition for UA V forestry remote sensing based on multi-source data fusion transformer framework,

    Hailin Feng, Qing Li, Wei Wang, Ali Kashif Bashir, Amit Kumar Singh, Jinshan Xu, and Kai Fang, “Security of target recognition for UA V forestry remote sensing based on multi-source data fusion transformer framework,” Information Fusion, vol. 112, no. 102555, 2024

  9. [17]

    V-A3tS: A rapid text steganalysis method based on position information and variable parameter multi-head self-attention controlled by length,

    Yihao Wang, Ru Zhang, Jianyi Liu, “V-A3tS: A rapid text steganalysis method based on position information and variable parameter multi-head self-attention controlled by length,” Journal of Information Security and Applications, vol. 75, no. 103512, 2023

  10. [18]

    A Survey on Intelligent In- ternet of Things: Applications, Security, Privacy, and Future Di- rections,

    Ons Aouedi, Thai-Hoc Vu, Alessio Sacco, Dinh C. Nguyen, Kan- daraj Piamrat, and Guido Marchetto, “A Survey on Intelligent In- ternet of Things: Applications, Security, Privacy, and Future Di- rections,” IEEE Communications Surveys & Tutorials , 2024. DOI: 10.1109/COMST.2024.3430368

  11. [19]

    Privacy-preserving adaptive traffic signal control in a connected vehicle environment,

    Chaopeng Tan and Kaidi Yang, “Privacy-preserving adaptive traffic signal control in a connected vehicle environment,” Transportation Research Part C: Emerging Technologies, vol. 158, no. 104453, 2024

  12. [20]

    Regulatory options for vehicle telematics devices: balancing driver safety, data privacy and data security,

    Jon Truby, Rafael Dean Brown, and Imad Antoine Ibrahim, “Regulatory options for vehicle telematics devices: balancing driver safety, data privacy and data security,” International Review of Law, Computers & Technology, vol. 38, no. 1, pp. 86–110, 2024

  13. [21]

    A Watermark for Large Language Mod- els,

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein, “A Watermark for Large Language Mod- els,” in Proceedings of the 40th International Conference on Machine Learning (ICML), pp. 17061–17084, 2023

  14. [22]

    Necessary and sufficient watermark for large language models,

    Yuki Takezawa, Ryoma Sato, Han Bao, Kenta Niwa, and Makoto Ya- mada, “Necessary and sufficient watermark for large language models,” arXiv preprint arXiv: 2310.00833 , 2023

  15. [23]

    REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models,

    Ruisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, and Farinaz Koushanfar, “REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models,” in Proceedings of the 33rd USENIX Security Symposium , 2024

  16. [24]

    A Survey of Text Watermarking in the Era of Large Language Models,

    Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Xi Zhang, Lijie Wen, Irwin King, Hui Xiong, and Philip S. Yu, “A Survey of Text Watermarking in the Era of Large Language Models,” ACM Computing Surveys, 2024

  17. [25]

    Adaptive Text Watermark for Large Language Models,

    Yepeng Liu and Yuheng Bu, “Adaptive Text Watermark for Large Language Models,” in Proceedings of the 41st International Conference on Machine Learning (ICML) , 2024

  18. [26]

    Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Water- mark for Large Language Models,

    Zhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, and Rui Wang, “Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Water- mark for Large Language Models,” in Proceedings of the 62nd Annual Meeting of the Associ...

  19. [27]

    A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models,

    Yihan Wu, Zhengmian Hu, Junfeng Guo, Hongyang Zhang, and Heng Huang, “A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models,” in Proceedings of the 41st International Conference on Machine Learning (ICML) , 2024

  20. [28]

    An Entropy-based Text Watermarking Detection Method,

    Yijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li, and Irwin King, “An Entropy-based Text Watermarking Detection Method,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL) , 2024

  21. [29]

    On the Reliability of Watermarks for Large Language Models,

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Gold- blum, and Tom Goldstein, “On the Reliability of Watermarks for Large Language Models,” in Proceedings of the Twelfth International Conference on Le...

  22. [30]

    Prov- able Robust Watermarking for AI-Generated Text,

    Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yuxiang Wang, “Prov- able Robust Watermarking for AI-Generated Text,” inProceedings of the Twelfth International Conference on Learning Representations (ICLR) , 2024

  23. [31]

    Undetectable Watermarks for Language Models,

    Miranda Christ, Sam Gunn, and Or Zamir, “Undetectable Watermarks for Language Models,” in Proceedings of the 37th Annual Conference on Learning Theory , vol. 196, pp. 1–15, 2024

  24. [32]

    An Unforgeable Publicly Verifiable Watermark for Large Language Models,

    Aiwei Liu, Leyi Pan, Xuming Hu, Shuang Li, Lijie Wen, Irwin King, and Philip S. Yu, “An Unforgeable Publicly Verifiable Watermark for Large Language Models,” in Proceedings of the Twelfth International Conference on Learning Representations (ICLR) , 2024

  25. [33]

    Provably Robust Multi-bit Watermarking for AI-generated Text via Error Correction Code,

    Wenjie Qu, Dong Yin, Zixin He, Wei Zou, Tianyang Tao, Jinyuan Jia, and Jiaheng Zhang, “Provably Robust Multi-bit Watermarking for AI-generated Text via Error Correction Code,” arXiv preprint arXiv:2401.16820, 2024

  26. [34]

    Towards Codable Watermarking for Injecting Multi-bits Information to LLMs,

    Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou, Yankai Lin, Fandong Meng, Jie Zhou, and Xu Sun, “Towards Codable Watermarking for Injecting Multi-bits Information to LLMs,” in Proceedings of the Twelfth International Conference on Learning Representations (ICLR) , 2024

  27. [35]

    Advancing Beyond Identification: Multi-bit Watermark for Large Language Models,

    KiYoon Yoo, Wonhyuk Ahn, and Nojun Kwak, “Advancing Beyond Identification: Multi-bit Watermark for Large Language Models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pp. 4031– 4055, 2024

  28. [36]

    Edge Intelligence in Intelligent Transportation Systems: A Survey,

    Taiyuan Gong, Li Zhu, F. Richard Yu, and Tao Tang, “Edge Intelligence in Intelligent Transportation Systems: A Survey,” IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 9, pp. 8919–8944, 2023

  29. [37]

    A review of 6G autonomous intelli- gent transportation systems: Mechanisms, applications and challenges,

    Xiaoheng Deng, Leilei Wang, Jinsong Gui, Ping Jiang, Xuechen Chen, Feng Zeng, and Shaohua Wan, “A review of 6G autonomous intelli- gent transportation systems: Mechanisms, applications and challenges,” Journal of Systems Architecture , vol. 142, no. 102929, 2023

  30. [38]

    Towards 5G- Enabled Self Adaptive Green and Reliable Communication in Intelligent Transportation System,

    Ali Hassan Sodhro, Sandeep Pirbhulal, Gul Hassan Sodhro, Muham- mad Muzammal, Luo Zongwei, and Andrei Gurtov, “Towards 5G- Enabled Self Adaptive Green and Reliable Communication in Intelligent Transportation System,”IEEE Transactions on Intelligent Transportation Systems, vol....

  31. [39]

    Big data algorithms and applications in intelligent transportation system: A review and IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS 14 bibliometric analysis,

    Sepideh Kaffash, An Truong Nguyen, and Joe Zhu, “Big data algorithms and applications in intelligent transportation system: A review and IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS 14 bibliometric analysis,” International Journal of Production Economics , vol. 231,...

  32. [40]

    A Taxonomy and Survey of Edge Cloud Computing for Intelligent Transportation Systems and Connected Vehi- cles,

    Peter Arthurs, Lee Gillam, Paul Krause, Ning Wang, Kaushik Halder, and Alexandros Mouzakitis, “A Taxonomy and Survey of Edge Cloud Computing for Intelligent Transportation Systems and Connected Vehi- cles,” IEEE Transactions on Intelligent Transportation Systems , vol. 23, no....

  33. [41]

    Traffic flow matrix-based graph neural network with attention mechanism for traffic flow prediction,

    Jian Chen, Li Zheng, Yuzhu Hu, Wei Wang, Hongxing Zhang, and Xiping Hu, “Traffic flow matrix-based graph neural network with attention mechanism for traffic flow prediction,” Information Fusion , vol. 104, no. 102146, 2024

  34. [42]

    Spatial-temporal graph convolution network model with traffic fundamental diagram information informed for network traffic flow prediction,

    Zhao Liu, Fan Ding, Yunqi Dai, Linchao Li, Tianyi Chen, and Huachun Tan, “Spatial-temporal graph convolution network model with traffic fundamental diagram information informed for network traffic flow prediction,” Expert Systems with Applications , vol. 249, no. 123543, 2024

  35. [43]

    Advancing Data-Driven Decision-Making in Smart Cities through Big Data Analytics: A Comprehensive Review of Existing Literature,

    Oluwaseun Oladeji Olaniyi, Olalekan J. Okunleye, and Samuel Oladiipo Olabanji, “Advancing Data-Driven Decision-Making in Smart Cities through Big Data Analytics: A Comprehensive Review of Existing Literature,” Current Journal of Applied Science and Technology , vol. 42, no. 25...

  36. [44]

    FD- TGCN: Fast and dynamic temporal graph convolution network for traffic flow prediction,

    Lijun Sun, Mingzhi Liu, Guanfeng Liu, Xiao Chen, and Xu Yu, “FD- TGCN: Fast and dynamic temporal graph convolution network for traffic flow prediction,” Information Fusion, vol. 106, no. 102291, 2024

  37. [45]

    Deep spatio-temporal 3D dilated dense neural network for traffic flow prediction,

    Rui He, Cuijuan Zhang, Yunpeng Xiao, Xingyu Lu, Song Zhang, and Yanbing Liu, “Deep spatio-temporal 3D dilated dense neural network for traffic flow prediction,” Expert Systems with Applications , vol. 237, no. 121394, 2024

  38. [46]

    Intelligent 3D Objects Classification for Vehicular Ad Hoc Network Based on Lidar and Deep Learning Approaches,

    Pedro Henrique Feij ´o de Sousa, Jefferson Silva Almeida, Elene Firmeza Ohata, Fabr´ıcio Gonzalez Nogueira, Bismark Claure Torrico, and Victor Hugo Costa de Albuquerque, “Intelligent 3D Objects Classification for Vehicular Ad Hoc Network Based on Lidar and Deep Learning Approa...

  39. [47]

    V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception,

    Runsheng Xu, Xin Xia, Jinlong Li, Hanzhao Li, Shuo Zhang, Zhengzhong Tu, Zonglin Meng, Hao Xiang, Xiaoyu Dong, Rui Song, Hongkai Yu, Bolei Zhou, and Jiaqi Ma, “V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception,” in Proceedings of the IEEE...

  40. [48]

    A multi-objectives framework for secure blockchain in fog–cloud network of vehicle-to-infrastructure applica- tions,

    Abdullah Lakhan, Mazin Abed Mohammed, Karrar Hameed Abdulka- reem, Muhammet Deveci, Haydar Abdulameer Marhoon, Jan Ne- doma, and Radek Martinek, “A multi-objectives framework for secure blockchain in fog–cloud network of vehicle-to-infrastructure applica- tions,” Knowledge-Bas...

  41. [49]

    Two- Way Reliable Forwarding Strategy of RIS Symbiotic Communications for Vehicular Named Data Networks,

    Kai Fang, Boyu Yang, Han Zhu, Zhihua Lin, and Zhuoran Wang, “Two- Way Reliable Forwarding Strategy of RIS Symbiotic Communications for Vehicular Named Data Networks,” IEEE Internet of Things Journal , vol. 10, no. 22, pp. 19385–19398, 2022

  42. [50]

    Blockchain-Empowered Resource Allocation in HAPS-assisted IoV Digital Twin Networks: A Federated DRL Approach,

    Hayla Nahom Abishu, Abegaz Mohammed Seid, Rutvij H. Jhaveri, Thippa Reddy Gadekallu, Aiman Erbad, and Mohsen Guizani, “Blockchain-Empowered Resource Allocation in HAPS-assisted IoV Digital Twin Networks: A Federated DRL Approach,” IEEE Transac- tions on Intelligent Vehicles , ...

  43. [51]

    Group’n Route: An Edge Learning-Based Clustering and Efficient Routing Scheme Leveraging Social Strength for the Internet of Vehicles,

    Naercio Magaia, Pedro Ferreira, Paulo Rog ´erio Pereira, Khan Muham- mad, Javier Del Ser, and Victor Hugo C. de Albuquerque, “Group’n Route: An Edge Learning-Based Clustering and Efficient Routing Scheme Leveraging Social Strength for the Internet of Vehicles,” IEEE Transactio...

  44. [52]

    Link Optimization in Software Defined IoV Driven Autonomous Transportation System,

    Ali Hassan Sodhro, Joel J. P. C. Rodrigues, Sandeep Pirbhulal, No- man Zahid, Ant ˆonio Roberto L. de Macedo, and Victor Hugo C. de Albuquerque, “Link Optimization in Software Defined IoV Driven Autonomous Transportation System,” IEEE Transactions on Intelligent Transportation...

  45. [53]

    PIG: Prompt Images Guidance for Night-Time Scene Parsing,

    Zhifeng Xie, Rui Qiu, Sen Wang, Xin Tan, Yuan Xie, and Lizhuang Ma, “PIG: Prompt Images Guidance for Night-Time Scene Parsing,” IEEE Transactions on Image Processing , vol. 33, pp. 3921–3934, 2024

  46. [54]

    Llama 2: Open Foundation and Fine- Tuned Chat Models,

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Alma- hairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al, “Llama 2: Open Foundation and Fine- Tuned Chat Models,” arXiv preprint arXiv:2307.09288 , 2023

  47. [55]

    ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools,

    Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, et al, “ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools,” arXiv preprint arXiv:2406.12793 , 2024

  48. [56]

    Towards Safe Autonomy in Hybrid Traffic: Detecting Unpredictable Abnormal Behaviors of Human Drivers via Information Sharing,

    Jiangwei Wang, Lili Su, Songyang Han, Dongjin Song, and Fei Miao, “Towards Safe Autonomy in Hybrid Traffic: Detecting Unpredictable Abnormal Behaviors of Human Drivers via Information Sharing,” ACM Transactions on Cyber-Physical Systems, vol. 8, no. 16, pp. 1–25, 2024

  49. [57]

    Generative Abnormal Data Detection for Enhancing Cellular Vehicle-to-Everything-Based Road Safety,

    Liang Zhao, Xu Fan, Ammar Hawbani, Lexi Xu, Keping Yu, and Zhi Liu, “Generative Abnormal Data Detection for Enhancing Cellular Vehicle-to-Everything-Based Road Safety,”IEEE Transactions on Green Communications and Networking , vol. 8, no. 4, pp. 1466–1478, 2024

  50. [58]

    ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,

    Shuyang Li, Talha Azfar, and Ruimin Ke, “ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,” IEEE Transactions on Intelligent Vehicles , pp. 1–12, 2024, DOI: 10.1109/TIV .2024.3508471

  51. [59]

    Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic- Aware Watermark Remedy,

    Yu Fu, Deyi Xiong, and Yue Dong, “Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic- Aware Watermark Remedy,” in Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI) , 2024

  52. [60]

    Who Wrote this Code? Watermarking for Code Generation,

    Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, and Gunhee Kim, “Who Wrote this Code? Watermarking for Code Generation,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, (Volume 1: Long Papers),...

  53. [61]

    Adversarial watermarking trans- former: Towards tracing text provenance with data hiding,

    Sahar Abdelnabi and Mario Fritz, “Adversarial watermarking trans- former: Towards tracing text provenance with data hiding,” in Proceed- ings of the 2021 IEEE Symposium on Security and Privacy (S&P) , pp. 121–140, 2021

  54. [62]

    Content- preserving text watermarking through unicode homoglyph substitution,

    Stefano Giovanni Rizzo, Flavio Bertini, and Danilo Montesi, “Content- preserving text watermarking through unicode homoglyph substitution,” in Proceedings of the 20th International Database Engineering & Applications Symposium, pp. 97–104, 2016

  55. [63]

    Robust multi-bit natural language watermarking through invariant features,

    KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak, “Robust multi-bit natural language watermarking through invariant features,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL) , pp. 2092–2115, 2023

  56. [64]

    Tracing text provenance via context-aware lexical substitution,

    Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, and Nenghai Yu, “Tracing text provenance via context-aware lexical substitution,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 36. pp.11613–11621, 2022

  57. [65]

    A Study of Situational Reasoning for Traffic Understanding,

    Jiarui Zhang, Filip Ilievski, Kaixin Ma, Aravinda Kollaa, Jonathan Francis, and Alessandro Oltramari, “A Study of Situational Reasoning for Traffic Understanding,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) , pp. 3262–3272, 2023

  58. [66]

    AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback,

    Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto, “AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback,” Advances in Neural Information Processing Systems ...

  59. [67]

    FinQA: A Dataset of Numerical Reasoning over Financial Data,

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang, “FinQA: A Dataset of Numerical Reasoning over Financial Data,” in Proceedings of the 2021 Conference on Empirical...

  60. [68]

    WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models,

    Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, and Juanzi Li, “WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1517–1542, 2024

  61. [69]

    Bertscore: Evaluating text generation with bert,

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi, “Bertscore: Evaluating text generation with bert,” in Pro- ceedings of the International Conference on Learning Representations (ICLR), 2020

  62. [70]

    ROUGE: A Package for Automatic Evaluation of Sum- maries

    Chin-Yew Lin. ROUGE: A Package for Automatic Evaluation of Sum- maries. Text summarization branches out , pp. 74–81, 2004

  63. [71]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach,

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint arXiv:1907.11692, 2019

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.