Pith. sign in

Paper Citation Record · LEDGER

Assessing the Robustness of LLM-based NLP Software via Automated Testing

As of 23 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2412.21016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21016 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:09:21.501802Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact1
  • verified fuzzy59
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33674812-ffab-4a95-976f-72c92510a7b6 · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.112333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.112333Z digest=sha256:f86eb3e6c35e86da3b40ff3b79f11928d9d6d42f44aa80e9195b46d14c0f36b4

Observation 5fc046a3-bdd2-4014-8b3e-1983ec2a78de · outbound

This paper cites Fuzz4all: Universal fuzzing with large language models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Fuzz4all: Universal fuzzing with large language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.930521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.117484Z digest=sha256:b57f43db555022cff31a373d8505178e4190d16f18f2616fc6d650cf56ae259b

Observation 3275661c-c057-4a9b-a995-da7dae0e8e3c · outbound

This paper cites Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.122598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.122598Z digest=sha256:95f7c64c35bf693245e4faa8af9e9b1d7910b5a9d90d7d520921214ec46914ae

Observation 9a76929f-de3e-4593-98c0-fe274ac3ceb8 · outbound

This paper cites Sentiment analysis for software engineering: How far can we go?.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Sentiment analysis for software engineering: How far can we go?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.128619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.128619Z digest=sha256:b2afc4c146e817e1e67f44aba6a945ac92bc35efbed09a01b5263db3ad73406b

Observation 6734b065-b4a1-4a6e-9210-d9360d842021 · outbound

This paper cites Mttm: Metamorphic testing for textual content modera- tion software,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Mttm: Metamorphic testing for textual content modera- tion software,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.906364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.135286Z digest=sha256:f9ba53621fb0e3d1e70b37ec1d8b2890bb1a4716997632e960b745bacc944de6

Observation d130626d-b2ee-462e-9257-46e59a1dc201 · outbound

This paper cites Unilog: Automatic logging via llm and in- context learning,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Unilog: Automatic logging via llm and in- context learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.893731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.140407Z digest=sha256:cd6f5866e405a0a58844b2e37e76e7b4c40061cd940e9e19cce3f56b334eecc3

Observation 79788340-7863-497e-873d-a19591d619e9 · outbound

This paper cites Managing extreme ai risks amid rapid progress,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Managing extreme ai risks amid rapid progress,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.145332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.145332Z digest=sha256:e39211b771099ea996f24629b5d3b50ac982d0ba79a1c83a096aaccaff943b67

Observation 3ef7f9ef-f075-4cd6-aa71-e3074c995016 · outbound

This paper cites Chatgpt incorrectness detection in software reviews,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Chatgpt incorrectness detection in software reviews,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.879241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.149682Z digest=sha256:30041d6b8030b7b6eebc08654abdc37f57e5f4ff92dc15d76074812968ff3950

Observation 7d33d601-7a80-4fd3-af4d-501597825469 · outbound

This paper cites Development in times of hype: How freelancers explore generative ai?.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Development in times of hype: How freelancers explore generative ai?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.865692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.154156Z digest=sha256:10f2acb4e6b62b4eb68ae579b31954548f20f26a7b893c054586bba691af7697

Observation 837bd9ab-c39e-446e-981c-43f9c694c2be · outbound

This paper cites How far are we? the triumphs and trials of generative ai in learning soft- ware engineering,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing How far are we? the triumphs and trials of generative ai in learning soft- ware engineering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.851908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.161003Z digest=sha256:93c79fe35871abf86ab9cc306080412a7d2a2010da5197d6efdeb78778f330cc

Observation fd1bcdfb-db8b-4d35-8f79-31a2200d7232 · outbound

This paper cites Text classification via large language models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Text classification via large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.838328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.167523Z digest=sha256:d6d5d691dfe9ee79c93ed9dfaf40cf37f488baf5d98a6b82e2eff4b079f06669

Observation 64384c7c-b40e-418f-b1d4-c50c6e9032ae · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Chatgpt outperforms crowd workers for text-annotation tasks,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.824801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.173570Z digest=sha256:e1ad03c0f4c9edc10177be0b31dc2790004ef5dd000f334f790d413d052bd144

Observation c055efff-a4c3-4010-baa4-bae9e7b02262 · outbound

This paper cites Can llms replace manual annotation of software engineering artifacts?.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Can llms replace manual annotation of software engineering artifacts?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.811293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.178258Z digest=sha256:3cc1c2618e81574e7b006083703f1ddb4d48135bed2842d0fb0f05fb7e62180c

Observation a84229fe-463e-4dbd-8c3b-ac72a8d2a595 · outbound

This paper cites Assessing the robustness of llm-based nlp software via automated testing,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Assessing the robustness of llm-based nlp software via automated testing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.798225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.183227Z digest=sha256:28a4a1dfe865dd73807ad23fc5c736b299ab3219b120518e491354d5bc767bb6

Observation 95d88076-b034-4506-8d6a-c8693001592a · outbound

This paper cites Nanofuzz: A usable tool for automatic test generation,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Nanofuzz: A usable tool for automatic test generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.784085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.187144Z digest=sha256:ab5481410776cf41a56864a1cb3e42c7c130145681c067b900b5bc011b35e8d9

Observation c6cdac78-0161-4d62-8492-52b18cba3161 · outbound

This paper cites Toward stealthy backdoor attacks against speech recognition via elements of sound,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Toward stealthy backdoor attacks against speech recognition via elements of sound,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.770094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.191723Z digest=sha256:6883203ade4b6d1417dcb9fcd9ddd147100a5810fb3f7221a3d2ef419f093b73

Observation f0cc6a6b-bf70-4b63-abc5-598523c33f69 · outbound

This paper cites Multitest: Physical-aware object insertion for testing multi-sensor fusion perception systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Multitest: Physical-aware object insertion for testing multi-sensor fusion perception systems,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.756528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.195394Z digest=sha256:75e889c899d5e68fbcdb98c29dce12264c2d06ae202bea1a7c83874c7444638f

Observation 83743ad6-c082-46a2-97b9-f349c64aa8f3 · outbound

This paper cites Glue-x: Evaluating natural language understanding models from an out-of-distribution generalization perspective,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Glue-x: Evaluating natural language understanding models from an out-of-distribution generalization perspective,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.741857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.200738Z digest=sha256:a8f685c5db9295607be9e709d7b58a839c4ce4f6136437691ca07fc142994bb9

Observation 7ffae674-a0e3-4f5f-9391-f2cc5313b7ee · outbound

This paper cites Black box adversarial prompting for foundation models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Black box adversarial prompting for foundation models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.728528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.206056Z digest=sha256:bcb9a99c550133aca17d8d54d624ce5b4aa72892ca10e310b9750438a0d5dabf

Observation 356e6372-d97a-48e8-8266-4f3741b08d8b · outbound

This paper cites On the robustness of chatgpt: An adversarial and out-of-distribution perspective,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing On the robustness of chatgpt: An adversarial and out-of-distribution perspective,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.714924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.210579Z digest=sha256:03a0da4257ecbf94e00fe9d79b54ffdd491cdba6d596aee660a6e174da92a307

Observation 418568d6-1722-4b50-8143-500f96e957ac · outbound

This paper cites Software testing with large language models: Survey, landscape, and vision,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Software testing with large language models: Survey, landscape, and vision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.215684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.215684Z digest=sha256:d87280b7c1888c51808844b08f0a5965acd4d9548b77fc1ac181130d50907f26

Observation e3204ae7-2767-41c9-94ec-6f5d62b0024a · outbound

This paper cites Llmeffichecker:understanding and testing efficiency degradation of large language models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Llmeffichecker:understanding and testing efficiency degradation of large language models,

Reference 22

Resolution
verified exact
doi, observed 2026-08-10T23:09:21.554208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.220908Z digest=sha256:795b7d8ad6a8d146ca6cbddc354463191e61d11313113b2fb2e471ef2ae880e7

Observation 1eda7d60-e6fd-4ff6-93e0-fafca9f7216c · outbound

This paper cites Deepatash: Focused test generation for deep learning systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Deepatash: Focused test generation for deep learning systems,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.692771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.226102Z digest=sha256:b2a9e510b26592bfaac2f003307c60bbee29304bde9ebc241b4d26618c133e0f

Observation 6703d7f1-45fe-4da1-ae5f-31e97dc19e4c · outbound

This paper cites Atom: Automated black-box testing of multi-label image classification systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Atom: Automated black-box testing of multi-label image classification systems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.676766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.230876Z digest=sha256:4616513f8eda0975c5cbc488f4c7ad55f953c6980dbf8dbee808ab52820c51c7

Observation cd36557c-9ad9-42e6-8f46-5780b88d30de · outbound

This paper cites Repairing failure-inducing inputs with input reflection,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Repairing failure-inducing inputs with input reflection,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.663064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.237814Z digest=sha256:f1056e244ca3cedb81f228a7085c20e286ea1e9cfbec1399c4129cba738743c7

Observation 7fe55874-cc23-415b-a653-00cf44943bac · outbound

This paper cites Generating natural language adversarial examples through probability weighted word saliency,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Generating natural language adversarial examples through probability weighted word saliency,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.242969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.242969Z digest=sha256:09946cc02a46c78e65ea6be0733acafc870060c1ddcc162fcc206c3af13b07f2

Observation 02524678-74b1-4ee3-bf31-77690262da23 · outbound

This paper cites Good debt or bad debt: Detecting semantic orientations in economic texts,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Good debt or bad debt: Detecting semantic orientations in economic texts,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.248451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.248451Z digest=sha256:202d5a67cc2f56d950b182f4bbc88df5a7fc837d5e7c5d874c4a83b62fe30bd9

Observation 1a6fa3f3-8177-4a6f-a6e9-f141eec723ba · outbound

This paper cites Character-level convolutional networks for text classification,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Character-level convolutional networks for text classification,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.629873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.253715Z digest=sha256:8a4d9e347e75bb1d75bfbd4b01cd8a84917c6044a18a2407080eaa0a0413d5bd

Observation 23b35352-b359-4bb7-a925-b8d464360702 · outbound

This paper cites Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.617183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.258746Z digest=sha256:3ba7a411b63749e9074920a246fbc947a318a23a69fc72a1599d87a75139948f

Observation 9db1dfea-966a-4440-934a-1f376c4fa517 · outbound

This paper cites Beam search: faster and monotonic,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Beam search: faster and monotonic,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.603298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.263321Z digest=sha256:b32a551bd3908fd2b8caa9f744aa906eede3878f9715c151caf7c48d0d1be838

Observation 6839bccf-15d0-4655-b4dc-20d562b05e0c · outbound

This paper cites Chatgpt- resistant screening instrument for identifying non-programmers,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Chatgpt- resistant screening instrument for identifying non-programmers,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.589874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.268488Z digest=sha256:4a39e76dff24a8c8533913d7bed9ecbb79e96ac901076cdc1ddfd4eb53f2f703

Observation 1ec37e3c-5686-439f-859f-92f35dd2dc7f · outbound

This paper cites Uncovering the causes of emotions in software developer communication using zero-shot llms,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Uncovering the causes of emotions in software developer communication using zero-shot llms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.576917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.273594Z digest=sha256:7970cc5d46c67a32d2836fc2bd525c48728b860e5352b6eb3d2cf184f9e1b892

Observation fdffdce0-926e-474c-a83d-34b258b5b19b · outbound

This paper cites Can automated text classification improve content analysis of software project data?.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Can automated text classification improve content analysis of software project data?

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.562736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.278981Z digest=sha256:11b4283dba9b4fc2e3b9840429e3978dac7340c3ee8be62cc87fad5ecbf1f4c6

Observation bf5a8710-661a-4804-9b04-dcfe1db9398f · outbound

This paper cites Hqa-attack: toward high quality black-box hard-label adversarial attack on text,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Hqa-attack: toward high quality black-box hard-label adversarial attack on text,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.549179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.283195Z digest=sha256:fc81200c1c58aae46cddf343f9a36f040347111477a1599aea9e109061a849d7

Observation 770485b4-bd9c-41ac-a857-1b3976cbc250 · outbound

This paper cites Limeattack: Local explainable method for textual hard-label adversarial attack,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Limeattack: Local explainable method for textual hard-label adversarial attack,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.535128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.287860Z digest=sha256:bcd022a438c86ca41164864bbba796134a70b0987f23c3a17583646fae6c00c1

Observation bb9810b9-a9c9-407c-8463-5f5d28578be6 · outbound

This paper cites Texthacker: Learning based hybrid local search algorithm for text hard-label adversarial attack,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Texthacker: Learning based hybrid local search algorithm for text hard-label adversarial attack,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.521916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.292989Z digest=sha256:3547cd5fc4eb3cfae9bd61d8c7e9adf4a2a4dc038a548fd298909f9dbf721c58

Observation 53cd349d-7b15-4647-8aad-59cfa80feb9c · outbound

This paper cites Natural language adversarial defense through synonym encoding,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Natural language adversarial defense through synonym encoding,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.509802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.297170Z digest=sha256:d090323a59d97f188b713047743020887f03b31a84313fbc98aa2dfacab92b9f

Observation 1aa2771a-3009-4c58-9d41-9ef9025fe74e · outbound

This paper cites Leap: Efficient and automated test method for nlp software,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Leap: Efficient and automated test method for nlp software,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.496569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.302028Z digest=sha256:882f8d3ec639070c691af2c979727434d377a4f73ff9b26ba4cf3315721cfb99

Observation 83b1b76b-6787-41e6-aa05-60c5b669d082 · outbound

This paper cites Beyond accuracy: Behavioral testing of nlp models with checklist (extended abstract),.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Beyond accuracy: Behavioral testing of nlp models with checklist (extended abstract),

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.483626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.307005Z digest=sha256:3824cc924336ea585bbc3d8c886a7886faacf37e1dda1c1b5f720a780f8d952d

Observation a4526533-c323-41d0-aeb1-765adb7d6204 · outbound

This paper cites Understanding the value of software engineering technologies,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Understanding the value of software engineering technologies,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.470028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.311384Z digest=sha256:efbb53f7d0ee04843ef0334f41920b03833140c8dcce73d4e7fba755f28e90dc

Observation db34cc07-6585-4a53-bc40-589bf21f99ba · outbound

This paper cites Enhancing speaker diarization with large language models: A contextual beam search approach,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Enhancing speaker diarization with large language models: A contextual beam search approach,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.456038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.315753Z digest=sha256:2104d3c02f9ff0fb372186f07bf926b2667346afe917dc002afdc2163abdf592

Observation d54dc19c-591e-428f-9eb5-5f4ed2d8e374 · outbound

This paper cites Conformal au- toregressive generation: Beam search with coverage guarantees,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Conformal au- toregressive generation: Beam search with coverage guarantees,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.443625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.321012Z digest=sha256:771e36ad2c6bfbd2be123fe33728a5153eeb77371725e04f6ad083a0557107b2

Observation aef45b66-761c-4b47-b4a6-59e01dafb0df · outbound

This paper cites Hybrid filtered beam search algorithm for the optimization of monitoring patrols,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Hybrid filtered beam search algorithm for the optimization of monitoring patrols,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.430249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.325846Z digest=sha256:5b79b70033bf615459fabd72f05507110c06387115136cf0fc6fc725f06e2c94

Observation 94c062c3-9f67-4575-86b2-44ed789ac295 · outbound

This paper cites Large- scale language model rescoring on long-form data,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Large- scale language model rescoring on long-form data,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.416829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.331963Z digest=sha256:3d9cf6a201df79e1f1b0c6cd2e02533ebf09eab6a0617c830072533ec53a3c5c

Observation 4b0f0987-71b3-4720-8c34-8720c6a882d5 · outbound

This paper cites Beamqa: Multi-hop knowledge graph question answering with sequence-to-sequence prediction and beam search,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Beamqa: Multi-hop knowledge graph question answering with sequence-to-sequence prediction and beam search,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.403722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.338073Z digest=sha256:06ba8f0bcf0cb294e38e2d84f664bdf6c2d281b67df2923cfd898a7b8c8288f8

Observation 2171988e-3c71-4cb6-92a8-9274de5bcff7 · outbound

This paper cites Iso/iec,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Iso/iec,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.390794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.344442Z digest=sha256:06023f97287a98f04fa59b7c462e0eed123d2483d1723f77e0fed2922b743e2f

Observation 8e89aa7e-b2b6-4294-97dc-df6a2c1acaf5 · outbound

This paper cites Intriguing properties of neural networks,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Intriguing properties of neural networks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.378054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.349539Z digest=sha256:67d27241f91523613108e5a68c905f525c9f040b98a1b1e25395d8b6a134d332

Observation 9a19dcab-594a-416e-ae87-c45b2e6a2933 · outbound

This paper cites Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.364737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.358164Z digest=sha256:2b7e4fbffb8172ce96eddade4186f0e38ccc11126c7ba9882df0f118acc6234b

Observation b548f926-232c-4463-8ccf-99137dfced67 · outbound

This paper cites Fast adversarial attacks on language models in one GPU minute,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Fast adversarial attacks on language models in one GPU minute,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.349577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.362862Z digest=sha256:9194b1c49d7e2a59ff74ca1925ef46e87eeffa7a2173dbc17c20015cfd51e2f9

Observation 3456b88d-b2b2-461e-ac4e-bc8292936efa · outbound

This paper cites Hotflip: White-box adversarial examples for text classification,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Hotflip: White-box adversarial examples for text classification,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.334432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.367184Z digest=sha256:39e3421792525dc88599cbd049bf4af6a615570d16a3f56079dc56c504e1eb37

Observation 1cb91ef3-1283-438c-b482-d6067edec186 · outbound

This paper cites Towards improving adversarial training of nlp models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Towards improving adversarial training of nlp models,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.316213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.371238Z digest=sha256:ceece977a97e7ce9f94d4e959ef276efade0bd3c07066756b7370e2f6404b24e

Observation a4fc34d3-11d1-403c-a56b-dc14ae58942a · outbound

This paper cites A recovering beam search algorithm for the one-machine dynamic total completion time scheduling problem,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing A recovering beam search algorithm for the one-machine dynamic total completion time scheduling problem,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.303326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.376016Z digest=sha256:ac4800918d07bd37bd7c98d171b90a5247a36e09cc38cc8ecdc6b7d31e65bd66

Observation 3d2f3aa9-135a-4f03-91ea-e5a519594840 · outbound

This paper cites Backtrack beam search for multiobjective scheduling problem,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Backtrack beam search for multiobjective scheduling problem,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.289417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.380919Z digest=sha256:e86f54535a24a00e057affb7163bdd34f612a4d0d9c015076031daf6b20d38d0

Observation c4591ce7-dbcf-48ba-b0b5-40512ade96cd · outbound

This paper cites Is bert really robust? a strong baseline for natural language attack on text classification and entailment,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Is bert really robust? a strong baseline for natural language attack on text classification and entailment,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.274372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.386940Z digest=sha256:00c86a18ab1b07424d19e51973364b6f8cc55652f82eae4a030ab28e132b6be7

Observation 9f2216c0-d92e-4a72-b0db-469daffdd1b9 · outbound

This paper cites Understanding Neural Networks through Representation Erasure.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Understanding Neural Networks through Representation Erasure

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.391726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.391726Z digest=sha256:6e73c0ebda4fd1295b6ee7d7f1fb2cdd7cece5d5c9dfd42c9d0ff60be8058adb

Observation c37a065d-974a-4ff8-a19e-5fab79d23ab7 · outbound

This paper cites Wordnet: A lexical database for english,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Wordnet: A lexical database for english,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.397121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.397121Z digest=sha256:01469904449f9a241c5b21171431e4f4d478f17151545c0cdcc0a97fa696397b

Observation 56b72cdf-06b7-4f4b-ae7a-3322a7ef246c · outbound

This paper cites Word-level textual adversarial attacking as combinatorial optimization,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Word-level textual adversarial attacking as combinatorial optimization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.261596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.400795Z digest=sha256:b0cc93f3e62d497e7f1b53a6cffe730b4ac11a804113b2a190eadbd551a3642c

Observation 16f620f4-327d-4154-b507-affe5b7aed6e · outbound

This paper cites Mistral 7B.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Mistral 7B

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.405749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.405749Z digest=sha256:58fae8c4feeade4be39213aeda032300e48884e29fbba5bd62835ffbebbca18b

Observation 64be3a36-ae2d-4a4c-b8a5-3dddb085b5d8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.411075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.411075Z digest=sha256:d3faf3bdf944572eaa91914921126e5311d9bd3a76a3d85d11291ad27d7e6d40

Observation 9d9c93eb-3fec-47a4-ba23-e26e00ff5fc9 · outbound

This paper cites Internlm2 technical report,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Internlm2 technical report,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.247527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.415492Z digest=sha256:4742b7b36c00684802f16e3baf54895ebe9647838e39ade52d1dff8e2c526164

Observation ba88e5d4-4337-467b-9f14-0d807046f222 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Yi: Open Foundation Models by 01.AI

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.419919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.419919Z digest=sha256:12d25fabe03960d788abeb6583ec69848d6c5981b653eb2d7cf2367f096f3903

Observation 6f53a992-0b25-4b12-a40b-0fee021bb096 · outbound

This paper cites Stress test evaluation for natural language inference,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Stress test evaluation for natural language inference,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.233591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.424990Z digest=sha256:9239c3c5564047ffeaa0ee1a6227187489d9ab63586c711f44e1544b58d1a774

Observation b502ec63-34fc-469d-838b-49a584c0565c · outbound

This paper cites Testing the limits: Unusual text inputs generation for mobile app crash detection with large language model,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Testing the limits: Unusual text inputs generation for mobile app crash detection with large language model,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.218959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.428948Z digest=sha256:995703a98527ab421e025353d93c7da25584175ae75f61f9b339ba24164d7049

Observation 70d22f26-bd95-44d1-be22-723f9ab587f2 · outbound

This paper cites Textbugger: Generating adversarial text against real-world applications,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Textbugger: Generating adversarial text against real-world applications,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.433205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.433205Z digest=sha256:5135e91342c7021d241479d5624e30505e15c1b6f7ac5fe369ea0867385b1d08

Observation abe60e63-e933-42f6-a14e-1903770c4391 · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.436910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.436910Z digest=sha256:1044d14a0bc0d46bc5d44c7fbbc16177e1fbedc2d409f74d044e9073bf2d2b1e

Observation e9bf3aa8-a061-45a8-bc64-1c5f44069dbd · outbound

This paper cites Augmenting LLMs with Knowledge: A survey on hallucination prevention.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Augmenting LLMs with Knowledge: A survey on hallucination prevention

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.441235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.441235Z digest=sha256:9795cdfcd4a88e9c02b9de3c01675760d04c0f380bec52842da48d420733ac16

Observation 2eb9e727-e280-4086-93dd-3f36968b2241 · outbound

This paper cites Coophance: Cooperative enhancement for robustness of deep learning systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Coophance: Cooperative enhancement for robustness of deep learning systems,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.196783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.446430Z digest=sha256:996130f7d54b2911c5b7f23e7c64efd8910ba281181ea370e1582931948c6c0e

Observation 326b48ef-266d-4ae1-b114-cff7df47625a · outbound

This paper cites Black-box testing of deep neural networks through test case diversity,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Black-box testing of deep neural networks through test case diversity,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.184012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.451171Z digest=sha256:29308924e657a2f264816d2dc75f1b79a397b3ee3bda4127d828ca5f5c91e011

Observation a852b0f6-d2d1-4547-9ea1-828af8349649 · outbound

This paper cites Dialtest: automated testing for recurrent- neural-network-driven dialogue systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Dialtest: automated testing for recurrent- neural-network-driven dialogue systems,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.169558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.456009Z digest=sha256:c54ed1dde8a18e1b904fca9c6df1550e8140807f3510906049082d22e2563a6d

Observation 3b9024e0-ef21-40ed-baff-0411a4256ad7 · outbound

This paper cites Keeper: Automated testing and fixing of machine learning software,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Keeper: Automated testing and fixing of machine learning software,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.460248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.460248Z digest=sha256:d7e02cb9e892485b9428fad78b39d623986dad619eacff66a145237ba3a855db

Observation be0407e1-9bf2-4f16-88f5-666ea11fda5f · outbound

This paper cites Automated testing and improvement of named entity recognition systems,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Automated testing and improvement of named entity recognition systems,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.156062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.464637Z digest=sha256:e0cabd4753bec3821d7ee733c19771f1d6147c3e9570180fdbc8a4435cb9d4ed

Observation 7721979d-88a7-452c-b106-0bc9ea327962 · outbound

This paper cites Keeping llms aligned after fine-tuning: The crucial role of prompt templates,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Keeping llms aligned after fine-tuning: The crucial role of prompt templates,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.142585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.468984Z digest=sha256:cbbd6cff29567fa467a60e57db0aa5b1b97218362fcf933034f995ecb5306de0

Observation a004e29f-f548-4591-95f4-3e135459b092 · outbound

This paper cites Look before you leap: An exploratory study of uncertainty analysis for large language models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Look before you leap: An exploratory study of uncertainty analysis for large language models,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.129508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.474107Z digest=sha256:9af074f9c1cd72e86e791468a40f75fc5936b1774c7efe6b649cb9abe29db82f

Observation 753f0aea-8d64-4766-82db-6ce0c996a62e · outbound

This paper cites Imperceptible content poisoning in llm-powered applications,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Imperceptible content poisoning in llm-powered applications,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.482667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.482667Z digest=sha256:ab57eca32683a0ef71aa6e449ebe5fb4969f2e3d83d28cc52a9f927ed0fd4d44

Observation c9456c30-064f-4ae7-971b-6105118bb1b8 · outbound

This paper cites Revisiting out-of-distribution robustness in nlp: Bench- marks, analysis, and llms evaluations,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Revisiting out-of-distribution robustness in nlp: Bench- marks, analysis, and llms evaluations,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.116338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.487088Z digest=sha256:ae697587efa7e8a0578e0bc77351b7be113008ea6c52aabca5e1dd0b2b606225

Observation cf59aa33-b82f-4c3d-9b43-7b11c05dd111 · outbound

This paper cites Revisit input perturbation problems for llms: A unified robustness evaluation framework for noisy slot filling task,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Revisit input perturbation problems for llms: A unified robustness evaluation framework for noisy slot filling task,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.102241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.492591Z digest=sha256:4e3a4c07a3f354b341ba26ac347b394d5a343c019ff2dd9b1d6b4885de742fe1

Observation 0ab261ff-e313-4c36-a2d3-e989f5cca9cd · outbound

This paper cites Ro- bustness over time: Understanding adversarial examples’ effectiveness 16 on longitudinal versions of large language models,.

Assessing the Robustness of LLM-based NLP Software via Automated Testing Ro- bustness over time: Understanding adversarial examples’ effectiveness 16 on longitudinal versions of large language models,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:21.497519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:21.497519Z digest=sha256:bff3c76691575cf6dbc7c86e69f82a0051e34342288a85d39da09f4e40c89a1b

Observation 6f7e4454-ddea-4549-97f0-3c53ed3ba49e · outbound

This paper cites degree in computer science and technology with the College of Computer Science and Software Engineering, Hohai University.

Assessing the Robustness of LLM-based NLP Software via Automated Testing degree in computer science and technology with the College of Computer Science and Software Engineering, Hohai University

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:09:22.088502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T23:09:21.501802Z digest=sha256:55cd93d7cb61799245b9ee75e36faa9090a5262667742b974b5930094b6b7735

Pith citing papers

No inbound Pith citation observations are available.