Pith. sign in

Paper Citation Record · LEDGER

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents

As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2607.06873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06873 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:02:58.383840Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact9
  • verified fuzzy50
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c9d1e7c-c348-422f-946b-62e9568cbc46 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ReAct: Synergizing reasoning and acting in language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.388221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fb81c0c36557cf96df75a07172f02d1d6eb2fe94d22b2c3242540a192835864f

Observation 72f10df0-056f-4523-8e0e-6aca88842e4b · outbound

This paper cites Toolformer: Language models can teach themselves to use tools,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Toolformer: Language models can teach themselves to use tools,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:26aaa833a35da8d7082af26cf3d6ec2ae6d735f2bb6ff89ed85dd2d83cc4858a

Observation 4ea94b3d-f9b9-4888-9852-dd8a529e6290 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.954013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e11cce9844e789b52db2176aa175f11fc026307e75f7aea8e3001f52f88a831b

Observation 4a4dfd04-19c2-4c43-a2fe-eafb3ce7c78b · outbound

This paper cites Preventing repeated real world AI failures by cataloging incidents: The AI incident database,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Preventing repeated real world AI failures by cataloging incidents: The AI incident database,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:8879d1e66bfc4c7409011e10858879c7b2ae1798a6da1ac17bd43ce749c4a410

Observation 4ffa99de-f735-4887-b6f9-e719cb61aae6 · outbound

This paper cites RealHarm: A collection of real-world language model application failures,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RealHarm: A collection of real-world language model application failures,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.353168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:47db442c8a8fbf123dd2fb64b7dfa5dc3ebe9846cd18f5378bcb9eefba0a1cde

Observation f27b15c7-3df5-4d10-bf94-f07e62d6107d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.972449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ee8a98dab8f70b88209149f1f533cf56ff733563f9bb973737ce292c3af40e52

Observation e592fa41-3f52-480f-a16a-9c01898fae04 · outbound

This paper cites WebArena: A realistic web environment for building autonomous agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents WebArena: A realistic web environment for building autonomous agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7854cb9bf5f60160d5ebc52dbeb5d9158ebb06a0ce295d0757bf17bff7dffcfe

Observation 0f3b2126-34aa-4a65-826f-ade350a7d323 · outbound

This paper cites GAIA: A benchmark for general AI assistants,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents GAIA: A benchmark for general AI assistants,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.381201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d5e1057db84ebccdd59457cbc93b24e8afbfa010d04375012411831d0ea0ea12

Observation 99a03855-90bb-4410-81dd-38744782f243 · outbound

This paper cites AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f9be4f3a09f63e10d67c0e54f3dcec2171d18b09c49f5bf6a89837c0c1def725

Observation c9221d46-79f0-4923-8502-f38c9021ec18 · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues?.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SWE-bench: Can language models resolve real-world GitHub issues?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.401263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f66ef8f68db070db6dac92bc5a2f5e7ed8f6b15087911d6b027adba1a87159a6

Observation 1f0725ff-4b9e-447f-b1c5-f27a6c1c8728 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.398826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f696ee1f15ef640e842b0fdfd3a4099f97f73f37d02bf9d2b9382161511bcd08

Observation 63c1db74-4e01-44c4-8c8c-8f522b8203c8 · outbound

This paper cites SpecOps: A fully automated AI agent testing framework in real-world GUI environments,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SpecOps: A fully automated AI agent testing framework in real-world GUI environments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:648463f2161715cad53047751ed87691afacf8f1f311eab2a450197d5f4577ce

Observation daea19d1-e896-4fbd-9314-21ce8a47e832 · outbound

This paper cites STELLAR: A search- based testing framework for large language model applications.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents STELLAR: A search- based testing framework for large language model applications

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.967164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:90d123b3163a76a94264a2559cd2021fc24545e3b9a3e507aa966dc92a041dab

Observation 6143456a-537e-4313-8166-7a7f273e6ff7 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.420670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:23a3d40be555d2c330ff5c26baf87959a745ba3b5b6407f6f77707600108b99d

Observation aec1170e-470f-4be9-8296-7486691ec429 · outbound

This paper cites A practitioner’s guide to process mining: Limitations of the directly-follows graph,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A practitioner’s guide to process mining: Limitations of the directly-follows graph,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.422481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c5d36e74d672c3e1c61bfecef2a3073b73087afd617aba35f5bdf87d66040034

Observation 57d39246-2796-4bff-b036-68890eadf922 · outbound

This paper cites τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.429430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7c179af947ed407a0de9f2937b7b4695bc9beac327cb56f316eb4512f1cc836f

Observation e5e93c59-a7ef-4463-9592-d07d6a3f5697 · outbound

This paper cites Event abstraction for process mining using supervised learning techniques,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction for process mining using supervised learning techniques,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.424048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4a1760fb7eaab329f810505acfc2109ee22018e0a268ded9a14ce67096790af6

Observation 94bedfb8-7f23-49ae-8e03-092a82ad01ac · outbound

This paper cites Event abstraction in process mining: Literature review and taxonomy,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction in process mining: Literature review and taxonomy,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:dededfa5ecaffbbb07b07dc54c4613dfab8f82137a74d23f04b6eca38f64748a

Observation a3e665c4-f53a-48ac-830b-15ce92ec3fb1 · outbound

This paper cites Boundary value exploration for software analysis,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Boundary value exploration for software analysis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.413741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fc15f6adaad1daccaa1532e0c3c6ef5495f2b9efdec64f6bd3cd2489a0f4a802

Observation 94441f16-9357-421f-bc16-0c2ed032bfb8 · outbound

This paper cites Automated robustness testing of off-the-shelf software components,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automated robustness testing of off-the-shelf software components,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.427615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a51fa1a23d707ec66083c78819fc28440db5db8d80bf9d2916673d9975900cd0

Observation 954f9e5c-be91-4ff4-8818-3fd3c6f4b45d · outbound

This paper cites τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.976247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f94fe4a0e0811dc68adb9766a71fabb147a49a4b8666cd632e27e0063ed84e08

Observation 528726a0-4bbc-43f0-82da-02e1076c96fc · outbound

This paper cites ToolLLM: Facilitating large language models to master 16000+ real-world APIs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ToolLLM: Facilitating large language models to master 16000+ real-world APIs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.391348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:deaa6ce9b93e2bcdd718195b572640a3f6311449b9e07abba65c74c2e4165356

Observation 313d6d59-d163-401c-9181-18e859c0f448 · outbound

This paper cites Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.963929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e09aaf795a28b51008e6a625c8df5e76f96fc1f3eec3345590c2b18d2ad44b6d

Observation 97062318-6ecb-476d-8c13-da94c9bc37b9 · outbound

This paper cites Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.975111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:78dc2a6fe9a24cb7be86c6e8ac2bb38020ce9024782d2fd6f8889460f9884966

Observation 86574250-54b4-41bf-a838-365082b4f56f · outbound

This paper cites Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:17b3cb6b52e5407c3a29b6a494acbde18bd40902551a4140b93d2e5f7b3e05e8

Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · outbound

This paper cites Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:06:37.956842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:058f459d5350e43aa2be8b4f076b86624cf2698fec015a9b2a2eb3623a350ad7

Observation bdd38177-ce3c-4b83-9b07-36928981b4da · outbound

This paper cites Mining specifications,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Mining specifications,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:18057951a6da02b5813015931edb616ec12ac39551ea517f5975237ed5a56049

Observation 84bb6543-3b44-4012-a793-c6147a4fb1f0 · outbound

This paper cites Discovering models of software processes from event-based data,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discovering models of software processes from event-based data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.395137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fccb7852821df9257567533bd8357d54854a186aaa57b28ef2be76ed961f145a

Observation f89dc7ec-340b-4e20-b62f-18d00ca50c54 · outbound

This paper cites Automatic generation of software behavioral models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automatic generation of software behavioral models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c7446474ef6ce5628d68f7dd6ff063478d922caac350ea3c7254bbe9b99ccd70

Observation 2dc6c576-bf93-4be9-bcc0-4dc77de44aad · outbound

This paper cites Inferring models of concurrent systems from logs of their behavior with CSight,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Inferring models of concurrent systems from logs of their behavior with CSight,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d11ccf3403f768b9394eb6594b456b23ca9cebd4e8afaeb45323763f25462bb3

Observation a0c731b5-5311-457b-9da7-7e14474b33d1 · outbound

This paper cites Workflow mining: Discovering process models from event logs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Workflow mining: Discovering process models from event logs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:49636b63605e4ce89f0531286d9c925b3a155f29c6d4b03ad4e7da26bdd58fe4

Observation 18aa5651-0695-48e9-822c-646501335467 · outbound

This paper cites Discov- ering block-structured process models from event logs—a constructive approach,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discov- ering block-structured process models from event logs—a constructive approach,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.425884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:74f6109dfa5504a45974890d9ebf872986452a797e30f761dd1a6efc83158404

Observation 1f2e5127-a857-461b-abc0-1219399aa20c · outbound

This paper cites Applying graph reduction techniques for identifying structural conflicts in process models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Applying graph reduction techniques for identifying structural conflicts in process models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:2e202d62ca026299dc650cfad6738d1969be98b023c22193314caf0f2994aa22

Observation d81ef0ce-87de-4eb1-a2a2-1871b5a90b68 · outbound

This paper cites Learning regular sets from queries and counterexamples,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Learning regular sets from queries and counterexamples,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.431693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0d8470bb7c2c6d4430fe7175ab715649d5088e3f8d4f21f3778d918567d43bf4

Observation 9e496994-d97f-432f-b6da-538d382c9579 · outbound

This paper cites Unsupervised dialog structure learning,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unsupervised dialog structure learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d8bacff9dfb927354f887626f175a883df072cb61cd2b7373119e0f2d810df40

Observation 8b21782c-e4f6-4cf1-8c76-f9e32d4497df · outbound

This paper cites Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:221c55c03085e11543ab2028f878b504cf4daa975b188b7d1d32a0c7f55d093d

Observation 47d02b58-9eb9-4c9d-8561-ca85fee5e5e5 · outbound

This paper cites Agenda-based user simulation for bootstrapping a POMDP dialogue system,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Agenda-based user simulation for bootstrapping a POMDP dialogue system,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.404999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9205a4140ee8ba0c8dff90745a2538473cae798ac3107f6f288762a40212bca5

Observation 8b43514d-7a55-415b-94e7-19673430fb4d · outbound

This paper cites ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a3bff8a3b30556ca9b6b0ae2fa297751bb6d8ecc78db383f9d31c422d1affd34

Observation 2fc55fc5-d849-4b04-aa23-93c0f960bcdb · outbound

This paper cites A taxonomy of model- based testing approaches,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A taxonomy of model- based testing approaches,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.385993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:85799f6bb887203071e979d8f2d403d68a549f03a99b541a38f0d41722494331

Observation 75ce28d8-4ae0-40ca-ba6d-2e5f3c4d9ca9 · outbound

This paper cites Principles and methods of testing finite state machines—a survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Principles and methods of testing finite state machines—a survey,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a4283560f66a850659cbbf3a89f656d4c1aae783b33ec8f259e1ad5a3b3b1cc3

Observation f73ca176-6c76-43e6-9d34-280c4bf33ab5 · outbound

This paper cites Testing software design modeled by finite-state machines,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing software design modeled by finite-state machines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.396696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:8cea3e43fd93fa2fb28e65877d2e54af0aa2a983a32a81f65a68adaf06e3d9c5

Observation 25e0e7b8-c0d7-4dfd-bdc5-0355e81b625d · outbound

This paper cites RESTler: Stateful REST API fuzzing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTler: Stateful REST API fuzzing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:cf9fccda40d1c00d9059ce7a01366de5d89d2a3542061c531b492fe25794c85f

Observation 413c8bf1-30d6-41dd-9201-7823baa165db · outbound

This paper cites RESTful API automated test case generation with Evo- Master,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTful API automated test case generation with Evo- Master,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.373884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c5d5c9e3c5e743ed3f48058907f0dd31b315ab0a838d7bbad79e02d60b1235b3

Observation 9c71df55-6f0d-40e5-86cb-c486edd2c362 · outbound

This paper cites Morest: Model-based RESTful API testing with execution feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Morest: Model-based RESTful API testing with execution feedback,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.351133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3286ec5fd9839772196a3a21cdf7b659b69d7c740270eafb5023983ad77986bc

Observation fdfc078e-4aac-477b-a422-f5bbcb06dd86 · outbound

This paper cites KAT: Dependency-aware automated API testing with large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents KAT: Dependency-aware automated API testing with large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.375692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e7bc2e8168fc7c4ec50d6db9c92529be57346497a48901dfe22fe8324b529841

Observation 5ccd9764-a37b-474b-b12d-7afdcb0079dd · outbound

This paper cites Testing RESTful APIs: A survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing RESTful APIs: A survey,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3fbdbf65257797f90146cee6df3d57a45d6c8ebbbd11112cd677e0c9bcefd715

Observation 2b3b289f-f2e8-4e78-98d5-f07c0cd671c2 · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents An empirical evaluation of using large language models for automated unit test generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0aaed769a5440a90c93a506717d6b58ec6618210db189745bfd80080b9092886

Observation 5352081e-274f-4b47-bb01-3889218a77e5 · outbound

This paper cites CoverUp: Effective high coverage test generation for Python,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CoverUp: Effective high coverage test generation for Python,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ae633dd64c74af320b68660666dfd36dfea9445b814b98f670348f4c775e018a

Observation a19faacc-39ec-42e5-9c38-01a65f71e8cd · outbound

This paper cites Evaluating and improving ChatGPT for unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Evaluating and improving ChatGPT for unit test generation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:63a94ccb06ddd1026ecf78568af9027ed51d31e7a3d357112f4159c745f2f623

Observation c6508104-bd2e-4132-b3b5-e0026d11b7f7 · outbound

This paper cites CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:5765bf29bd1dd43e312f74076e45dfbfddf080f1bd72d8910ebabe820b0b2e7a

Observation 74a8218e-fa88-4365-9149-1b2c1b93a272 · outbound

This paper cites TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:48ae9b9184b4027db4a50480c3d5ef25da7f2986c5b44941aed3a7c39d754134

Observation d2ee0e6d-23e2-41d1-865d-8836ec6aaeef · outbound

This paper cites The oracle problem in software testing: A survey.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents The oracle problem in software testing: A survey

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.362723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ad2afc3ea05e8c286b72e6641cab6ead948fb810e8c2f412ec71f1804ea83802

Observation 91b24357-a3aa-40bb-a95e-781f191ad93b · outbound

This paper cites Pseudo-oracles for non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Pseudo-oracles for non-testable programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.368635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:1336c503f0aaefadf830265fadd953e1c4594ef37e9464ba9744c0e9065f21b3

Observation e755b481-bc90-4bad-9642-d31ac5d8b987 · outbound

This paper cites On testing non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On testing non-testable programs,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.356276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c67b4b2ce0bcaaa3f7ddb10dbe9c6360764223868e2f4c161014055be80255c3

Observation 5280139e-b263-4a46-9432-063d6ccc5e10 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.379148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c922d25ba2ef7d1543c6a1d3b3f40e8a982c13bf7921b49cb1250469113132b0

Observation 8919763f-e676-4082-b6e1-3691f34f621c · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Large language models cannot self-correct reasoning yet

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:31698c6cc899f7a46313011d43e5b0c36e43d9fddd034ac7a8c23a049dd8fc17

Observation 56f5958e-caa8-4320-9086-ead614e1fd1c · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:22fce48b9960cc3d4456ad39735afe605b004cc8ac17919c4e1c92653b5274e5

Observation fc6e1223-e3df-467e-90f8-89dcda18765c · outbound

This paper cites Self-Refine: Iterative refinement with self-feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Self-Refine: Iterative refinement with self-feedback,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f900e558196bb9db2a7dd55e5f52d217268b51d6d821cd44aea677025484e5b0

Observation ffd8ae57-7bb3-41f1-a318-d4876be7b2df · outbound

This paper cites CRITIC: Large language models can self-correct with tool-interactive critiquing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CRITIC: Large language models can self-correct with tool-interactive critiquing,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:32ad5aa73961a829e52bbdec3398af907e7c4b4c0332b00b93992c580255fb76

Observation 4b153557-3da1-4ffe-b1b5-bd2cff717884 · outbound

This paper cites Metamorphic testing: A review of challenges and opportunities,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Metamorphic testing: A review of challenges and opportunities,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.398672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b3618cd7c10dc10f302642fb78fb618f65d6879bf80212e5fb6c09542e1115c0

Observation 45b02c00-7cdd-4acc-b72b-106b5381a916 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:1ed0096d56894c98342613b754211f3711fd7af60f9c8708ae9909b2b6e6843d

Observation 587a72f6-11a9-45de-b57e-fa05e453fc52 · outbound

This paper cites Ignore previous prompt: Attack techniques for language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Ignore previous prompt: Attack techniques for language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.419028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e9716423cdc086f3d70901c16b468c2b74b0da836ffc59fbcfcd5e155ad4d8fb

Pith citing papers

No inbound Pith citation observations are available.