Pith. sign in

Paper Citation Record · LEDGER

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents

As of 4 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2607.06873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06873 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T00:02:58.383840Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact9
  • verified fuzzy50
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c9d1e7c-c348-422f-946b-62e9568cbc46 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ReAct: Synergizing reasoning and acting in language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.388221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c6c0b339d7808379eb0a18aa42a9fbaee3269602039fad0fdaac9df9c1e7f7d6

Observation 72f10df0-056f-4523-8e0e-6aca88842e4b · outbound

This paper cites Toolformer: Language models can teach themselves to use tools,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Toolformer: Language models can teach themselves to use tools,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a56739f7fe34b17120353140646d6b960bf74ca795a6a97c1d1d4e9ee2e34efd

Observation 4ea94b3d-f9b9-4888-9852-dd8a529e6290 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.954013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:450a7002da7895215f6fddb96f76c370f7e1953c2c7d0d1442b28049f5e92b01

Observation 4a4dfd04-19c2-4c43-a2fe-eafb3ce7c78b · outbound

This paper cites Preventing repeated real world AI failures by cataloging incidents: The AI incident database,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Preventing repeated real world AI failures by cataloging incidents: The AI incident database,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:12aab4ea739856862b388201042bb3ea01dab38892ffd4a52f3ac95d2b91f452

Observation 4ffa99de-f735-4887-b6f9-e719cb61aae6 · outbound

This paper cites RealHarm: A collection of real-world language model application failures,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RealHarm: A collection of real-world language model application failures,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.353168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:175cab7687ff5eae6a07c221ce4ce130b567b79394b762d8b257d8551f94d59d

Observation f27b15c7-3df5-4d10-bf94-f07e62d6107d · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.972449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:14594f3a4b48d5f5c90cc39b67e0f02247d6703ceeb76083146f1c18580d4452

Observation e592fa41-3f52-480f-a16a-9c01898fae04 · outbound

This paper cites WebArena: A realistic web environment for building autonomous agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents WebArena: A realistic web environment for building autonomous agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:c1bc7e3740afeb5d629066f488b6fb01342e3ce0ca181b7bf165e3b0c54e9925

Observation 0f3b2126-34aa-4a65-826f-ade350a7d323 · outbound

This paper cites GAIA: A benchmark for general AI assistants,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents GAIA: A benchmark for general AI assistants,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.381201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:846906b421a044ece1ff3a9cd0807d67cd7099614328490acc1c824c58db1fd3

Observation 99a03855-90bb-4410-81dd-38744782f243 · outbound

This paper cites AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents AppWorld: A controllable world of apps and people for benchmarking interactive coding agents,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b02082faab8c94816fc5ec7615e38fdc604e628d9dd8c24ac4265f12ad6c83c5

Observation c9221d46-79f0-4923-8502-f38c9021ec18 · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues?.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SWE-bench: Can language models resolve real-world GitHub issues?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.401263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7f838f57aeff9543e35ce49e1f84d6d62b584cf9e6088ae97c2cbdc970f5e8f8

Observation 1f0725ff-4b9e-447f-b1c5-f27a6c1c8728 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.398826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d4abd03180a66624810e62b707d5f73a96406f69af804b1a928bb53cfecf6923

Observation 63c1db74-4e01-44c4-8c8c-8f522b8203c8 · outbound

This paper cites SpecOps: A fully automated AI agent testing framework in real-world GUI environments,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents SpecOps: A fully automated AI agent testing framework in real-world GUI environments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:2d1c8a18094fac8efa69dbac5a8ea5cc8bb1ae9291e5fd8707e56f6aa9c0157f

Observation daea19d1-e896-4fbd-9314-21ce8a47e832 · outbound

This paper cites STELLAR: A search- based testing framework for large language model applications.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents STELLAR: A search- based testing framework for large language model applications

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.967164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:991410139a5f80b109e24f3b41744198f758fc27180e7d24aad961277a8814a3

Observation 6143456a-537e-4313-8166-7a7f273e6ff7 · outbound

This paper cites an unresolved cited work.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-10T00:06:38.420670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9f9d7fc26a7a1d3407e8f6a8e6d26a0ab328e4dace65b6c7796bafe4115c45b8

Observation aec1170e-470f-4be9-8296-7486691ec429 · outbound

This paper cites A practitioner’s guide to process mining: Limitations of the directly-follows graph,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A practitioner’s guide to process mining: Limitations of the directly-follows graph,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.422481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:50310ed902f27889a367b30f9376f66c4c40dd61192b9db0af56add25998c8af

Observation 57d39246-2796-4bff-b036-68890eadf922 · outbound

This paper cites τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ 3-bench: From text-only to multimodal, knowledge- aware agent evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.429430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ca4af5697303270075bac2f912b3b7ed5291b5b1f0f109f21aa8bba467b88b25

Observation e5e93c59-a7ef-4463-9592-d07d6a3f5697 · outbound

This paper cites Event abstraction for process mining using supervised learning techniques,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction for process mining using supervised learning techniques,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.424048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7ac882d56d9c2918e2d698a0a3b6cbcbceac368510407ec24757568ed286a535

Observation 94bedfb8-7f23-49ae-8e03-092a82ad01ac · outbound

This paper cites Event abstraction in process mining: Literature review and taxonomy,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Event abstraction in process mining: Literature review and taxonomy,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:40a72cc56b64f2c80d2c4446438e5efa4b64f304ae52bad43ef6bdc834b0b55f

Observation a3e665c4-f53a-48ac-830b-15ce92ec3fb1 · outbound

This paper cites Boundary value exploration for software analysis,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Boundary value exploration for software analysis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.413741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:e9251c02253f4f6f01ce4b48f17c8abb1714e1f0618ad9250e39c39af01cde87

Observation 94441f16-9357-421f-bc16-0c2ed032bfb8 · outbound

This paper cites Automated robustness testing of off-the-shelf software components,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automated robustness testing of off-the-shelf software components,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.427615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4b2cf45df2cd6cfe14ab3bbc82ce8f4708cfcfd5c7954c6c8a191b79de178699

Observation 954f9e5c-be91-4ff4-8818-3fd3c6f4b45d · outbound

This paper cites τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents τ- knowledge: Evaluating conversational agents over unstructured knowl- edge,

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.976247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:11534a2c99cba890f6de9f03af239b42c6b2737d87664916b21ec878e2941cdd

Observation 528726a0-4bbc-43f0-82da-02e1076c96fc · outbound

This paper cites ToolLLM: Facilitating large language models to master 16000+ real-world APIs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ToolLLM: Facilitating large language models to master 16000+ real-world APIs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.391348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:46e1b1f8be8cb4d7a3e5dd0b986c615a4009ca1498c40eae1ff0f5a67c51bb8d

Observation 313d6d59-d163-401c-9181-18e859c0f448 · outbound

This paper cites Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Towards self-evolving benchmarks: Synthesizing agent trajectories via test-time exploration under validate-by-reproduce paradigm

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.963929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d330866c5b170cb31665021640867df1609939a492a4f4bbb2fbff4d3ded1385

Observation 97062318-6ecb-476d-8c13-da94c9bc37b9 · outbound

This paper cites Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Revisiting benchmark and assessment: An agent-based exploratory dynamic evaluation framework for llms

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.975111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:00d5619006ec9e83367e01af0484bda7338ca59b2343c74931d2ec69930b5dae

Observation 86574250-54b4-41bf-a838-365082b4f56f · outbound

This paper cites Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Graph2Eval: Automatic multimodal task generation for agents via knowledge graphs,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-10T00:06:37.973251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:5ba95bfc6d84870f8c73c9bd8d3024033907ef0ea11e40d30c91edc9ed1e6dcb

Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · outbound

This paper cites Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:06:37.956842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f07afc23ff28ba80f46c5b866e4e8b924b9ef6ff0b6daeee4f2ea3c0244fac54

Observation bdd38177-ce3c-4b83-9b07-36928981b4da · outbound

This paper cites Mining specifications,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Mining specifications,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:3a6a58bea56a303abf9b5999ae5240d0492f17479cf17f794f19a76671f0babc

Observation 84bb6543-3b44-4012-a793-c6147a4fb1f0 · outbound

This paper cites Discovering models of software processes from event-based data,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discovering models of software processes from event-based data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.395137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:eec1c26d210cb1e92275ae2b1b43555256b329040b0a96d26e8e4a8bf77d17e3

Observation f89dc7ec-340b-4e20-b62f-18d00ca50c54 · outbound

This paper cites Automatic generation of software behavioral models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Automatic generation of software behavioral models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:71fb393b6eeb51b9964381c533c39cd0ae81453f078640b42063a25982102c22

Observation 2dc6c576-bf93-4be9-bcc0-4dc77de44aad · outbound

This paper cites Inferring models of concurrent systems from logs of their behavior with CSight,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Inferring models of concurrent systems from logs of their behavior with CSight,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:b0ee6ad848342677e52f6dd7b09bb3416d7a0833f7cf63044e5e81727bc2375b

Observation a0c731b5-5311-457b-9da7-7e14474b33d1 · outbound

This paper cites Workflow mining: Discovering process models from event logs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Workflow mining: Discovering process models from event logs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:10b22b84ae6c3b18013ebe8405468abda2acf7c9a676300150b489049ec030f3

Observation 18aa5651-0695-48e9-822c-646501335467 · outbound

This paper cites Discov- ering block-structured process models from event logs—a constructive approach,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Discov- ering block-structured process models from event logs—a constructive approach,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.425884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:88f32afa6ef1a152a513ee13c3efa6f5d78d12468a61c37d219028229ffc1c70

Observation 1f2e5127-a857-461b-abc0-1219399aa20c · outbound

This paper cites Applying graph reduction techniques for identifying structural conflicts in process models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Applying graph reduction techniques for identifying structural conflicts in process models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:afd8fe07a8df4b7dc1600c295432844cbf36112d19d1e357936c5ce7f625ff30

Observation d81ef0ce-87de-4eb1-a2a2-1871b5a90b68 · outbound

This paper cites Learning regular sets from queries and counterexamples,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Learning regular sets from queries and counterexamples,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.431693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:f51c5a4de4e37a0ce15bc4dc29bedd4cdfdfbf49a5d0ea4d4fc9b063082defe8

Observation 9e496994-d97f-432f-b6da-538d382c9579 · outbound

This paper cites Unsupervised dialog structure learning,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Unsupervised dialog structure learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:4abe2b3dba460886a6b233cfac2d3458b23b890b2cd70d6272198b6925896986

Observation 8b21782c-e4f6-4cf1-8c76-f9e32d4497df · outbound

This paper cites Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Dialog2Flow: Pre-training soft-contrastive action-driven sentence embeddings for automatic dia- log flow extraction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.393236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d088b3b08089cf7581e406125aecfd47ae6cd788b63acf7e3ef9ae2b358f08e9

Observation 47d02b58-9eb9-4c9d-8561-ca85fee5e5e5 · outbound

This paper cites Agenda-based user simulation for bootstrapping a POMDP dialogue system,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Agenda-based user simulation for bootstrapping a POMDP dialogue system,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.404999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:68f4903bf0068104afdf5169fe182b546b4af56bad567114e862c5961c0d3cef

Observation 8b43514d-7a55-415b-94e7-19673430fb4d · outbound

This paper cites ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents ConvLab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.407241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:fec715486c4c855e8977a21e1772fea91ad4624c46efd0a674e40e250396a587

Observation 2fc55fc5-d849-4b04-aa23-93c0f960bcdb · outbound

This paper cites A taxonomy of model- based testing approaches,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents A taxonomy of model- based testing approaches,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.385993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7651caf4f5d21c572bbd6bf1902e2dc32786bac06ccd7b5d295834cd7fbf0119

Observation 75ce28d8-4ae0-40ca-ba6d-2e5f3c4d9ca9 · outbound

This paper cites Principles and methods of testing finite state machines—a survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Principles and methods of testing finite state machines—a survey,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.389665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:47536486338b40bbbd9583d0f192223537dc97d37cbbdd97485aaa39df90934a

Observation f73ca176-6c76-43e6-9d34-280c4bf33ab5 · outbound

This paper cites Testing software design modeled by finite-state machines,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing software design modeled by finite-state machines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.396696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ce2af597230ccc22cfdccf79c02367236d087752871a799832a648d99eb488c1

Observation 25e0e7b8-c0d7-4dfd-bdc5-0355e81b625d · outbound

This paper cites RESTler: Stateful REST API fuzzing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTler: Stateful REST API fuzzing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.384367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:86e18e5c8c4c5d1171eadb3f9790ac2b821996c5c9594b4ace1094a2b0e80581

Observation 413c8bf1-30d6-41dd-9201-7823baa165db · outbound

This paper cites RESTful API automated test case generation with Evo- Master,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents RESTful API automated test case generation with Evo- Master,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.373884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:d2023d87625e48d96def7438e5eb0d2b3a8256a068ef0a19e30fadebde9c343a

Observation 9c71df55-6f0d-40e5-86cb-c486edd2c362 · outbound

This paper cites Morest: Model-based RESTful API testing with execution feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Morest: Model-based RESTful API testing with execution feedback,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.351133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:bc47978c1e1d802a513dc8c8766634d7a3f2311964edd557eb17844897bf7ab2

Observation fdfc078e-4aac-477b-a422-f5bbcb06dd86 · outbound

This paper cites KAT: Dependency-aware automated API testing with large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents KAT: Dependency-aware automated API testing with large language models,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.375692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0daa4af15c3e7183a30c9fb6a6d88e9ed102b76cf54ce67bd93a49ecf3326c0a

Observation 5ccd9764-a37b-474b-b12d-7afdcb0079dd · outbound

This paper cites Testing RESTful APIs: A survey,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Testing RESTful APIs: A survey,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.408976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:386e636bc49b866907a16e3e64a2a2ea0087c0ee1ee5750938ad6c7fbcb25066

Observation 2b3b289f-f2e8-4e78-98d5-f07c0cd671c2 · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents An empirical evaluation of using large language models for automated unit test generation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.412345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:0e627f7e14b8d792b07ec47565ba285b5648a4b1443a7b7b3297c22b1a013f97

Observation 5352081e-274f-4b47-bb01-3889218a77e5 · outbound

This paper cites CoverUp: Effective high coverage test generation for Python,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CoverUp: Effective high coverage test generation for Python,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a1607fe6f669b239b84ec91710b60208d7c000679a366890fce3b16f66e7aafc

Observation a19faacc-39ec-42e5-9c38-01a65f71e8cd · outbound

This paper cites Evaluating and improving ChatGPT for unit test generation,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Evaluating and improving ChatGPT for unit test generation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.417180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:de54bf2c51317e2a43ed4e8ea91e9eb6af7414724bc0d1fa4e079d2180bac607

Observation c6508104-bd2e-4132-b3b5-e0026d11b7f7 · outbound

This paper cites CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CodaMosa: Escaping coverage plateaus in test generation with pre-trained large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.410662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a27d21dad7c615f01c683650923709d0d77441f4515c63e3b2dea28bc097c26d

Observation 74a8218e-fa88-4365-9149-1b2c1b93a272 · outbound

This paper cites TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:299d7e47f53fe5c7955e994fc5e796ed5e12db5d8125b6fb5872be00234e4a55

Observation d2ee0e6d-23e2-41d1-865d-8836ec6aaeef · outbound

This paper cites The oracle problem in software testing: A survey.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents The oracle problem in software testing: A survey

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.362723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:7a59795a4aba2dbb561e0ec4c9f6fcb0ed620cf547038c72430e0cfd888d340e

Observation 91b24357-a3aa-40bb-a95e-781f191ad93b · outbound

This paper cites Pseudo-oracles for non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Pseudo-oracles for non-testable programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.368635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:a2587a6de7666faf48569f4926b8839ffe21ed2d491eb72c4666f80abe06e1dc

Observation e755b481-bc90-4bad-9642-d31ac5d8b987 · outbound

This paper cites On testing non-testable programs,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On testing non-testable programs,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.356276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:eded76b5a34bbf4be21008d1fff076bfc6cb36f4fb912c30f34b6de14c78772f

Observation 5280139e-b263-4a46-9432-063d6ccc5e10 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.379148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:776e51d5846c44c667ba516c43d7c22b59dfe1c00429a0a96ed4b6aea895a20a

Observation 8919763f-e676-4082-b6e1-3691f34f621c · outbound

This paper cites Large language models cannot self-correct reasoning yet.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Large language models cannot self-correct reasoning yet

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.415797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:aa320b487166dbde806b0f960d76dcee45270ecd6554ed44b9cec7b37ee08528

Observation 56f5958e-caa8-4320-9086-ead614e1fd1c · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:06:37.970034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:61c36706047478854fbc416fe39ede2aa2494d47eac2e8ff89605b625a33ac70

Observation fc6e1223-e3df-467e-90f8-89dcda18765c · outbound

This paper cites Self-Refine: Iterative refinement with self-feedback,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Self-Refine: Iterative refinement with self-feedback,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.372182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9e8728434f24ba382c83471b1219f9eef9cad98e5abdf1df70044e420f56e447

Observation ffd8ae57-7bb3-41f1-a318-d4876be7b2df · outbound

This paper cites CRITIC: Large language models can self-correct with tool-interactive critiquing,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents CRITIC: Large language models can self-correct with tool-interactive critiquing,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.377391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:ffed36c72cdf1275bc75a407380ee5a21c932a95ee6eb730e0b0e161418323d9

Observation 4b153557-3da1-4ffe-b1b5-bd2cff717884 · outbound

This paper cites Metamorphic testing: A review of challenges and opportunities,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Metamorphic testing: A review of challenges and opportunities,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.398672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9bb436c9a048c9c05f8d0670adeeb2be806733cb5280d51baebe81d87cf4d66d

Observation 45b02c00-7cdd-4acc-b72b-106b5381a916 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Not what you’ve signed up for: Compromising real-world LLM- integrated applications with indirect prompt injection,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.403085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:8037808aa6a7e462066cd87b881e2bbd1f9970c75ef33a27e3c24806d39217a1

Observation 587a72f6-11a9-45de-b57e-fa05e453fc52 · outbound

This paper cites Ignore previous prompt: Attack techniques for language models.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Ignore previous prompt: Attack techniques for language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T00:06:38.419028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:5a4f39fc37c0965aa1d5bd3b5147408b11b086d79a72f664795a7771f81f59e8

Pith citing papers

No inbound Pith citation observations are available.