A GenAI-powered causal inference framework uses deep generative model representations as learned deconfounders to identify dynamic causal effects of video features on real-time outcomes.
J., Ting, D
12 Pith papers cite this work, alongside 3,325 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 12roles
background 2polarities
background 2representative citing papers
A GenAI-based method extracts representations from unstructured data and uses a neural network to fit marginal structural models that recover causal effects of treatment feature sequences including their positions.
A novel FMECA-based framework was developed and validated for systematic assessment of patient safety risks in LLM-generated clinical discharge summaries, demonstrating moderate-to-substantial inter-rater agreement and good usability.
Large language models exhibit normative conformity in addition to informational conformity, and subtle social context can direct which group they conform to.
LLMs drop from 71.1% to 38.0% accuracy on medical questions when misleading context is injected, measured via new MedMisBench benchmark with 10,932 items.
CuraView detects sentence-level faithfulness hallucinations in medical discharge summaries via GraphRAG knowledge graphs and multi-agent evidence grading, achieving 0.831 F1 on critical contradictions with a fine-tuned Qwen3-14B model and 50% relative improvement over baselines.
LLM-powered conversational voice sleep diaries achieved higher adherence and richer contextual reports than text-based diaries, with a noted trade-off in structured field completeness.
Lightweight LLMs lose 7.2 percentage points of accuracy when false health claims are injected into prompts, but only 1.4 points when medical jargon is replaced with everyday language.
The system integrates a Neo4j knowledge graph, four-stage symptom matching with LLM verification, genetic-algorithm-optimized proactive questioning, and multimodal evidence-based visualizations to improve diagnostic transparency and treatment interpretability in TCM, reporting 32% fewer non-standard
ARSM-Agent cuts attack success rate to 8.7% and reaches 0.91 knowledge consistency in medical tasks by linking risk perception, evidence retrieval, consistency verification, and confidence reweighting.
In a Dutch academic hospital pilot, LLM drafts were copied into 58.5% of discharge summaries, 86.9% of users reported reduced documentation time, and 91.3% intended to keep using the tool.
A hybrid RAG system with retrieval, Cohere reranking, and claim-level LLM judgment achieves 100% grounding accuracy on 200 claims from 25 biomedical queries in a pilot study.
citing papers explorer
-
Causal Inference with Video Features as Treatments
A GenAI-powered causal inference framework uses deep generative model representations as learned deconfounders to identify dynamic causal effects of video features on real-time outcomes.
-
GenAI Powered Dynamic Causal Inference with Unstructured Data
A GenAI-based method extracts representations from unstructured data and uses a neural network to fit marginal structural models that recover causal effects of treatment feature sequences including their positions.
-
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content
A novel FMECA-based framework was developed and validated for systematic assessment of patient safety risks in LLM-generated clinical discharge summaries, demonstrating moderate-to-substantial inter-rater agreement and good usability.
-
Large Language Models Exhibit Normative Conformity
Large language models exhibit normative conformity in addition to informational conformity, and subtle social context can direct which group they conform to.
-
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
LLMs drop from 71.1% to 38.0% accuracy on medical questions when misleading context is injected, measured via new MedMisBench benchmark with 10,932 items.
-
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
CuraView detects sentence-level faithfulness hallucinations in medical discharge summaries via GraphRAG knowledge graphs and multi-agent evidence grading, achieving 0.831 F1 on critical contradictions with a fine-tuned Qwen3-14B model and 50% relative improvement over baselines.
-
Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep
LLM-powered conversational voice sleep diaries achieved higher adherence and richer contextual reports than text-based diaries, with a noted trade-off in structured field completeness.
-
Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications
Lightweight LLMs lose 7.2 percentage points of accuracy when false health claims are injected into prompts, but only 1.4 points when medical jargon is replaced with everyday language.
-
Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation
The system integrates a Neo4j knowledge graph, four-stage symptom matching with LLM verification, genetic-algorithm-optimized proactive questioning, and multimodal evidence-based visualizations to improve diagnostic transparency and treatment interpretability in TCM, reporting 32% fewer non-standard
-
Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks
ARSM-Agent cuts attack success rate to 8.7% and reaches 0.91 knowledge consistency in medical tasks by linking risk perception, evidence retrieval, consistency verification, and confidence reweighting.
-
Phase 1 Implementation of LLM-generated Discharge Summaries showing high Adoption in a Dutch Academic Hospital
In a Dutch academic hospital pilot, LLM drafts were copied into 58.5% of discharge summaries, 86.9% of users reported reduced documentation time, and 91.3% intended to keep using the tool.
-
A Hybrid Retrieval and Reranking Framework for Evidence-Grounded Retrieval-Augmented Generation
A hybrid RAG system with retrieval, Cohere reranking, and claim-level LLM judgment achieves 100% grounding accuracy on 200 claims from 25 biomedical queries in a pilot study.