H-SAL erases latent concepts from text profiles using self-descriptions as implicit debiasing signals and shows competitive performance on a new multi-domain Stack Exchange helpfulness benchmark.
arXiv preprint arXiv:2306.05949 , year=
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Audit of ChatGPT, Copilot, Gemini and Perplexity finds ~16% of cited sources are AI-generated across 712 queries on politics, health and environment.
Big AI captures regulation through 27 mechanisms in five categories, most commonly via discourse influence and law elusion, often justified by narratives that regulation stifles innovation or serves national interest.
TICoE achieves more precise and faithful concept erasure in text-to-image models by collaborating text and image data through a convex manifold and hierarchical learning, outperforming prior methods.
Analysis of 499 generative AI incidents shows use-related failures predominate and frequently harm non-users, producing a distinct risk profile from traditional AI.
SEAL uses semantic embeddings and locality-sensitive hashing to create distortion-free, database-free watermarks for generative images that are conditioned on content for improved forgery resistance.
Interviews reveal a four-stage vibe coding workflow that accelerates prototyping while introducing tensions between quick efficiency and reflective design intention, plus asymmetries in trust and ownership.
Survey organizes LLM trustworthiness into seven categories and 29 sub-categories, measures eight sub-categories on popular models, and finds that more aligned models generally score higher but with varying effectiveness.
Frontier AI safety policies have a structural coordination gap caused by diffuse benefits and concentrated costs, which can be addressed by adapting precommitment and shared response protocols from other high-risk domains.
Stable Diffusion augments limited indoor scene datasets for better recognition models, and DIRE detects the generated images with 100% accuracy using lightweight classifiers.
citing papers explorer
-
Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles
H-SAL erases latent concepts from text profiles using self-descriptions as implicit debiasing signals and shows competitive performance on a new multi-domain Stack Exchange helpfulness benchmark.
-
Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
Audit of ChatGPT, Copilot, Gemini and Perplexity finds ~16% of cited sources are AI-generated across 712 queries on politics, health and environment.
-
Big AI's Regulatory Capture: Mapping Industry Interference and Government Complicity
Big AI captures regulation through 27 mechanisms in five categories, most commonly via discourse influence and law elusion, often justified by narratives that regulation stifles innovation or serves national interest.
-
Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration
TICoE achieves more precise and faithful concept erasure in text-to-image models by collaborating text and image data through a convex manifold and hierarchical learning, outperforming prior methods.
-
A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents
Analysis of 499 generative AI incidents shows use-related failures predominate and frequently harm non-users, producing a distinct risk profile from traditional AI.
-
SEAL: Semantic Aware Image Watermarking
SEAL uses semantic embeddings and locality-sensitive hashing to create distortion-free, database-free watermarks for generative images that are conditioned on content for improved forgery resistance.
-
Vibe Coding in Product Teams: Reconfiguring AI-Assisted Workflows, Prototyping, and Collaboration
Interviews reveal a four-stage vibe coding workflow that accelerates prototyping while introducing tensions between quick efficiency and reflective design intention, plus asymmetries in trust and ownership.
-
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Survey organizes LLM trustworthiness into seven categories and 29 sub-categories, measures eight sub-categories on popular models, and finds that more aligned models generally score higher but with varying effectiveness.
-
The coordination gap in frontier AI safety policies
Frontier AI safety policies have a structural coordination gap caused by diffuse benefits and concentrated costs, which can be addressed by adapting precommitment and shared response protocols from other high-risk domains.
-
Rethinking Text-to-Image as Semantic-Aware Data Augmentation for Indoor Scene Recognition
Stable Diffusion augments limited indoor scene datasets for better recognition models, and DIRE detects the generated images with 100% accuracy using lightweight classifiers.