Density-based outliers in labor market text act as leading indicators of new occupational clusters, with an extended Emerging Occupation Score predicting formation 2 quarters ahead at F1=0.74 on 84,988 postings.
hub Mixed citations
Warren, Lu Cheng, Haidar M
Mixed citation behavior. Most common role is background (56%).
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Domain adaptation via synthetic manuscript images improves OMR performance on real-world piano manuscripts without requiring in-domain symbols.
PackSELL packs delta-encoded indices and values into single words with tunable bit allocation, delivering up to 1.63x faster FP16 SpMV and FP32-accurate performance exceeding FP16 cuSPARSE while reducing memory traffic.
Seven clinician-informed safety criteria enable LLM-as-a-Judge to reach substantial agreement with human consensus (Cohen's κ up to 0.75) on evaluating LLM responses to users demonstrating psychosis.
AlphaEvolve is an LLM-orchestrated evolutionary coding agent that discovered a 4x4 complex matrix multiplication algorithm using 48 scalar multiplications, the first improvement over Strassen's algorithm in 56 years, plus optimizations for Google data centers and hardware.
A new dataset of 1,200 dynamic hologram attack videos and a background-subtraction-based verification method achieve state-of-the-art detection of unseen dynamic attacks on identity documents.
Hamm-grams are a new class of fixed-length regular expressions over bytes with single-character wildcards, mined efficiently with LSH and clustering to yield more robust features than n-grams for malware classification and detection.
TabPATE applies a PATE-style private aggregation to synthetic tabular queries generated from feature ranges, enabling private in-context learning with near-random membership inference success while keeping competitive utility.
LLM compression of filings and earnings calls often changes the source-implied bear/neutral/bull decision; agentic multi-candidate auditing against the source reduces those flips.
Cross-lingual prompt exploration improves factual recall and consistency in LLMs across 17 languages more efficiently than native-language scaling.
Automatically optimizing agent skill files on a branching lakehouse improved held-out validation accuracy by 31.9% on 25 synthetic-but-trace-anchored tasks.
PIPER retrieves and ranks tabular datasets by profiling their content and using LLM-generated queries for dense vector search, outperforming metadata baselines and TableQA methods in low-metadata settings.
Macro uses DPO on composite preference pairs to raise validity of multilingual self-generated counterfactual explanations by 12.55% on average over chain-of-thought while preserving minimality.
SPARK improves LLM-based test code fault localization by retrieving similar past faults and selectively annotating suspicious lines in new failing tests.
Methods for constructing Hypergraphs of Text are proposed with a new effort ratio metric where TF-IDF baselines match LLM methods in experiments.
A new catalog classifying 35 data error types into missing, incorrect, and redundant categories for tabular data, with definitions and examples to improve data quality management.
MONETA is the first multimodal benchmark for industry classification using text and geographic sources, with MLLM baselines at 62-74% accuracy and up to 22.8% gains from multi-turn context enrichment and explanations.
Fine-tuned LLaMA-3.1-8B generates valid, plausible counterfactuals for sensor-based stress prediction, and augmenting with them recovers ~20% of F1 under label scarcity.
The paper introduces the InsideOut benchmark to quantify insider-outsider bias in LLM-generated interview scripts across 10 cultures and shows that multi-agent mitigation frameworks substantially reduce the bias on metrics like Cultural Alignment Gap.
Single-agent LLM frameworks outperform naive multi-agent systems in multimodal clinical risk prediction tasks and are better calibrated.
Context-mediated domain adaptation treats user modifications to AI artifacts as implicit domain specifications that reshape LLM-powered multi-agent reasoning, demonstrated via the Seedentia system which extracted 46 domain knowledge entries from expert edits.
Weakly supervised ML classifier and hypothesis-testing signature mining detect LDAP reconnaissance at 65% TPR and 81.48% field precision.
Proposes an AI-driven synthetic data generation framework to create realistic cybersecurity datasets for smart city research where real data is scarce or sensitive.
This perspective paper calls for a research program treating LLMs as consequential social actors whose outputs influence human decisions, norms, and collective dynamics.
citing papers explorer
-
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization
Macro uses DPO on composite preference pairs to raise validity of multilingual self-generated counterfactual explanations by 12.55% on average over chain-of-thought while preserving minimality.