Large-scale analysis of unfiltered search queries shows geospatial intent at 18% of total, dominated by transactional categories outside traditional GIS scope.
hub
arXiv preprint arXiv:2209.11055 , year=
14 Pith papers cite this work, alongside 101 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Hybrid quantum training discovers parity bases that improve accuracy 24-42% on binary tasks and recover performance on text benchmarks, with all inference remaining classical.
An empirical study of 4chan and Reddit data shows that users of varying technical skill share primary resources for SNCII creation and secondary resources for dissemination, with knowledge transfer from experts to newcomers enabling spread.
A pipeline classifies conversational AI questions to curriculum topics via a prerequisite graph, achieving 80% accuracy on 1,340 questions and correlating topic question volume with self-reported difficulty (rho=0.491).
CR4T is a model-agnostic framework using lightweight risk detection and domain-conditioned rewriting to convert unsafe or refusal-style LLM responses into developmentally appropriate guidance for adolescents.
Topic models can act as binary classifiers for retrieving news on extreme climate events in German media, with performance varying by hazard type and boosted by keyword probabilities.
Trapped-ion quantum fine-tuning of AI models shows linear energy scaling and 24% better classification error than classical logistic regression or SVM baselines, with a projected energy break-even at 34 qubits.
FlaXifyer applies few-shot learning on pre-trained language models to categorize intermittent CI job failures from logs at 84.3% Macro F1 and 92.0% Top-2 accuracy using 12 examples per category, with LogSift reducing log review effort by 74.4%.
ModernBERT is a new bidirectional encoder model achieving SOTA performance on diverse classification and retrieval benchmarks while offering superior speed and memory efficiency for long-context inference.
Fine-tuned compact models achieve strong multilingual performance and large efficiency gains over LLMs on production data from 114 languages for claim detection and 28 for veracity prediction.
Opir introduces efficient multi-task encoder models trained on a 996-category safety taxonomy that match or exceed larger baselines on most safety benchmarks while using under 100M parameters for edge variants.
The paper introduces a multi-agent LLM framework using ReAct-style agents with self-reflection for classifying telecom queries, anonymizing PII via k-anonymity and differential privacy, and translating expert responses for end users.
LoRA-MME ensembles LoRA-adapted UniXcoder, CodeBERT, GraphCodeBERT, and CodeBERTa with learned weights to reach 0.7906 weighted F1 and 0.6867 macro F1 on code comment classification.
Decomposing automotive query understanding into a lightweight classification stage followed by specialized entity extraction yields better accuracy and lower latency than joint single-step processing.
citing papers explorer
-
Much of Geospatial Web Search Is Beyond Traditional GIS
Large-scale analysis of unfiltered search queries shows geospatial intent at 18% of total, dominated by transactional categories outside traditional GIS scope.
-
Quantum Parity Representations: Learnable Basis Discovery, Encoders, and Shadow Deployment
Hybrid quantum training discovers parity bases that improve accuracy 24-42% on binary tasks and recover performance on text benchmarks, with all inference remaining classical.
-
Characterizing Resource Sharing Practices on Underground Internet Forum Synthetic Non-Consensual Intimate Image Content Creation Communities
An empirical study of 4chan and Reddit data shows that users of varying technical skill share primary resources for SNCII creation and secondary resources for dissemination, with knowledge transfer from experts to newcomers enabling spread.
-
Detecting Knowledge Gaps from Conversational AI Interactions Using Curriculum Prerequisite Graphs
A pipeline classifies conversational AI questions to curriculum topics via a prerequisite graph, achieving 80% accuracy on 1,340 questions and correlating topic question volume with self-reported difficulty (rho=0.491).
-
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
CR4T is a model-agnostic framework using lightweight risk detection and domain-conditioned rewriting to convert unsafe or refusal-style LLM responses into developmentally appropriate guidance for adolescents.
-
Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News
Topic models can act as binary classifiers for retrieving news on extreme climate events in German media, with performance varying by hazard type and boosted by keyword probabilities.
-
Measuring Accuracy and Energy-to-Solution of Quantum Fine-Tuning of Foundational AI Models
Trapped-ion quantum fine-tuning of AI models shows linear energy scaling and 24% better classification error than classical logistic regression or SVM baselines, with a projected energy break-even at 34 qubits.
-
Predicting Intermittent Job Failure Categories for Diagnosis Using Few-Shot Fine-Tuned Language Models
FlaXifyer applies few-shot learning on pre-trained language models to categorize intermittent CI job failures from logs at 84.3% Macro F1 and 92.0% Top-2 accuracy using 12 examples per category, with LogSift reducing log review effort by 74.4%.
-
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
ModernBERT is a new bidirectional encoder model achieving SOTA performance on diverse classification and retrieval benchmarks while offering superior speed and memory efficiency for long-context inference.
-
Multilingual Fact-Checking at Scale: Fine-Tuned Compact Models vs LLMs
Fine-tuned compact models achieve strong multilingual performance and large efficiency gains over LLMs on production data from 114 languages for claim detection and 28 for veracity prediction.
-
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
Opir introduces efficient multi-task encoder models trained on a 996-category safety taxonomy that match or exceed larger baselines on most safety benchmarks while using under 100M parameters for edge variants.
-
Cross-Domain Query Translation for Network Troubleshooting: A Multi-Agent LLM Framework with Privacy Preservation and Self-Reflection
The paper introduces a multi-agent LLM framework using ReAct-style agents with self-reflection for classifying telecom queries, anonymizing PII via k-anonymity and differential privacy, and translating expert responses for end users.
-
LoRA-MME: Multi-Model Ensemble of LoRA-Tuned Encoders for Code Comment Classification
LoRA-MME ensembles LoRA-adapted UniXcoder, CodeBERT, GraphCodeBERT, and CodeBERTa with learned weights to reach 0.7906 weighted F1 and 0.6867 macro F1 on code comment classification.
-
Domain-Specific Query Understanding for Automotive Applications: A Modular and Scalable Approach
Decomposing automotive query understanding into a lightweight classification stage followed by specialized entity extraction yields better accuracy and lower latency than joint single-step processing.