Prosa demonstrates that rubric-based binary scoring with multi-judge filtering yields full agreement on 16 LLM rankings across judges on Brazilian Portuguese chats, compared to only 7/16 under holistic scoring, while widening score gaps by 47%.
Title resolution pending
7 Pith papers cite this work, alongside 12 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Uncertainty estimation and regularization on weak positive pairs improves mAP by 3.06%, 3.55%, and 6.94% on CUHK-PEDES, RSTPReid, and ICFG-PEDES respectively.
LLMs achieve only 59.7% role identification accuracy in Secret Hitler versus 86.7% for rule-based agents, show negative impact as fascists, and produce 40% shorter games due to failed deception.
Computational analysis of Rossini's multiple settings of 'Mi lagnerò tacendo' uses parsing, data mining, and graph theory to explore melodic, harmonic, and textual choices as a foundation for philological research and generative models.
FedKLPR introduces KL-divergence-guided training, pruning-aware weighted aggregation, and cross-round recovery to achieve 40-42% communication reduction on ResNet-50 while preserving competitive accuracy in federated person re-identification across eight datasets.
The paper introduces a taxonomy of AI safety for LLMs organized into Trustworthy AI, Responsible AI, and Safe AI perspectives, accompanied by a review of state-of-the-art methods, challenges, and future directions.
Random Forest achieves 99.9% accuracy, precision, recall and F1-score for fraud detection on a 101k-record telecom CDR dataset after Min-Max scaling and SMOTE.
citing papers explorer
-
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
Prosa demonstrates that rubric-based binary scoring with multi-judge filtering yields full agreement on 16 LLM rankings across judges on Brazilian Portuguese chats, compared to only 7/16 under holistic scoring, while widening score gaps by 47%.
-
Harnessing Weak Pair Uncertainty for Text-based Person Search
Uncertainty estimation and regularization on weak positive pairs improves mAP by 3.06%, 3.55%, and 6.94% on CUHK-PEDES, RSTPReid, and ICFG-PEDES respectively.
-
Evaluating Large Language Models in a Complex Hidden Role Game
LLMs achieve only 59.7% role identification accuracy in Secret Hitler versus 86.7% for rule-based agents, show negative impact as fascists, and produce 40% shorter games due to failed deception.
-
Advanced Scientific Methodology Plays Rossini
Computational analysis of Rossini's multiple settings of 'Mi lagnerò tacendo' uses parsing, data mining, and graph theory to explore melodic, harmonic, and textual choices as a foundation for philological research and generative models.
-
FedKLPR: KL-Guided Pruning-Aware Federated Learning for Person Re-Identification
FedKLPR introduces KL-divergence-guided training, pruning-aware weighted aggregation, and cross-round recovery to achieve 40-42% communication reduction on ResNet-50 while preserving competitive accuracy in federated person re-identification across eight datasets.
-
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions
The paper introduces a taxonomy of AI safety for LLMs organized into Trustworthy AI, Responsible AI, and Safe AI perspectives, accompanied by a review of state-of-the-art methods, challenges, and future directions.
-
An Efficient Machine Learning-based Framework for Detection and Prevention of Frauds in Telecom Networks
Random Forest achieves 99.9% accuracy, precision, recall and F1-score for fraud detection on a 101k-record telecom CDR dataset after Min-Max scaling and SMOTE.