Pandora's Regret is a closed-form pairwise scoring rule derived from expected optimal search costs that elicits true probabilities and outperforms log loss, accuracy, and F1 at predicting diagnostic costs on MedMNIST models.
hub
Preference Learning with Gaussian Processes
16 Pith papers cite this work, alongside 2,759 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
unclear 1representative citing papers
LOGICA adds context to pretrained biological LMs via logit-space contrastive alignment with gated adapters, improving AUC on held-out drug-resistance mutation ranking from ~0.55 to ~0.65 while preserving token likelihoods.
RankGuard is a decentralized OLTR system that filters model updates using local click data for poisoning resistance and supplies the first formal convergence guarantee for decentralized OLTR.
TAPS converts diffusion marginal probabilities into path-conditioned acceptance estimates to select prefix-closed subtrees under a fixed verification budget, achieving up to 7.9x end-to-end speedup over autoregressive decoding.
Latent Heuristic Search performs continuous optimization over learned embeddings of heuristics, using normalizing flows and LLM prompting to discover competitive solvers for TSP, CVRP, KSP, and OBP.
Acoustic features from narration show a robust association with audiobook appeal independent of title effects, based on analysis of LibriVox data and proprietary metrics.
Scoring functions are sub-optimal for all utility-fairness trade-offs in ranking under a generic fairness formulation, but semi-greedy post-processing can approach the performance of exhaustive post-processing.
Adaptive Prompt Elicitation (APE) uses an information-theoretic framework to generate visual queries that elicit and compile user intent into better prompts for text-to-image models, showing improved alignment in benchmarks and a user study.
PREFAB applies preference learning grounded in the peak-end rule to let users annotate only key affective change segments while interpolating the rest, reducing workload and improving confidence in a 25-participant study.
LILO integrates LLMs to translate natural language feedback into preference signals for Gaussian process-based Bayesian optimization, outperforming standard preference BO and LLM-only methods on benchmarks.
A bilevel Proxy-FEA diagnostic framework is introduced and tested on a simplified LDED32 stripe benchmark to reveal proxy misalignment with FEA labels and a stress-distortion trade-off in RL-guided scan-order optimization.
Trust calibration in agentic tool use is cast as preferential Bayesian optimization over a latent human risk-tolerance function observed through binary approve/deny feedback with a probit likelihood.
The paper formalizes query answering with soft constraints on knowledge graphs and introduces two lightweight methods (parameter tuning or small neural network) to incorporate them while preserving original rankings.
InternLM2 is a new open-source LLM that outperforms prior versions on 30 benchmarks and long-context tasks through scaled pre-training to 32k tokens and a conditional online RLHF alignment strategy.
Tutorial on a GP-based framework for preference and choice learning that unifies random utility models, limits of discernment, and multi-utility scenarios via customized likelihoods for object and label preferences.
A comprehensive survey of knowledge distillation for LLMs structured around algorithms, skill enhancement, and vertical applications, highlighting data augmentation as a key enabler.
citing papers explorer
-
Pandora's Regret: A Proper Scoring Rule for Evaluating Sequential Search
Pandora's Regret is a closed-form pairwise scoring rule derived from expected optimal search costs that elicits true probabilities and outperforms log loss, accuracy, and F1 at predicting diagnostic costs on MedMNIST models.
-
Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment
LOGICA adds context to pretrained biological LMs via logit-space contrastive alignment with gated adapters, improving AUC on held-out drug-resistance mutation ranking from ~0.55 to ~0.65 while preserving token likelihoods.
-
Efficient and Robust Online Learning to Rank in Decentralized Systems
RankGuard is a decentralized OLTR system that filters model updates using local click data for poisoning resistance and supplies the first formal convergence guarantee for decentralized OLTR.
-
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
TAPS converts diffusion marginal probabilities into path-conditioned acceptance estimates to select prefix-closed subtrees under a fixed verification budget, achieving up to 7.9x end-to-end speedup over autoregressive decoding.
-
Latent Heuristic Search: Continuous Optimization for Automated Algorithm Design
Latent Heuristic Search performs continuous optimization over learned embeddings of heuristics, using normalizing flows and LLM prompting to discover competitive solvers for TSP, CVRP, KSP, and OBP.
-
Audio-Based Understanding of Audiobook Narration Appeal
Acoustic features from narration show a robust association with audiobook appeal independent of title effects, based on analysis of LibriVox data and proprietary metrics.
-
Scoring Is Not Enough: Addressing Gaps in Utility-fairness Trade-offs for Ranking
Scoring functions are sub-optimal for all utility-fairness trade-offs in ranking under a generic fairness formulation, but semi-greedy post-processing can approach the performance of exhaustive post-processing.
-
Adaptive Prompt Elicitation for Text-to-Image Generation
Adaptive Prompt Elicitation (APE) uses an information-theoretic framework to generate visual queries that elicit and compile user intent into better prompts for text-to-image models, showing improved alignment in benchmarks and a user study.
-
PREFAB: PREFerence-based Affective Modeling for Low-Budget Self-Annotation
PREFAB applies preference learning grounded in the peak-end rule to let users annotate only key affective change segments while interpolating the rest, reducing workload and improving confidence in a 25-participant study.
-
LILO: Bayesian Optimization with Natural Language Feedback
LILO integrates LLMs to translate natural language feedback into preference signals for Gaussian process-based Bayesian optimization, outperforming standard preference BO and LLM-only methods on benchmarks.
-
Reinforcement Learning for Laser Additive Manufacturing Scan-Order Optimisation: A Bilevel Proxy--FEA Diagnostic Framework for Reward and World-Model Diagnosis
A bilevel Proxy-FEA diagnostic framework is introduced and tested on a simplified LDED32 stripe benchmark to reveal proxy misalignment with FEA labels and a stress-distortion trade-off in RL-guided scan-order optimization.
-
Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use
Trust calibration in agentic tool use is cast as preferential Bayesian optimization over a latent human risk-tolerance function observed through binary approve/deny feedback with a probit likelihood.
-
Interactive Query Answering on Knowledge Graphs with Soft Entity Constraints
The paper formalizes query answering with soft constraints on knowledge graphs and introduces two lightweight methods (parameter tuning or small neural network) to incorporate them while preserving original rankings.
-
InternLM2 Technical Report
InternLM2 is a new open-source LLM that outperforms prior versions on 30 benchmarks and long-context tasks through scaled pre-training to 32k tokens and a conditional online RLHF alignment strategy.
-
A tutorial on learning from preferences and choices with Gaussian Processes
Tutorial on a GP-based framework for preference and choice learning that unifies random utility models, limits of discernment, and multi-utility scenarios via customized likelihoods for object and label preferences.
-
A Survey on Knowledge Distillation of Large Language Models
A comprehensive survey of knowledge distillation for LLMs structured around algorithms, skill enhancement, and vertical applications, highlighting data augmentation as a key enabler.