Pith. sign in

Paper Citation Record · LEDGER

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe

As of 11 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 2 inbound Pith citation observations for arXiv:2412.19396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19396 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:36.072796Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:54.880840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T06:31:24.706841Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73a48c5f-5d88-4ccf-8d49-4b7204c7fffe · outbound

This paper cites Principal compo- nent analysis.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Principal compo- nent analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.839768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.871160Z digest=sha256:1897b81c8f63544195db8ed0d9403bec82563458a01ef258d4583b1084b2307d

Observation 930b6a8d-30f9-46fc-abed-c980215c43b9 · outbound

This paper cites Learning user interaction models for predicting web search result preferences.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Learning user interaction models for predicting web search result preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.830495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.874789Z digest=sha256:4e3e39daf199a6f285a88baff0040b31e57822e4bf09d1edce4cef3d268572c5

Observation 505cf66b-75ec-4ff0-9ba5-2e5a9976b835 · outbound

This paper cites A first-order algorithm for the a-optimal experimental design problem: a mathe- matical programming approach.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe A first-order algorithm for the a-optimal experimental design problem: a mathe- matical programming approach

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.821746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.878190Z digest=sha256:a673d65b1635da13b03678d6347a1eea514fce96ceb14e3bfdd7cda02eb863da

Observation 946624a0-79a0-4e96-83a5-7760dff213b0 · outbound

This paper cites Best arm identification in multi-armed ban- dits.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Best arm identification in multi-armed ban- dits

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.811310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.881815Z digest=sha256:1c459854449b7df4212255141731880fcd901a9fe1cd3aa97deb3bdf60eb1ff0

Observation eba99011-a564-4f1b-8118-efc839964e14 · outbound

This paper cites Fixed-budget best-arm iden- tification in structured bandits.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Fixed-budget best-arm iden- tification in structured bandits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.801691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.885352Z digest=sha256:3a59427620f6b028fe03defaa6b1f71ca5925ac12c0559a48847307792c9002c

Observation 518c7b9b-46f6-4128-a748-2c4ca279c30b · outbound

This paper cites Two-point step size gradient methods.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Two-point step size gradient methods

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.888835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.888835Z digest=sha256:229fb456364f7be3fe3250e34cd9916405fe1b9aa08bb1263f64fd4cf6c05601

Observation 242af294-ce23-4769-8b92-bcf3a2761bce · outbound

This paper cites Convex Op- timization.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Convex Op- timization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.787695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.892372Z digest=sha256:2d68714e83abd57986c06a87cd34f5e253cd46db79b8fb39772cd7b2e7e7f3a6

Observation dd716cd0-d739-42d2-8e5a-a4b51bcaaf46 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Rank analysis of incomplete block designs: I

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.779321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.895440Z digest=sha256:df5aa855352cc5ee85640884e10b176e108994b3da8896c422b87db9b6aeb414

Observation 55068812-9180-4dbe-9901-fd1d64eff91a · outbound

This paper cites Pure exploration in multi-armed bandits problems.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Pure exploration in multi-armed bandits problems

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.771043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.898781Z digest=sha256:2243897ed0448f8099b5d4f39ddd278500de22f4e65eaeabc62d5cb0ab7a4d83

Observation 453f6f5d-cbc5-4d96-a03a-fd10d120fdb0 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.901723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.901723Z digest=sha256:7784399d16819f9f39923e799120c5517e45eeac3a9e44be182c83d558592350

Observation 5b93440a-468d-44c5-8166-ecf173d41ca1 · outbound

This paper cites Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Bge m3-embedding: Multi- lingual, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.762019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.905476Z digest=sha256:0331a9289d3ff184dc4382c8466ca49e51495fe505ad126e8fb5c92396d730f3

Observation b4a2617b-30a4-45ac-877f-3102566abb44 · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Active Preference Optimization for Sample Efficient RLHF

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.909262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.909262Z digest=sha256:94360ea58e232c17fbd20651cb2e04ac34616e45cafda6dfc810bc799fd9fc1f

Observation 73b7c139-d7c9-40f5-92c9-4dab7ec5908b · outbound

This paper cites CVXPY: A Python-embedded modeling language for convex opti- mization.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe CVXPY: A Python-embedded modeling language for convex opti- mization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.753350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.912807Z digest=sha256:b6890276c6d7c1e578ca4fa9e77527c215d61be017cef455b9a28d921a4456aa

Observation 3efbfe66-296e-47ea-8d1a-f3cd32f3de94 · outbound

This paper cites Conditional gradi- ent algorithms with open loop step size rules.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Conditional gradi- ent algorithms with open loop step size rules

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.744562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.915897Z digest=sha256:b76606e93d1a309b457758c78c4b0a0ca7646707fcce190b9061b8e23650181d

Observation 896a7941-36f1-4bc1-9846-4cd874d7fe21 · outbound

This paper cites A density-based algorithm for discov- ering clusters in large spatial databases with noise.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe A density-based algorithm for discov- ering clusters in large spatial databases with noise

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.735126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.919152Z digest=sha256:e47a9be2b42a99ddd29bffbe7d84bbe7cd73db4b397f18269b2d1e2f9b463fbd

Observation 2d5d74a0-768c-4498-b597-ddbf5bf53581 · outbound

This paper cites Elsevier, 2013.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Elsevier, 2013

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.726375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.922214Z digest=sha256:947aa1921c12e8c2d0820080922359ba4b277662805bd6000843298f9d49ec17

Observation a6e6bf70-f454-47e5-9693-42513eb6a3e8 · outbound

This paper cites An algorithm for quadratic programming.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe An algorithm for quadratic programming

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.717847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.925273Z digest=sha256:7334c1efa38ae99f33cd9b28f57693eec694369e5518f7f8bb6da0f501aa5cb3

Observation 6bc78a44-46ee-4350-a2fb-79d7a598df04 · outbound

This paper cites Enlargement methods for comput- ing the inverse matrix.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Enlargement methods for comput- ing the inverse matrix

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.708804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.928458Z digest=sha256:166dd373b144d9bbf96593c1eeff08adcecc7dd4d4c2edefb9ff9f173d9c3c09

Observation 4d8c5305-11ed-4043-bbda-77be463f6f77 · outbound

This paper cites Solving the optimal experiment design problem with mixed-integer convex methods.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Solving the optimal experiment design problem with mixed-integer convex methods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.931597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.931597Z digest=sha256:a9c50ed47e06bf43ecc8b3f33c623cb7def73882dfa96bad4521e06dbcc4a2a1

Observation 40039037-f07d-4763-89e9-aee4a989c958 · outbound

This paper cites Cascading linear submodular bandits: Accounting for position bias and diversity in online learning to rank.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Cascading linear submodular bandits: Accounting for position bias and diversity in online learning to rank

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.701093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.934809Z digest=sha256:c791f13b34d2fd62e76e93d033d0ffd6a9ba2753a230e193ad75d9bdfbbe55ee

Observation 58265661-b81e-4c07-bc48-77ddc47b4e8d · outbound

This paper cites Fidelity, soundness, and efficiency of interleaved comparison methods.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Fidelity, soundness, and efficiency of interleaved comparison methods

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.693085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.937969Z digest=sha256:4f5b85d598cdcedacbb26758c6ab4243545da7e6a1ea0dc80dbfe69dce34c676

Observation 53628f8a-55d0-46c2-8da6-242d6ae83d4d · outbound

This paper cites On- line evaluation for information retrieval.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe On- line evaluation for information retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.685233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.941071Z digest=sha256:185d38d781fa03c1e45dff8a6335eecdae220111f7ff424275dfa5c734410f6f

Observation db83ee6a-2423-4227-b040-5ed8893f223d · outbound

This paper cites Revisiting frank-wolfe: Projection-free sparse convex optimization.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Revisiting frank-wolfe: Projection-free sparse convex optimization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.676563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.944014Z digest=sha256:af80483087cf04a6f47338c04f0c88c849b6589f59fb10bb3e1a80b594d4e6d0

Observation 885bae25-db46-4367-8cd3-36e84751bed5 · outbound

This paper cites Is pes- simism provably efficient for offline rl? In Interna- tional Conference on Machine Learning, pages 5084–.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Is pes- simism provably efficient for offline rl? In Interna- tional Conference on Machine Learning, pages 5084–

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.667230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.946677Z digest=sha256:308f31832c2a1188735e4786c60ad8055cc5d062260468dc6aad45370865079a

Observation b77ac3fd-0a00-4d17-acf3-b3352b88eeda · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Beyond Reward: Offline Preference-guided Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.949355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.949355Z digest=sha256:86478ef51fde9db2c8152c1d749e41283f5637ff8391ed0299d155e60439404e

Observation 0a84b77a-833b-4536-8e6a-3b17e1b965f2 · outbound

This paper cites Rank correlation methods.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Rank correlation methods

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.658085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.952389Z digest=sha256:fa4381bb0e086a1f19c189c6243be27de00c182061e7691df9a5269fe731490f

Observation b0321da2-070e-4837-bbef-b6d75704ba14 · outbound

This paper cites Frank-wolfe with subsampling oracle.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Frank-wolfe with subsampling oracle

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.648716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.955165Z digest=sha256:27edcad3755c276ef27a5b21aa88c2df6f3992c03d26997540e8f853b42ee2ca

Observation 75f78b15-1355-4537-b17d-bd3ecead8dc0 · outbound

This paper cites Rounding of polytopes in the real number model of computation.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Rounding of polytopes in the real number model of computation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.639459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.958030Z digest=sha256:a1931b764da61e8c2c3e8b4fe069ae8d3e04b4983a89cb937e4f733f669478c5

Observation a95202aa-7b24-47f2-bd62-7c4f724bf185 · outbound

This paper cites Data-driven rank breaking for efficient rank aggregation.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Data-driven rank breaking for efficient rank aggregation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.630624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.960696Z digest=sha256:5863f5a5d88f7dcff76a9b771446146f47c78b2856ce51bd8652eab552e842a5

Observation c228a665-3c0c-4161-be11-17d7c8226ccb · outbound

This paper cites Sequential minimax search for a max- imum.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Sequential minimax search for a max- imum

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.621105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.963483Z digest=sha256:3c915ee0c6dc1bca5262dc8bea7a5133aeb2e3afb31561c40344b75c01caa9b1

Observation c8c58908-8490-4ebe-b6cd-f3ddc1ca5a6c · outbound

This paper cites The equivalence of two extremum problems.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe The equivalence of two extremum problems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.612166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.966102Z digest=sha256:f704a4c7a3c0e9a17f73c3ec67165ec05eb5464e953efca87808746d8887dd5a

Observation 81d7585d-2499-49d0-b9ce-dc41842cd85f · outbound

This paper cites Cascading bandits: Learning to rank in the cascade model.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Cascading bandits: Learning to rank in the cascade model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.603867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.968972Z digest=sha256:fae0366a99823ee850bfe5e977b719092b741e2c2ea59e80aba76a2399e0c175

Observation d7256de0-abf7-4d62-a53d-ae2f66a2204c · outbound

This paper cites Multiple-play bandits in the position-based model.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Multiple-play bandits in the position-based model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.594973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.972027Z digest=sha256:086cf41a12d1dfa82a4ea90b3eac0992763e7b3414a4f48b923740d1763d72f4

Observation 0f594595-3ccb-47fa-b77f-3bb9416bb218 · outbound

This paper cites Bandit Algo- rithms.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Bandit Algo- rithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.586252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.975123Z digest=sha256:5b94135b2625e915b12732f6770976fb1742a2ba6e7134d0444397d5052350df

Observation 0005eb18-6d47-425f-9e3a-8c77aa53033d · outbound

This paper cites Contextual combinatorial cascading bandits.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Contextual combinatorial cascading bandits

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.578251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.978154Z digest=sha256:cdd2e354a4bd4135db5eef1be98bd580351a562bd53d143e864fcec55885b029

Observation 35f8f7c1-8e37-4012-94bf-f12d5cfd6414 · outbound

This paper cites Individual Choice Behavior: A Theoretical Analysis.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Individual Choice Behavior: A Theoretical Analysis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.569569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.981611Z digest=sha256:e19c636d73f93034f594e454027ae300a67a9eece994e1eb80c68c1b5569c948

Observation 5aba3326-f61b-435e-bedb-ce745c656172 · outbound

This paper cites Introduction to Information Retrieval.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Introduction to Information Retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.560738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.984719Z digest=sha256:8abd99d17d132f552b76259f17a712c486592153ec486cb4dda8f3c52b9f14f9

Observation 6305a891-33ef-45e0-ac8b-b36061d7c992 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.987938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.987938Z digest=sha256:9d9043ef35389166964ae03a918267725071328e65684e5cb36d3d6b57dcd850

Observation 629ce119-4754-4d82-91a5-90639aa00243 · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.991585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.991585Z digest=sha256:96bdc107ad1328f6ecfe2ba618c5b64efd4a8a74fa7bdd02b3505ffd3b29efe0

Observation daaef99a-2146-452a-9261-2c47fc04d65b · outbound

This paper cites Optimal design for human feed- back.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Optimal design for human feed- back

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-11T00:45:36.257795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.996239Z digest=sha256:5cb0f3a61200c9f751c8826278828ebed7440d6d2ae80588f9f216ab79095447

Observation dbe72c99-cf2e-4cbb-b308-0cf7d3b4a84e · outbound

This paper cites Learning from comparisons and choices.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Learning from comparisons and choices

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.551469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:35.999518Z digest=sha256:ed8e636e453a79a0d424f71299c3caa87cc16971dce48423de4e75b684ef6422

Observation db94ce64-9903-4d6f-ad11-3d5cd5db18d2 · outbound

This paper cites Introductory Lectures on Convex Opti- mization: A Basic Course.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Introductory Lectures on Convex Opti- mization: A Basic Course

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.541349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.003046Z digest=sha256:71bcee3a9151ec13a3bb6f55e7580dfc6b738218a64528a0fa81389658ff8936

Observation 22d89ca2-4e61-4eed-a4a5-0e9e029d52b9 · outbound

This paper cites Training lan- guage models to follow instructions with human feed- back.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Training lan- guage models to follow instructions with human feed- back

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.532595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.006440Z digest=sha256:26aedbd1154b5f13b682f73ea40187ac9ca0d72e67f93634f0b1261cdef2469d

Observation 8013951a-7260-4d49-a43e-a7e756199bb2 · outbound

This paper cites Fast and robust rank aggregation against model misspecification.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Fast and robust rank aggregation against model misspecification

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.523272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.010012Z digest=sha256:4f449bf6303b7982daba788d842cfc66bb21c785c9986678fd306116180128e8

Observation ba98eb98-a162-47a8-8bab-5931f97520b2 · outbound

This paper cites The analysis of permutations.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe The analysis of permutations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.513734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.013208Z digest=sha256:c96d41e150a51f36413b2f209761f9a84805646daee5536f7857e4313a3ba20a

Observation f1e55345-883c-446e-baf1-ffe34d2f3407 · outbound

This paper cites An introduction to grids, graphs, and networks.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe An introduction to grids, graphs, and networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.504296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.016453Z digest=sha256:995c1d7ecb577fba1dca0aa92bfe6eb2c8ea692a8ea036adb3f96d8087140312

Observation b744c86b-d170-4cbb-a89b-c276b9940195 · outbound

This paper cites Optimal Design of Experiments.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Optimal Design of Experiments

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.495480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.019718Z digest=sha256:3db39ab2a1a52240d32c2f293680cabce613e83928589acc9d07d1be0b52bfda

Observation 3a33d967-b7f3-4f51-b3d1-e8936b6bbef5 · outbound

This paper cites Learning diverse rankings with multi-armed bandits.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Learning diverse rankings with multi-armed bandits

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.486650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.023153Z digest=sha256:f29e08b90db1eb4773f9f3d94bf66b7ab577c6d58e5369a3d3e2656c544c3d91

Observation 3a91af31-d1ec-4f5e-8266-73876c6eb0d1 · outbound

This paper cites Di- rect preference optimization: Your language model is secretly a reward model.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Di- rect preference optimization: Your language model is secretly a reward model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.477419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.026091Z digest=sha256:a0e810845f86912856ba71df9c5557482b983f1215faebf1f49768cfb50aa919

Observation 23f54b48-01a5-4875-a344-503d550d5ac9 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:36.029134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:36.029134Z digest=sha256:5486c6ba6cd1ca7b2bbe798ac01926c3eeb5b45db035cb03914774e47a555182

Observation 10f40224-b8a0-45b0-b975-7f726413eceb · outbound

This paper cites Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:36.032487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:36.032487Z digest=sha256:06dce29177501c5b2e6d4e956646a88dd788442b116b03b24ffe42d81e81ce80

Observation a9e2cdee-d4a7-4423-a6b1-ddd1378e5cb0 · outbound

This paper cites Embedding-Aligned Language Models.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Embedding-Aligned Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:45:36.132016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.035816Z digest=sha256:6480c86336beb39a6a937bbcb56a4099562e80e8fb4a3881b1547bc04636a7b2

Observation 34a7ec85-c897-437b-8ca4-083a0534283a · outbound

This paper cites Judgment under uncertainty: Heuristics and biases.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Judgment under uncertainty: Heuristics and biases

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.467848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.039148Z digest=sha256:a701720a43ece9a42f83811c10e5a5df2e9e8529522c3f3aedf73d6bc82f8654

Observation 4c291af1-e943-46da-bc9a-a2109ad1261d · outbound

This paper cites Inverting modified matrices.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Inverting modified matrices

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.458538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.041960Z digest=sha256:444ecba78330556323f4a353f45cba65cc1025b7fc714e2223809cacf2514216

Observation b87e07d4-3892-4a15-8bde-c8022158949d · outbound

This paper cites Provably Efficient Offline Reinforcement Learning with Trajectory-Wise Reward.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Provably Efficient Offline Reinforcement Learning with Trajectory-Wise Reward

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:45:36.118267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.044547Z digest=sha256:a93287466515fe697249e48e6561d1d44295eaade9eb101086885cf51ac28280

Observation f533508f-d760-4962-bf9e-f85264d302e9 · outbound

This paper cites Minimax optimal fixed- budget best arm identification in linear bandits.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Minimax optimal fixed- budget best arm identification in linear bandits

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.449241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.047216Z digest=sha256:238c5664551aed8d078dc1344af982d3bd41893e0ea37c582b1a03cc90ce1f6a

Observation 3e4d5ad2-c588-4d51-b592-88b113e50751 · outbound

This paper cites When is realizability sufficient for off- policy reinforcement learning? In International Con- ference on Machine Learning , pages 40637–40668.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe When is realizability sufficient for off- policy reinforcement learning? In International Con- ference on Machine Learning , pages 40637–40668

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.440991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.049721Z digest=sha256:06f27df1a648824e8d735b16135af57d0ab43933562ffc6a7c0e97fdefbf4088

Observation c7597923-dc27-4358-8e31-ba70322dc6f7 · outbound

This paper cites Analysis of the frank–wolfe method for convex composite optimiza- tion involving a logarithmically-homogeneous barrier.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Analysis of the frank–wolfe method for convex composite optimiza- tion involving a logarithmically-homogeneous barrier

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.432037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.052288Z digest=sha256:46427fc93b2c21241bb4e08b2325c781fa3314fea51262c4d19a6d6bd1003a39

Observation 46ab2f0c-9e43-4587-8c00-7ad648c2f25c · outbound

This paper cites Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:36.054778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:36.054778Z digest=sha256:132ea3fbace408efc7b5623aa3a8bfa6d821beeab68d2b710deced84e84a7fe7

Observation 77556b48-5f7d-4d2f-8e33-327a942b51a0 · outbound

This paper cites Cascading bandits for large-scale recommendation problems.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Cascading bandits for large-scale recommendation problems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.423698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.057600Z digest=sha256:8699451ae7eec40727e7c57c6f5e50f412da439e3c7f82cbc86d1705691d6872

Observation 74bcd6d3-60db-4ad6-bfd5-2653dbf767d8 · outbound

This paper cites None of these works focus on ranking problems.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe None of these works focus on ranking problems

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.414579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.060012Z digest=sha256:7bd99d1d1b13bba8ae8736c336cfe3c5cea40b72a406cacb7c08128706398554

Observation d88e2feb-cc1a-4ee4-a7a5-71b5025ff525 · outbound

This paper cites an unresolved cited work.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:45:36.405534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.063358Z digest=sha256:31e3c1baa05a80ad6911276cf8231aa291bb90fc9d323b255d44fa8f0526049e

Observation 760e450b-aa44-45b0-8caf-129e035a4a91 · outbound

This paper cites f (tu) = f (u) − θ log t, for all u ∈ int(K) and t >0 for some θ ≥ 1,.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe f (tu) = f (u) − θ log t, for all u ∈ int(K) and t >0 for some θ ≥ 1,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:36.396582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.066381Z digest=sha256:417d2adc3a56aa26099f321ef6313c671d13bb60f81590c5bc1ef29afb96e3d7

Observation f64f5873-fd28-43ca-9ae7-51804e038a6b · outbound

This paper cites an unresolved cited work.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:45:36.387401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.069642Z digest=sha256:2db8bb7908b850c757a81ed86ee5dd1029097326e7d19dbaffe9224583953ebb

Observation 138dc4f5-a130-4bff-aeed-8b2d6fce9908 · outbound

This paper cites descent lemma.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe descent lemma

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:45:36.378115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:45:36.072796Z digest=sha256:39572e829447756df2c9771ab18c6780269f261d503cf7c7d366df060ea86dd6

Pith citing papers

Observation 64413cef-0d91-4807-9ca8-f8e52ee96539 · inbound

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain cites this paper.

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:54.880840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:54.880840Z digest=sha256:52e71b61f3410aeb7bf17dc30dd641767dcc3fc9320caf5617235707f95616b8

Observation 3dc6952f-33d2-4a0c-8a49-5617d87ef2a4 · inbound

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization cites this paper.

MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:24.712722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:14:37.374346Z digest=sha256:286e0d4a66541d1a406cdaa5299a8ff49e5ecb6bdd071e7eb19689ef53233b2a