REVIEW 2 major objections 5 minor 52 references
FAIRY: A Framework for Understanding Relationships between Users' Actions and their Social Feeds
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Paths through a user's own actions can explain why feed items appear
desk verdict A genuinely new framing for feed explainability, but the evaluation's train/test split likely leaks and the data aren't released, so the accuracy claims need a clean re-run before they can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the interaction graph, a heterogeneous information network built only from information a user can see: nodes represent users, content items, and categories; edges are directed, weighted, timestamped actions such as follows, asks, upvotes, scrobbles, loves, and belongs-to relations, with inverse edges added so traversal can go both ways. An explanation path is defined as any path from the user to a feed item for which every edge timestamp is strictly earlier than the time the item was seen, guaranteeing that the explanation only cites actions already on record; when the graph is not connected, the platform's topic taxonomy is overlaid to route paths through category ancestors. The ranking engine is a pairwise learning-to-rank model trained with ordinal regression on user preference judgments, using five feature groups—user, category, item, path instance, and path pattern—chosen deliberately so that, in principle, a user could look at a path and see why it was ranked. This machinery converts a massive, hidden feed-generation process into a small set of readable human-scale explanations.
What would settle it
On a platform with an open recommender or logged ranking reasons, collect the top-ranked FAIRY paths for a sample of feed items and compare them with the logged reasons; if the majority of top paths do not match or subsume those reasons, or if randomly re-timestamped paths perform equally well in user preference tests, the central claim is refuted.
Extended reading notes
Core claim
The paper's central discovery is that explanation paths—sequences of timestamped edges connecting the user to the feed item—are a valid operationalization of why a feed item appeared, and that user judgments of relevance and surprisal on these paths are learnable from lightweight, user-visible features such as user influence, category specificity, item engagement, path length and recency, pattern frequency, and edge-type counts. The paper demonstrates this by collecting thousands of pairwise preference judgments from real users and training ordinal-regression ranking models that predict the preferred explanation. On one platform the model reaches 60.33% accuracy for relevance and 60.38% for surprisal; on the other it reaches 56.24% and 54.21%, each time statistically above the strongest of three baseline relationship-discovery methods. The paper also finds that path-pattern features carry much of the signal on one platform, user-specific models often beat one global model, and users are consistent in their judgments for 74–80% of transitive triplets.
Load-bearing premise
The load-bearing premise is that a feed item's appearance is always traceable to paths in the user's own visible interaction graph, so if the platform uses hidden signals, global trends, or actions outside the user's view, FAIRY's explanations are plausible-sounding but not true reasons.
Editorial extensions
If this is right
- A user-side tool can surface the top few explanation paths for any feed item, letting users audit their feeds without any cooperation from the platform.
- Because the features are user-visible and interpretable, a user can trace why one path was ranked above another and can see which of her own actions contributed.
- The same interaction-graph construction and ranking pipeline transfers to other platforms whose data can be laid out as nodes and timestamped action edges, not just the two studied here.
- Pairwise preference judgments from users are a workable gold standard for explanation quality: users were consistent in 80% of transitive triplets on one platform and 74% on the other.
- Embedding-based similarity between feed items and path nodes improves ranking over taxonomy-distance alone, so content signals belong in explanation ranking.
Reading between the lines
- If platforms ever open their ranking logs, FAIRY-style paths could serve as a user-side audit that checks whether visible explanations match the real algorithmic reasons; mismatches would flag hidden personalization.
- The surprisal ranking could double as a filter-bubble detector: paths that reveal unexpected connections between the user and incoming content expose associations the user did not consciously make.
- A natural extension is to turn the ranking model around: instead of explaining a given item, find the user action whose removal most changes the set of top explanation paths, giving a prescriptive 'do this to see less of X' recommendation.
- The pairwise preference protocol itself could be reused to collect explanation judgments at scale via crowd platforms, since users found the path-pair task cognitively manageable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FAIRY, a user-side framework that constructs a heterogeneous interaction graph from a user's visible platform actions and treats paths from the user to a feed item as candidate explanations. It defines timestamp-constrained explanation paths (Definition 2.2), represents them with simple interpretable features (Section 3.1), and ranks them with SVMRank trained on pairwise user judgments of relevance and surprisal. User studies on Quora and Last.fm with 20 paid users per platform yield pairwise ranking accuracies of about 60% and 54-56% (Table 2), which are claimed to significantly outperform adapted ESPRESSO, REX, and PRA baselines. The paper also reports user-specific models, ablations, perturbation analyses, transitivity checks, and anecdotal examples.
Significance. If the reported accuracies are trustworthy, FAIRY is a useful step toward user-side feed transparency: the idea of grounding explanations in a user-visible interaction graph is well motivated, the features are deliberately interpretable, and the two-platform user study is a genuine attempt to collect gold judgments. I credit the authors for releasing code, for the transitivity sanity check, and for the perturbation analysis that probes what drives judgments. The main reservation is that the central empirical claim currently rests on a fragile evaluation protocol; without a query-disjoint split, the significance of the numbers in Table 2 cannot be assessed.
major comments (2)
- [Section 5 (Datasets and LTR)] The train/test split is described only as '80% training, 10% development, and 10% test sets' (Section 5), without specifying the split unit. This is consequential because the judgment data are heavily overlapping: 11,677 Quora pairs are drawn from only 459 distinct (u,f) queries, and each explanation path appears in about 1.9 pairs on average; the corresponding Last.fm numbers are 4,791 pairs, 235 queries, and 1.7 pairs per path. A random pair-level split therefore places the same path and the same (u,f) query in both training and test, and since the LTR features are properties of a path in the context of a specific feed item, the test instances are near-duplicates of training instances. The unsupervised baselines in Table 2 do not train on labels and so do not benefit from this leakage, which can inflate FAIRY's reported 5-10 point gains. Please re-run the evaluation with grouped splits (e.g., leave-one-query-out, leave-one-user-out, or a split with no shared paths between train and test) and report whether the accuracies in Table 2 survive; the current p-values from a paired t-test do not address this form of non-independence.
- [Section 4 (User Studies) and Section 6.2 (Perturbation analysis)] The evaluation measures only pairwise preference accuracy on randomly sampled path pairs, not the quality of the final top-k explanation lists a user would see. The perturbation results in Table 4 show that when pairs differ only in the feed item, accuracy falls to about 50-52% on both platforms, which suggests that the model is strongly tied to the specific path-pair distribution used for training and may not transfer to the realistic setting of ranking all candidate paths for a new feed item. Please add an evaluation of full ranking quality (e.g., nDCG or precision at k on the complete candidate path set per (u,f) query) or otherwise justify that pairwise accuracy on random pairs supports the claimed usefulness of the ranked list.
minor comments (5)
- [Abstract and Section 1] The abstract and introduction say FAIRY 'explains' why items were shown, but the framework computes plausible user-side paths and the evaluation measures user preference, not agreement with the platform's actual feed algorithm. Please consistently use 'plausible explanation' or 'user-side explanation' to avoid overclaiming, as the Introduction itself acknowledges these relationships are 'a proxy that the users could find plausible'.
- [Section 6.2 (Ablation study)] The text states that removing path pattern features 'does not affect accuracies of models on Last.fm', but Table 3 shows a drop from 56.24 to 54.32 for relevance and from 54.21 to 53.65 for surprisal; please adjust the wording to say that the effect is smaller than on Quora rather than absent.
- [Table 5] Several path examples contain text-encoding artifacts (e.g., 'Cookinд', 'Enдekbert', 'sin дs') that obscure the examples; please ensure the final PDF uses correct glyphs.
- [Figures 3-6] The user-specific model accuracies are reported without error bars or confidence intervals; with only 20 users per platform, please state how many judgments per user were used and whether the observed differences are stable across split seeds.
- [Table 2 (statistical test)] The paired t-test should state the pairing unit (pair of paths, query, or user) and whether it accounts for the non-independence of pairs sharing paths or queries; as written, the significance claim is not verifiable from the reported information.
Circularity Check
No significant circularity: FAIRY's ranking model is supervised prediction from graph features to external user judgments, not a self-derived quantity.
full rationale
FAIRY's claimed derivation chain is empirical rather than definitional. The paper defines candidate explanation paths as temporal-constrained paths in a user-specific interaction graph, defines interpretable graph-structural features, collects pairwise user judgments of relevance and surprisal, trains a pairwise learning-to-rank model on those judgments, and evaluates on held-out pairs against unsupervised baselines. Relevance and surprisal are not defined as functions of the features; they are external user assessments, and the features are computed from graph structure, edge types, timestamps, and embeddings, not from the judgment labels. The only self-citation of note, [22] for the taxonomy-overlay connectivity strategy, is a construction detail and not load-bearing for the empirical claim. ESPRESSO [36] is used as a baseline, not as authority for FAIRY's design. The paper does not invoke a uniqueness theorem or prior author result to force its modeling choice. The possible absence of a query-disjoint or path-disjoint train/test split is an evaluation-validity concern, but it is not a reduction of the prediction to its inputs by construction and therefore does not constitute circularity under the stated criteria.
Assumptions & free parameters
free parameters (4)
- LTR model weights (SVMRank linear kernel) =
Learned from user judgments; not reported
- Sentence embedding dimension and Arora et al. model hyperparameters =
Not specified
- Path length cutoffs =
4 for Quora, 5 for Last.fm
- Number of path pairs sampled per feed item =
About 25
assumptions (4)
- domain assumption User-visible information is sufficient to construct an interaction graph that contains meaningful explanation paths.
- domain assumption Taxonomy overlay ensures connectivity between the user and the feed item.
- domain assumption Timestamp ordering, with all edges before tau(f), is a valid explanatory constraint.
- domain assumption Paid fresh-account users' judgments are representative of real feed recipients.
invented entities (1)
-
Explanation path
Cite this review
Pith. "Pith review of FAIRY: A Framework for Understanding Relationships between Users' Actions and their Social Feeds." pith.science (2026). https://pith.science/paper/CKF5Y6ZO
@misc{pith2026190803109,
author = {Pith},
title = {Pith review of: FAIRY: A Framework for Understanding Relationships between Users' Actions and their Social Feeds},
year = {2026},
howpublished = {\url{https://pith.science/paper/CKF5Y6ZO}},
note = {Machine review of arXiv:1908.03109}
}
read the original abstract
Users increasingly rely on social media feeds for consuming daily information. The items in a feed, such as news, questions, songs, etc., usually result from the complex interplay of a user's social contacts, her interests and her actions on the platform. The relationship of the user's own behavior and the received feed is often puzzling, and many users would like to have a clear explanation on why certain items were shown to them. Transparency and explainability are key concerns in the modern world of cognitive overload, filter bubbles, user tracking, and privacy risks. This paper presents FAIRY, a framework that systematically discovers, ranks, and explains relationships between users' actions and items in their social media feeds. We model the user's local neighborhood on the platform as an interaction graph, a form of heterogeneous information network constructed solely from information that is easily accessible to the concerned user. We posit that paths in this interaction graph connecting the user and her feed items can act as pertinent explanations for the user. These paths are scored with a learning-to-rank model that captures relevance and surprisal. User studies on two social platforms demonstrate the practical viability and user benefits of the FAIRY method.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Deepak Agarwal, Bee-Chung Chen, Qi He, Zhenhao Hua, Guy Lebanon, Yiming Ma, Pannagadatta Shivaswamy, Hsiao-Ping Tseng, Jaewon Yang, and Liang Zhang
-
[2]
Gummadi, Patrick Loiseau, and Alan Mislove
Athanasios Andreou, Giridhari Venkatadri, Oana Goga, Krishna P. Gummadi, Patrick Loiseau, and Alan Mislove. 2018. Investigating ad transparency mecha- nisms in social media: A case study of Facebook’s explanations. In NDSS
work page 2018
-
[3]
Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2016. A simple but tough-to-beat baseline for sentence embeddings. In ICLR
work page 2016
-
[4]
Hannah Bast and Elmar Haussmann. 2015. More accurate question answering on Freebase. In CIKM
work page 2015
-
[5]
Freya Behrens, Sebastian Bischoff, Pius Ladenburger, Julius Rückin, Laurenz Seidel, Fabian Stolp, Michael Vaichenker, Adrian Ziegler, Davide Mottin, Fatemeh Aghaei, et al. 2018. MetaExp: Interactive Explanation and Exploration of Large Knowledge Graphs. In WWW
work page 2018
-
[6]
Federico Bianchi, Matteo Palmonari, Marco Cremaschi, and Elisabetta Fersini
-
[7]
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to rank: From pairwise approach to listwise approach. In ICML
work page 2007
-
[8]
Ben Carterette, Paul N Bennett, David Maxwell Chickering, and Susan T Dumais
Show all 52 references
-
[9]
Kelley Cotter, Janghee Cho, and Emilee Rader. 2017. Explaining the news feed algorithm: An analysis of the News Feed FYI blog. In CHI
2017
-
[10]
Hongbo Deng, Jiawei Han, Bo Zhao, Yintao Yu, and Cindy Xide Lin. 2011. Prob- abilistic topic models with biased propagation on heterogeneous information networks. In KDD
2011
-
[11]
Yuxiao Dong, Nitesh V Chawla, and Ananthram Swami. 2017. metapath2vec: Scalable representation learning for heterogeneous networks. In KDD
2017
-
[12]
Chad Edwards, Autumn Edwards, Patric R Spence, and Ashleigh K Shelton
-
[13]
I always assumed that I wasn’t really that close to [her]
Motahhare Eslami, Aimee Rickman, Kristen Vaccaro, Amirhossein Aleyasen, Andy Vuong, Karrie Karahalios, Kevin Hamilton, and Christian Sandvig. 2015. “I always assumed that I wasn’t really that close to [her]”: Reasoning about Invisible Algorithms in News Feeds. In CHI
2015
-
[14]
Lujun Fang, Anish Das Sarma, Cong Yu, and Philip Bohannon. 2011. REX: Explaining relationships between entity pairs. In VLDB
2011
-
[15]
Jill Freyne, Shlomo Berkovsky, Elizabeth M Daly, and Werner Geyer. 2010. Social networking feeds: recommending items of interest. In RecSys. Quora: Relevance 1 Shrey f oll ows −−−−−−−→Ali f oll ows −−−−−−−→Social Psychology belongs to−1 −−−−−−−−−→What are the things you should...
2010
-
[16]
Yuri Gurevich, Efim Hudis, and Jeannette M Wing. 2016. Inverse privacy. CACM (2016)
2016
-
[17]
Kevin Hamilton, Karrie Karahalios, Christian Sandvig, and Motahhare Eslami
-
[18]
Liangjie Hong, Ron Bekkerman, Joseph Adler, and Brian D Davison. 2012. Learn- ing to rank social update streams. In SIGIR
2012
-
[19]
Binbin Hu, Chuan Shi, Wayne Xin Zhao, and Philip S Yu. 2018. Leveraging Meta- path based Context for Top-N Recommendation with A Neural Co-Attention Model. In KDD
2018
-
[20]
Robert Jäschke, Leandro Marinho, Andreas Hotho, Lars Schmidt-Thieme, and Gerd Stumme. 2007. Tag recommendations in folksonomies. In PKDD
2007
-
[21]
Thorsten Joachims. 2006. Training linear SVMs in linear time. In KDD
2006
-
[22]
A path to understanding the effects of algorithm awareness. In CHI
-
[23]
Xiangnan Kong, Bokai Cao, and Philip S Yu. 2013. Multi-label classification by mining label and instance correlations from heterogeneous information networks. In KDD
2013
-
[24]
Ni Lao and William W Cohen. 2010. Relational retrieval using a combination of path-constrained random walks. Machine learning (2010)
2010
-
[25]
Ni Lao, Tom Mitchell, and William W Cohen. 2011. Random walk inference and learning in a large scale knowledge base. In EMNLP
2011
-
[26]
Mathias Lécuyer, Guillaume Ducoffe, Francis Lan, Andrei Papancea, Theofilos Petsios, Riley Spahn, Augustin Chaintreau, and Roxana Geambasu. 2014. XRay: Enhancing the Web’s Transparency with Differential Correlation. In USENIX Security Symposium
2014
-
[27]
Gjergji Kasneci, Maya Ramanath, Mauro Sozio, Fabian M Suchanek, and Gerhard Weikum. 2009. STAR: Steiner-tree approximation in relationship graphs. InICDE
2009
-
[28]
Sangkeun Lee, Sungchan Park, Minsuk Kahng, and Sang-Goo Lee. 2013. PathRank: Ranking nodes on a heterogeneous graph for flexible hybrid recommender sys- tems. Expert Systems with Applications (2013)
2013
-
[29]
Jiongqian Liang, Deepak Ajwani, Patrick K Nicholson, Alessandra Sala, and Srinivasan Parthasarathy. 2016. What links Alice and Bob?: Matching and ranking semantic patterns in heterogeneous networks. In WWW
2016
-
[30]
Xiaozhong Liu, Yingying Yu, Chun Guo, and Yizhou Sun. 2014. Meta-path-based ranking with pseudo relevance feedback on heterogeneous graph for citation recommendation. In CIKM
2014
-
[31]
Javier Parra-Arnau, Jagdish Prasad Achara, and Claude Castelluccia. 2017. MyAd- Choices: Bringing transparency and control to online advertising. TWeb (2017)
2017
-
[32]
Mathias Lecuyer, Riley Spahn, Yannis Spiliopolous, Augustin Chaintreau, Roxana Geambasu, and Daniel Hsu. 2015. Sunlight: Fine-grained targeting detection at scale with statistical confidence. In SIGSAC
2015
-
[33]
Filip Radlinski and Thorsten Joachims. 2005. Query chains: Learning to rank from implicit feedback. In KDD
2005
-
[34]
Cartic Ramakrishnan, William H Milnor, Matthew Perry, and Amit P Sheth. 2005. Discovering informative connection subgraphs in multi-relational graphs. ACM SIGKDD Explorations Newsletter (2005)
2005
-
[35]
Michael Schuhmacher, Laura Dietz, and Simone Paolo Ponzetto. 2015. Ranking entities for Web queries through text and knowledge. In CIKM. ACM
2015
-
[36]
Bedathur, Sarath Kumar Kondreddi, Patrick Ernst, and Gerhard Weikum
Stephan Seufert, Klaus Berberich, Srikanta J. Bedathur, Sarath Kumar Kondreddi, Patrick Ernst, and Gerhard Weikum. 2016. ESPRESSO: Explaining Relationships between Entity Sets. In CIKM
2016
-
[37]
Giuseppe Pirrò. 2015. Explaining and suggesting relatedness in knowledge graphs. In ISWC
2015
-
[38]
Ping-Han Soh, Yu-Chieh Lin, and Ming-Syan Chen. 2013. Recommendation for online social feeds by exploiting user response behavior. In WWW
2013
-
[39]
PK Srijith, Michal Lukasik, Kalina Bontcheva, and Trevor Cohn. 2017. Longitudi- nal modeling of social media with Hawkes process based on users and networks. In ASONAM
2017
-
[40]
Yizhou Sun and Jiawei Han. 2013. Mining heterogeneous information networks: A structural analysis approach. ACM SIGKDD Explorations Newsletter (2013)
2013
-
[41]
Yizhou Sun, Jiawei Han, Charu C Aggarwal, and Nitesh V Chawla. 2012. When will it happen?: relationship prediction in heterogeneous information networks. In WSDM
2012
-
[42]
Dominic Seyler, Praveen Chandar, and Matthew Davis. 2018. An Information Retrieval Framework for Contextual Suggestion Based on Heterogeneous Infor- mation Network Embeddings. In SIGIR
2018
-
[43]
Gang Wang, Konark Gill, Manish Mohanlal, Haitao Zheng, and Ben Y Zhao. 2013. Wisdom in the social crowd: An analysis of Quora. In WWW
2013
-
[44]
Xiao Yu, Xiang Ren, Yizhou Sun, Quanquan Gu, Bradley Sturt, Urvashi Khandel- wal, Brandon Norick, and Jiawei Han. 2014. Personalized entity recommendation: A heterogeneous information network approach. In WSDM
2014
-
[45]
Chuxu Zhang, Chao Huang, Lu Yu, Xiangliang Zhang, and Nitesh V Chawla
-
[46]
Jiawei Zhang, Philip S Yu, and Zhi-Hua Zhou. 2014. Meta-path based multi- network collective link prediction. In KDD
2014
-
[47]
Yizhou Sun, Jiawei Han, Xifeng Yan, Philip S Yu, and Tianyi Wu. 2011. Pathsim: Meta path-based top-k similarity search in heterogeneous information networks. In VLDB
2011
-
[2008]
Here or there: Here or There: Preference Judgments for Relevance. InECIR
-
[2014]
Computers in Human Behavior (2014)
Is that a bot running the social media feed? Testing the differences in perceptions of communication quality for a human agent and a bot agent on Twitter. Computers in Human Behavior (2014)
2014
-
[2015]
Personalizing linkedin feed. In KDD
-
[2017]
Actively learning to rank semantic associations for personalized contextual exploration of knowledge graphs. In ESWC
-
[2018]
Camel: Content-Aware and Meta-path Augmented Metric Learning for Author Identification. In WWW
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.