{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:KVMXKZIACYBGHPMA2GWYU66VDF","short_pith_number":"pith:KVMXKZIA","schema_version":"1.0","canonical_sha256":"5559756500160263bd80d1ad8a7bd51963bc56d71f63f00aac9f2d10378ab105","source":{"kind":"arxiv","id":"2310.17567","version":1},"attestation_state":"computed","paper":{"title":"Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG","cs.NE"],"primary_cat":"cs.CL","authors_text":"Anirudh Goyal, Arushi Gupta, Dingli Yu, Jonah Brown-Cohen, Sanjeev Arora, Simran Kaur","submitted_at":"2023-10-26T16:55:05Z","abstract_excerpt":"With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023).\n  This work introduces Skill-Mix, a new evaluation to measure ability to combine skills. Using a list of $N$ skills the evaluator repeatedly picks random subsets of $k$ skills and asks the LLM to produce te"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.17567","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2023-10-26T16:55:05Z","cross_cats_sorted":["cs.AI","cs.LG","cs.NE"],"title_canon_sha256":"09f1f8821838538d0155ede5e1be926f9b8c38f2b4846a1f4bbfc5a2f1c313aa","abstract_canon_sha256":"d54a1fb6615df161f60ad2337999022294ba326a3a4dfe80d933b8e0c81e0af0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:05:36.248974Z","signature_b64":"yT6/2sjFS5yLMsREBY+41jfOe4xvcuS0NIVvnPr7znINDtAYRVic/XOKciWmhHxK7JqT+6g0xcHOmtYEwjkcBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5559756500160263bd80d1ad8a7bd51963bc56d71f63f00aac9f2d10378ab105","last_reissued_at":"2026-07-05T07:05:36.248440Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:05:36.248440Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG","cs.NE"],"primary_cat":"cs.CL","authors_text":"Anirudh Goyal, Arushi Gupta, Dingli Yu, Jonah Brown-Cohen, Sanjeev Arora, Simran Kaur","submitted_at":"2023-10-26T16:55:05Z","abstract_excerpt":"With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023).\n  This work introduces Skill-Mix, a new evaluation to measure ability to combine skills. Using a list of $N$ skills the evaluator repeatedly picks random subsets of $k$ skills and asks the LLM to produce te"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.17567","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.17567/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.17567","created_at":"2026-07-05T07:05:36.248507+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.17567v1","created_at":"2026-07-05T07:05:36.248507+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.17567","created_at":"2026-07-05T07:05:36.248507+00:00"},{"alias_kind":"pith_short_12","alias_value":"KVMXKZIACYBG","created_at":"2026-07-05T07:05:36.248507+00:00"},{"alias_kind":"pith_short_16","alias_value":"KVMXKZIACYBGHPMA","created_at":"2026-07-05T07:05:36.248507+00:00"},{"alias_kind":"pith_short_8","alias_value":"KVMXKZIA","created_at":"2026-07-05T07:05:36.248507+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.18089","citing_title":"From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning","ref_index":169,"is_internal_anchor":false},{"citing_arxiv_id":"2606.32025","citing_title":"Generative Skill Composition for LLM Agents","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14769","citing_title":"Composable Crystals: Controllable Materials Discovery via Concept Learning","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2501.02378","citing_title":"A ghost mechanism: An analytical model of abrupt learning in recurrent networks","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2503.18562","citing_title":"Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2505.13742","citing_title":"Understanding Task Representations in Neural Networks via Bayesian Ablation","ref_index":13,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF","json":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF.json","graph_json":"https://pith.science/api/pith-number/KVMXKZIACYBGHPMA2GWYU66VDF/graph.json","events_json":"https://pith.science/api/pith-number/KVMXKZIACYBGHPMA2GWYU66VDF/events.json","paper":"https://pith.science/paper/KVMXKZIA"},"agent_actions":{"view_html":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF","download_json":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF.json","view_paper":"https://pith.science/paper/KVMXKZIA","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.17567&json=true","fetch_graph":"https://pith.science/api/pith-number/KVMXKZIACYBGHPMA2GWYU66VDF/graph.json","fetch_events":"https://pith.science/api/pith-number/KVMXKZIACYBGHPMA2GWYU66VDF/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF/action/storage_attestation","attest_author":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF/action/author_attestation","sign_citation":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF/action/citation_signature","submit_replication":"https://pith.science/pith/KVMXKZIACYBGHPMA2GWYU66VDF/action/replication_record"}},"created_at":"2026-07-05T07:05:36.248507+00:00","updated_at":"2026-07-05T07:05:36.248507+00:00"}