{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:7A7JGKLHBGDVDEQH3BRLTECOBH","short_pith_number":"pith:7A7JGKLH","schema_version":"1.0","canonical_sha256":"f83e9329670987519207d862b9904e09f83e9a16dab1437f24a06443e31e19a7","source":{"kind":"arxiv","id":"2406.00060","version":1},"attestation_state":"computed","paper":{"title":"Cascade-Aware Training of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CL","authors_text":"Aditya Krishna Menon, Alec Go, Ankit Singh Rawat, Congchao Wang, Harikrishna Narasimhan, Keith Rush, Sean Augenstein, Wittawat Jitkrittum","submitted_at":"2024-05-29T22:28:46Z","abstract_excerpt":"Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications. To address this, cascades of LMs offer an effective solution that conditionally employ smaller models for simpler queries. Cascaded systems are typically built with independently trained models, neglecting the advantages of considering inference-time interactions of the cascaded LMs during training. In this paper, we present cascade-aware training(CAT), an approach to optimizing the overall quality-cost performance tradeoff of a cascade of LMs. We achieve inference-t"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.00060","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-05-29T22:28:46Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"d4e1cb0306470b40ac8ffe3a0b2b7f1c4249fdbb985022bb5d549991f8be8f93","abstract_canon_sha256":"12e544920dbab94152dbfc319b6a6ba04e6c587b792ff637427173a36bd58b99"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:25:56.198935Z","signature_b64":"G8Lbn+g+q9oMLwUCUEamDt9tuuCRbwqlzR8/Dz0W6OtrlV8kqv/L29kYBw/uCSToJukKl2kkK92EFqcYYzRrAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"f83e9329670987519207d862b9904e09f83e9a16dab1437f24a06443e31e19a7","last_reissued_at":"2026-07-05T08:25:56.198459Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:25:56.198459Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Cascade-Aware Training of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CL","authors_text":"Aditya Krishna Menon, Alec Go, Ankit Singh Rawat, Congchao Wang, Harikrishna Narasimhan, Keith Rush, Sean Augenstein, Wittawat Jitkrittum","submitted_at":"2024-05-29T22:28:46Z","abstract_excerpt":"Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications. To address this, cascades of LMs offer an effective solution that conditionally employ smaller models for simpler queries. Cascaded systems are typically built with independently trained models, neglecting the advantages of considering inference-time interactions of the cascaded LMs during training. In this paper, we present cascade-aware training(CAT), an approach to optimizing the overall quality-cost performance tradeoff of a cascade of LMs. We achieve inference-t"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.00060","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.00060/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.00060","created_at":"2026-07-05T08:25:56.198521+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.00060v1","created_at":"2026-07-05T08:25:56.198521+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.00060","created_at":"2026-07-05T08:25:56.198521+00:00"},{"alias_kind":"pith_short_12","alias_value":"7A7JGKLHBGDV","created_at":"2026-07-05T08:25:56.198521+00:00"},{"alias_kind":"pith_short_16","alias_value":"7A7JGKLHBGDVDEQH","created_at":"2026-07-05T08:25:56.198521+00:00"},{"alias_kind":"pith_short_8","alias_value":"7A7JGKLH","created_at":"2026-07-05T08:25:56.198521+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25871","citing_title":"AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sponsored Search","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH","json":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH.json","graph_json":"https://pith.science/api/pith-number/7A7JGKLHBGDVDEQH3BRLTECOBH/graph.json","events_json":"https://pith.science/api/pith-number/7A7JGKLHBGDVDEQH3BRLTECOBH/events.json","paper":"https://pith.science/paper/7A7JGKLH"},"agent_actions":{"view_html":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH","download_json":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH.json","view_paper":"https://pith.science/paper/7A7JGKLH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.00060&json=true","fetch_graph":"https://pith.science/api/pith-number/7A7JGKLHBGDVDEQH3BRLTECOBH/graph.json","fetch_events":"https://pith.science/api/pith-number/7A7JGKLHBGDVDEQH3BRLTECOBH/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH/action/storage_attestation","attest_author":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH/action/author_attestation","sign_citation":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH/action/citation_signature","submit_replication":"https://pith.science/pith/7A7JGKLHBGDVDEQH3BRLTECOBH/action/replication_record"}},"created_at":"2026-07-05T08:25:56.198521+00:00","updated_at":"2026-07-05T08:25:56.198521+00:00"}