GRAM selectively trains auxiliary modules so that ablating one at inference removes a targeted capability while preserving the rest, closely tracking data-filtered models at 5x lower cost across 5 capability profiles.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
CONDITIONAL 3representative citing papers
Clinical narrative format beats raw JSON for LLMs up to 8B parameters on medication reconciliation but raw JSON wins at 70B scale, with omissions as the main error type.
Diffusion language models encode a decodable, steerable representation of denoising progress (the fraction of unmasked tokens) in their residual streams.
citing papers explorer
-
Modular Pretraining Enables Access Control
GRAM selectively trains auxiliary modules so that ablating one at inference removes a targeted capability while preserving the rest, closely tracking data-filtered models at 5x lower cost across 5 capability profiles.
-
Serialisation Strategy Matters: How FHIR Data Format Affects LLM Medication Reconciliation
Clinical narrative format beats raw JSON for LLMs up to 8B parameters on medication reconciliation but raw JSON wins at 70B scale, with omissions as the main error type.
-
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Diffusion language models encode a decodable, steerable representation of denoising progress (the fraction of unmasked tokens) in their residual streams.