Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Ontology-amplified distillation of a 27B student matches GPT-5 grounding on Vietnamese financial tasks, while contextuality audits return zero residual contextuality for routing.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 09:22 UTC pith:XS6QH2PH

load-bearing objection Honest, underpowered pilot: local 27B student matches GPT-5 grounding counts on 40 Vietnamese finance tasks after ontology-amplified distillation from 47 synthetic pairs, plus a clean negative contextuality result—claims stay inside their own hedges. the 4 major comments →

arxiv 2607.11948 v1 pith:XS6QH2PH submitted 2026-07-11 cs.AI cs.CLcs.LGcs.MA

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

classification cs.AI cs.CLcs.LGcs.MA
keywords ontology-amplified distillationsovereign enterprise language modelsdirect preference optimizationcontextuality auditingContextuality-by-Defaultfinancial-domain groundingagent routing governancetenant-owned LLMs
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Regulated financial institutions need language models they can own and run inside their own perimeter. This paper pairs a reduced-power proof-of-mechanism for ontology-amplified distillation with a negative-results contextuality audit for agent routing. A Qwen3.6-27B student is adapted to a Foundation AgenticOS ontology by supervised fine-tuning on frontier-teacher trajectories plus ontology-grounded direct preference optimization, all from 47 synthetic English cross-domain preference pairs trained locally on one Apple M5 Max. On 40 held-out Vietnamese financial-domain tasks the student grounds 36 of 40 (rate 0.90, mean ontology term-coverage 0.95 on a metric floored at 0.50), matching the GPT-5 frontier baseline exactly, yet the study is underpowered to claim equivalence and does not show the pre-registered amplification that the student should exceed the frontier. A separate pilot finds the corrected Contextuality-by-Default degree is zero for every Phase 1.3 group, so the useful routing signal is direct influence and construct coupling rather than residual contextuality. The combined evidence is offered as a mechanism-plus-governance diagnostic, not as a claim of deployability, safety, superiority, or statistical equivalence.

Core claim

Ontology-amplified distillation can bring a 27B student to the same ontology-grounding rate as a GPT-5 frontier teacher on held-out Vietnamese financial tasks (36/40 each), while a contextuality audit of the same agent-routing setting yields zero residual Contextuality-by-Default degree, indicating that direct influence and construct coupling, not surviving contextuality, should drive governance decisions.

What carries the argument

Ontology-amplified distillation (supervised fine-tuning on frontier-teacher trajectories followed by ontology-grounded DPO on synthetic preference pairs) together with the corrected canonical Contextuality-by-Default degree as a governance diagnostic for when apparent agent disagreement warrants standardization, multi-agent synthesis, or human review.

Load-bearing premise

That 47 synthetic English-language preference pairs plus frontier-teacher trajectories, trained on a single laptop, form a sufficient and unbiased signal for ontology grounding that generalizes cleanly to held-out Vietnamese financial tasks without synthetic-data artifacts or floor effects from the coverage metric.

What would settle it

Re-run the identical 40-task Vietnamese financial evaluation with a larger, non-synthetic, Vietnamese-native preference set and a coverage metric that is not floored at 0.50; if the student then falls well below the GPT-5 grounding rate or residual Contextuality-by-Default becomes reliably positive, the claimed mechanism and diagnostic both fail.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript combines two FAOS studies into a single mechanism-and-control article. First, it reports a reduced-power proof-of-mechanism for ontology-amplified distillation: a Qwen3.6-27B student is adapted via supervised fine-tuning on frontier-teacher trajectories and ontology-grounded DPO from 47 synthetic English cross-domain preference pairs, trained locally on one Apple M5 Max. On 40 held-out Vietnamese financial-domain tasks the student grounds 36/40 (rate 0.90; mean r_onto = 0.95 on a metric floored at 0.50), matching GPT-5; the design is underpowered for equivalence (paired-difference 95% CI spans ±4 tasks) and does not test or confirm the pre-registered amplification prediction that the student should exceed the frontier. Second, a separate negative-results pilot of a corrected canonical Contextuality-by-Default audit finds zero residual contextuality for all Phase 1.3 groups in both a local-Qwen run and a Gemma replication; the useful signal is direct influence and construct coupling. The abstract explicitly states that the evidence supports neither deployability, safety, superiority, statistical equivalence, nor a contextuality-positive routing rule.

Significance. If the distillation mechanism generalizes under fuller evaluation, ontology-grounded local adaptation of open-weight models would be practically relevant for regulated financial institutions under data-residency constraints. Pairing that mechanism with a contextuality-audit diagnostic for enterprise-agent routing is a coherent methodological contribution: it links model-building to a governance decision rule (prompt standardization, multi-agent synthesis, or human review). Explicit credit is due for pre-registering the amplification prediction, for reporting a failed prediction and underpowered CI without overclaiming, for the negative contextuality result, and for the clear non-claims about deployability and safety. The work is best read as a carefully scoped proof-of-mechanism plus negative-results pilot rather than a definitive empirical or safety result.

major comments (4)
  1. [Abstract (training setup and evaluation)] The load-bearing training signal is 47 synthetic English-language cross-domain preference pairs plus frontier trajectories, evaluated on held-out Vietnamese financial-domain tasks. The abstract’s descriptive match (36/40 vs GPT-5) is interesting but does not, by itself, establish that the ontology-grounding mechanism transfers rather than reflecting synthetic-data artifacts, teacher-trajectory leakage, or domain mismatch. A major revision should add ablations or controls (e.g., non-ontology DPO, language-matched pairs, or artifact audits) that isolate ontology amplification from these confounds; without them the mechanism claim remains under-supported even as a proof-of-mechanism.
  2. [Abstract (r_onto definition and 36/40 result)] Mean r_onto = 0.95 is reported on a metric floored at 0.50. Flooring compresses the lower tail and can inflate both the mean and the apparent grounding quality; the 36/40 grounded count is also tied to this ontology term-coverage construction, which aligns with ontology-grounded DPO training. Report unfloored r_onto (mean, distribution, and per-task values), the fraction of tasks affected by the floor, and a sensitivity analysis with an alternative grounding criterion that is not term-coverage of the training ontology. This is load-bearing for interpreting the student–frontier match.
  3. [Title; Abstract (pre-registered amplification prediction)] The title and framing use “Ontology-Amplified Distillation,” yet the abstract states that the run does not test or show the pre-registered amplification prediction (student should exceed the frontier) and that the outcome is underpowered for equivalence. Either temper the title/framing to “ontology-grounded” or “ontology-conditioned” distillation, or add a dedicated subsection that treats the failed amplification prediction as a primary result and revises the mechanism claim accordingly. Leaving “amplified” in the title while reporting a null on amplification is a load-bearing framing inconsistency.
  4. [Abstract (combined studies framing)] The two studies (distillation proof-of-mechanism; contextuality-audit negative pilot) are presented as a combined mechanism-and-control article, but the abstract does not show a shared experimental spine, shared tasks, or a decision rule that uses the audit to gate the distilled model. Clarify the logical link: either integrate them on the same task suite with an explicit routing policy, or restructure as two loosely coupled contributions with separate claims. As written, the “combined” framing risks overstating unity without a load-bearing bridge.
minor comments (5)
  1. [Abstract (grounded rate 0.90)] State the exact operational definition of a “grounded” task (threshold on r_onto or other rule) so the 36/40 count is reproducible from the metric description alone.
  2. [Abstract (baselines)] Name the GPT-5 and Gemma model snapshots/dates and the decoding settings used for the frontier baseline and the replication check.
  3. [Abstract (contextuality audit)] Expand the one-line description of the “corrected canonical Contextuality-by-Default degree (Phase 1.3)” with a pointer to the formula or prior definition so readers can interpret the zero result without external FAOS lore.
  4. [Abstract (underpowered equivalence statement)] Report the paired-difference point estimate alongside the ±4-task 95% CI, not only the interval width.
  5. [Abstract / data availability] If code, preference pairs, or audit scripts will be released, state the license and repository plan; if not, say so explicitly given the reproducibility emphasis of a proof-of-mechanism study.

Circularity Check

0 steps flagged

No significant circularity; empirical report with explicit hedges, not a derivation that reduces to its inputs by construction.

full rationale

Only the abstract is available. The load-bearing claims are empirical measurements on 40 held-out Vietnamese financial-domain tasks (student grounds 36/40, r_onto mean 0.95 floored at 0.50, matching GPT-5's 36/40) after SFT+DPO on 47 synthetic English ontology-grounded preference pairs, plus a separate negative-results contextuality pilot reporting corrected degree zero. The paper itself states the outcome is underpowered (paired-difference 95% CI spans ±4 tasks), fails the pre-registered amplification prediction that the student should exceed the frontier, and supports neither equivalence, deployability, nor a contextuality-positive routing rule. Training on ontology-grounded DPO and evaluating ontology term-coverage is intentional objective alignment common in distillation work; it does not make the 36/40 count or the failed amplification claim equivalent to the inputs by construction, nor is any free parameter fitted and then renamed as a first-principles prediction. No uniqueness theorem, self-citation chain, or ansatz is invoked as load-bearing in the abstract. Per the rules, mild metric–training relatedness without a quotable reduction does not raise the score. Finding: no significant circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 2 invented entities

Abstract-only review: free parameters and axioms are those explicitly named or implied by the abstract’s training and metric setup. No invented physical entities; the main constructs are methodological (ontology-amplified distillation recipe and Contextuality-by-Default audit). Synthetic-pair generation and the 0.50 floor on r_onto are the load-bearing unstated or lightly stated choices.

free parameters (2)
  • r_onto floor = 0.50
    Ontology term-coverage metric is floored at 0.50, which directly raises the reported mean r_onto = 0.95 and can mask low-coverage failures.
  • synthetic preference-pair set size and content = 47 pairs
    47 synthetic English cross-domain preference pairs are the entire DPO signal; their generation process and sampling are not specified in the abstract and act as an effective free design choice.
axioms (4)
  • domain assumption Ontology term-coverage (r_onto) and binary grounding on held-out tasks are valid proxies for ontology-amplified capability.
    Central evaluation rests on these metrics equaling the GPT-5 baseline; abstract does not justify why coverage/grounding imply deployable financial competence.
  • domain assumption Synthetic English frontier-teacher trajectories transfer to Vietnamese financial-domain held-out tasks.
    Training language and domain differ from evaluation language and domain; transfer is assumed without reported controls.
  • domain assumption Corrected Contextuality-by-Default degree is the right residual after removing direct influence and construct coupling.
    Negative-results pilot concludes residual contextuality is zero; validity depends on the correction procedure being complete.
  • standard math Standard SFT + DPO optimization dynamics apply on a single Apple M5 Max for a 27B student.
    Uses conventional supervised fine-tuning and direct preference optimization; no new optimizer claimed.
invented entities (2)
  • Ontology-amplified distillation (FAOS-specific recipe) no independent evidence
    purpose: Name the combined SFT-on-teacher-trajectories + ontology-grounded DPO procedure used to adapt the student to the Foundation AgenticOS ontology.
    Methodological construct rather than a new physical entity; independent evidence is limited to the single underpowered run reported.
  • Corrected canonical Contextuality-by-Default degree (Phase 1.3) no independent evidence
    purpose: Quantify residual contextuality for enterprise-agent routing after accounting for direct influence and construct coupling.
    Diagnostic quantity whose value is reported as zero; no external falsifiable handle beyond the two pilot runs.

pith-pipeline@v1.1.0-grok45 · 6274 in / 3420 out tokens · 26692 ms · 2026-07-15T09:22:09.735333+00:00 · methodology

0 comments
read the original abstract

Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter. This paper combines two related FAOS studies into one mechanism-and-control article. First, it reports a reduced-power proof-of-mechanism study of ontology-amplified distillation: a Qwen3.6-27B student is adapted to the Foundation AgenticOS ontology through supervised fine-tuning on frontier-teacher trajectories and ontology-grounded direct preference optimization (DPO), trained locally on a single Apple M5 Max from 47 synthetic, English-language, cross-domain preference pairs. On 40 held-out Vietnamese financial-domain tasks, the distilled student grounds 36 of 40 tasks (grounded rate 0.90; mean ontology term-coverage r_onto = 0.95 on a metric floored at 0.50), equal to the GPT-5 frontier baseline, which also grounds 36 of 40. The outcome is underpowered to establish equivalence: the paired-difference 95% confidence interval spans +/-4 tasks, and the run does not test or show the pre-registered amplification prediction that the student should exceed the frontier. Second, the paper consolidates a contextuality-audit method for enterprise-agent routing. In a separate negative-results pilot, the corrected canonical Contextuality-by-Default degree is zero for all Phase 1.3 groups in both the local-Qwen run and an explicitly labeled Gemma replication check; the useful signal is direct influence and construct coupling, not surviving residual contextuality. Together, the studies pair an ontology-grounded model-building mechanism with a governance diagnostic for deciding when apparent disagreement should trigger prompt standardization, multi-agent synthesis, or human review. The evidence supports neither deployability, safety, superiority, statistical equivalence, nor a contextuality-positive routing rule.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

    cs.CL 2026-07 conditional novelty 7.0

    Forced-binary next-token probabilities from an instruction-tuned LLM were saturated (near-deterministic) for 17/18 question pairs, so QQ-equality verdicts did not identify a response mechanism; saturation screening sh...