Pith. sign in

REVIEW 3 cited by

Countering Language Drift via Visual Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.04499 v1 pith:Z77LW6MA submitted 2019-09-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagecommunicationdriftnaturalagentsconstraintseasilygrounding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Emergent multi-agent communication protocols are very different from natural language and not easily interpretable by humans. We find that agents that were initially pretrained to produce natural language can also experience detrimental language drift: when a non-linguistic reward is used in a goal-based task, e.g. some scalar success metric, the communication protocol may easily and radically diverge from natural language. We recast translation as a multi-agent communication game and examine auxiliary training constraints for their effectiveness in mitigating language drift. We show that a combination of syntactic (language model likelihood) and semantic (visual grounding) constraints gives the best communication performance, allowing pre-trained agents to retain English syntax while learning to accurately convey the intended meaning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    SceneBooth keeps a provided subject image untouched and paints a new background around it, guided by a caption, object labels, and a predicted scene layout.

  2. DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    DreamBlend guides an overfit fine-tuned checkpoint with cross-attention maps from an underfit checkpoint, improving subject fidelity, prompt fidelity, and diversity in personalized text-to-image generation.

  3. MyTimeMachine: Personalized Facial Age Transformation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A personalized facial age transformation method that uses an adapter network on top of the SAM global aging model, trained with 10 to 50 photos of one person, to produce re-aged images that resemble that person's actu...

Pith tools