Chirality emerges in SMILES translation models through an abrupt encoder-centered reorganization of representations after a long plateau, identified via checkpoint analysis and ablation.
Canonical reference
Title resolution pending
Canonical reference. 80% of citing Pith papers cite this work as background.
citation-role summary
citation-polarity summary
representative citing papers
A planner-executor multi-agent system using gpt-oss-120b and Parsl orchestrates scalable high-throughput MOF screening on the Aurora supercomputer with low overhead.
A lightweight VLM inspecting SMAC policy videos produces open-ended multi-agent curricula that outperform text-only ablations and PLR scalar-score methods on held-out maps.
Ahoy enables LLM agents to select and enact multiple declarative interaction protocols concurrently without specialized training to achieve goals.
Maistros 8B is a new state-of-the-art open-weights Greek LLM built via knowledge distillation from large reasoning models on the CulturaQA dataset.
Curtailing diversity in candidate pools for test-time scaling increases unsafe LLM outputs, as demonstrated by a reference-guided reduction protocol that evades standard safety classifiers across open and closed models.
CLAP is a closed-loop framework for domain agent post-training that integrates data processing, training, evaluation, and release gating, demonstrated on manufacturing data with modest gains in 3 of 5 batches.
Bash-Commenter applies CPT, SFT, and Syntax-Aware Preference Optimization (SAPO) via AST atomic operations to LLaMA-3.1-8B, reporting higher BLEU-4/METEOR/ROUGE-L scores than baselines on single-line and multi-line Bash comment generation tasks.
LLMs trained via rubric-based self-rewarding RL with GRPO enhanced feeling expression and sycophancy robustness but degraded truthful QA performance.
CTEM framework links behavioral history to evolving emotional states with user feedback updates, instantiated as Auri agent and tested in a 21-day study showing gains in naturalness, coherence, and emotional harmony.
Youth-authored synthesis argues LLM chatbots can temporarily reduce adolescent loneliness for some subgroups but risk deepening it for others, yielding three population-sensitive design implications.
GPT-4o exhibits daily and weekly periodic fluctuations in performance on a fixed physics task, accounting for about 20% of observed variance.
Users adjust AI agent personalities differently by task context, forming distinct profiles that increase perceived anthropomorphism, autonomy, and trust.
A survey of LLMs for graph computation introduces a role-based taxonomy of executors versus planners and concludes that current models suit simple small-scale tasks but remain unreliable for large-scale exact computation.
Qualitative studies show creatives prefer self-experimentation over structured guidance for GenAI image tools to preserve creative autonomy despite terminology barriers.
A survey synthesizing LLM and MM-LLM uses in transportation operations, mobility services, and decision support while noting challenges like data heterogeneity and real-time needs.
citing papers explorer
-
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
Chirality emerges in SMILES translation models through an abrupt encoder-centered reorganization of representations after a long plateau, identified via checkpoint analysis and ablation.
-
Multi-Agent Orchestration for High-Throughput Materials Screening on a Leadership-Class System
A planner-executor multi-agent system using gpt-oss-120b and Parsl orchestrates scalable high-throughput MOF screening on the Aurora supercomputer with low overhead.
-
Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs
A lightweight VLM inspecting SMAC policy videos produces open-ended multi-agent curricula that outperform text-only ablations and PLR scalar-score methods on held-out maps.
-
Ahoy: LLMs Enacting Multiagent Interaction Protocols
Ahoy enables LLM agents to select and enact multiple declarative interaction protocols concurrently without specialized training to achieve goals.
-
Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models
Maistros 8B is a new state-of-the-art open-weights Greek LLM built via knowledge distillation from large reasoning models on the CulturaQA dataset.
-
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
Curtailing diversity in candidate pools for test-time scaling increases unsafe LLM outputs, as demonstrated by a reference-guided reduction protocol that evades standard safety classifiers across open and closed models.
-
CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training
CLAP is a closed-loop framework for domain agent post-training that integrates data processing, training, evaluation, and release gating, demonstrated on manufacturing data with modest gains in 3 of 5 batches.
-
Bash-Commenter: Leveraging Syntax-Aware Preference Optimization to Reinforce Large Language Model for Bash Code Comment Generation
Bash-Commenter applies CPT, SFT, and Syntax-Aware Preference Optimization (SAPO) via AST atomic operations to LLaMA-3.1-8B, reporting higher BLEU-4/METEOR/ROUGE-L scores than baselines on single-line and multi-line Bash comment generation tasks.
-
When AI Says It Feels
LLMs trained via rubric-based self-rewarding RL with GRPO enhanced feeling expression and sycophancy robustness but degraded truthful QA performance.
-
Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
CTEM framework links behavioral history to evolving emotional states with user feedback updates, instantiated as Auri agent and tested in a 21-day study showing gains in naturalness, coherence, and emotional harmony.
-
Messages in a Digital Bottle: A Youth-Coauthored Perspective on LLM Chatbots and Adolescent Loneliness
Youth-authored synthesis argues LLM chatbots can temporarily reduce adolescent loneliness for some subgroups but risk deepening it for others, yielding three population-sensitive design implications.
-
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
GPT-4o exhibits daily and weekly periodic fluctuations in performance on a fixed physics task, accounting for about 20% of observed variance.
-
From Fixed to Flexible: Shaping AI Personality in Context-Sensitive Interaction
Users adjust AI agent personalities differently by task context, forming distinct profiles that increase perceived anthropomorphism, autonomy, and trust.
-
Are Large Language Models Suitable for Graph Computation? Progress and Prospects
A survey of LLMs for graph computation introduces a role-based taxonomy of executors versus planners and concludes that current models suit simple small-scale tasks but remain unreliable for large-scale exact computation.
-
How Creatives Approach GenAI Image Generation: Tensions Between Structured Guidance, Self-Experimentation, and Creative Autonomy
Qualitative studies show creatives prefer self-experimentation over structured guidance for GenAI image tools to preserve creative autonomy despite terminology barriers.
-
Large Language Models in Transportation Systems Management and Operations: From Text Reasoning to Multi-modal Decision Support
A survey synthesizing LLM and MM-LLM uses in transportation operations, mobility services, and decision support while noting challenges like data heterogeneity and real-time needs.