REVIEW 20 cited by
Hermes 3 Technical Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Instruct (or "chat") tuned models have become the primary way in which most people interact with large language models. As opposed to "base" or "foundation" models, instruct-tuned models are optimized to respond to imperative statements. We present Hermes 3, a neutrally-aligned generalist instruct and tool use model with strong reasoning and creative abilities. Its largest version, Hermes 3 405B, achieves state of the art performance among open weight models on several public benchmarks.
Forward citations
Cited by 20 Pith papers
-
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Adding 12% cross-repository code dependency contexts to the long-context fine-tuning mix improves long-range retrieval, state tracking, repo code understanding, and agentic tool use, while largely preserving short-con...
-
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers
A trace-grounded, effect-scored benchmark framework shows that even the strongest LLM agents solve only ~half of live MCP tasks, with accuracy collapsing on longer tool chains.
-
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
A framework using statistical dissimilarity and LLM judges quantifies what fraction of the behavioral transition during fine-tuning is captured by each order parameter.
-
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
MELAC introduces 19 Persian and Iranian-culture evaluation datasets and benchmarks 41 LLMs, showing weak performance on Iranian-specific content.
-
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Safety refusal in Llama 3.1 8B is governed by a dominant activation direction plus smaller interpretable directions, and removing prompt tokens that activate these secondary directions can bypass fine-tuned safety.
-
It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers
Instruction-tuning a BERT-style encoder to answer with a single masked token turns its MLM head into a competitive zero-shot and fine-tuned classifier at 0.4B parameters.
-
Hermes 4 Technical Report
Hermes 4 releases three open-weight reasoning models (14B, 70B, 405B) trained with synthetic data and a length-control SFT stage, evaluated on mathematics, code, knowledge, and alignment benchmarks.
-
The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
Clustering-based routing plus self-consistency voting among ten 7B open models reportedly outranks GPT-4.1 and GPT-4.5 on average over 15 diverse benchmarks.
-
A Systematic Analysis of Base Model Choice for Reward Modeling
Reward model quality changes by up to 14% depending on the base language model chosen, and a combination of five standard benchmarks can partially guide that choice.
-
Fast Proxies for LLM Robustness Evaluation
Simple prompt-based and embedding-space attacks predict, with rank correlations up to 0.94, how open-source LLMs fare against a six-attack red-teaming ensemble, at roughly one thousandth of the compute.
-
Guided Code Generation with LLMs: A Multi-Agent Framework for Complex Code Tasks
A structured multi-agent framework reports 56.2% Pass@1 on HumanEval with Llama 3.1 8B int4, versus 45.4% for one-shot generation.
-
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
A noise-robust SFT framework that detects noisy responses via multi-expert LLM consensus, relabels them with context-enhanced reasoning, and filters low-confidence samples, improving LLM performance on five benchmarks.
-
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Adding a special refuse token to a fine-tuned LLM lets a developer tune refusal rates at inference time by thresholding the token's probability, with per-category control.
-
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
Fine-tuning experiments show instruction-following data improves LLM function-calling accuracy and relevance detection, a Decision Token plus synthetic negative examples helps non-relevant cases, and a tailored transl...
-
LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data
A lightweight RGB-D cross-attention network is proposed for rail defect detection, but the SOTA accuracy and generalization claims are internally inconsistent and the implementation is not public.
-
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection
A position paper proposing molecular communication as the link layer for epidemic-control bio-nano networks, with an ORF3a-based mutation identification simulation; the provided manuscript body does not match this abstract.
-
HASHIRU: Hierarchical Agent System for Hybrid Intelligent Resource Utilization
A hierarchical AI agent framework that dynamically hires and fires specialist models and creates tools reports strong benchmark numbers, though its gains may come from tool use rather than the architecture.
-
Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method
Fine-tuning Stable Diffusion with DreamBooth-style knowledge and hypernetwork-guided crack control maps can synthesize substation meter defect images that boost a YOLOv8 defect detector's mAP when added to the training set.
-
Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery
Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.
-
Reinforcement Learning Enhanced LLMs: A Survey
A survey that catalogs RL-enhanced LLMs and organizes alignment methods into RLHF, RLAIF, and DPO, without presenting new results.
Discussion (0). Continue with ORCID to comment.