Pith. sign in

REVIEW 9 cited by

Hermes 3 Technical Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11857 v1 pith:TDTFIP4P submitted 2024-08-15 cs.CL

classification cs.CL
keywords modelshermesinstructabilitiesachievesbasebecomebenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruct (or "chat") tuned models have become the primary way in which most people interact with large language models. As opposed to "base" or "foundation" models, instruct-tuned models are optimized to respond to imperative statements. We present Hermes 3, a neutrally-aligned generalist instruct and tool use model with strong reasoning and creative abilities. Its largest version, Hermes 3 405B, achieves state of the art performance among open weight models on several public benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Adding 12% cross-repository code dependency contexts to the long-context fine-tuning mix improves long-range retrieval, state tracking, repo code understanding, and agentic tool use, while largely preserving short-con...

  2. DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A trace-grounded, effect-scored benchmark framework shows that even the strongest LLM agents solve only ~half of live MCP tasks, with accuracy collapsing on longer tool chains.

  3. Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A framework using statistical dissimilarity and LLM judges quantifies what fraction of the behavioral transition during fine-tuning is captured by each order parameter.

  4. MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language

    cs.CL 2025-08 conditional novelty 6.0 of 10

    MELAC introduces 19 Persian and Iranian-culture evaluation datasets and benchmarks 41 LLMs, showing weak performance on Iranian-specific content.

  5. Hermes 4 Technical Report

    cs.AI 2025-08 conditional novelty 5.0 of 10

    Hermes 4 releases three open-weight reasoning models (14B, 70B, 405B) trained with synthetic data and a length-control SFT stage, evaluated on mathematics, code, knowledge, and alignment benchmarks.

  6. LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data

    cs.CV 2026-01 reject novelty 4.0 of 10

    A lightweight RGB-D cross-attention network is proposed for rail defect detection, but the SOTA accuracy and generalization claims are internally inconsistent and the implementation is not public.

  7. Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection

    cs.NI 2025-08 unverdicted novelty 4.0 of 10

    A position paper proposing molecular communication as the link layer for epidemic-control bio-nano networks, with an ORF3a-based mutation identification simulation; the provided manuscript body does not match this abstract.

  8. Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method

    cs.CV 2026-01 reject novelty 3.0 of 10

    Fine-tuning Stable Diffusion with DreamBooth-style knowledge and hypernetwork-guided crack control maps can synthesize substation meter defect images that boost a YOLOv8 defect detector's mAP when added to the training set.

  9. Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

    cs.CR 2025-07 reject novelty 3.0 of 10

    Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.

Pith tools