Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either respond directly or generate explicit <think>...</think> reasoning traces before answering. This 70B variant is trained with the expanded post-training corpus (~60B tokens) emphasizing verified reasoning data, leading to improvements in mathematics, coding, STEM, logic, and structured outputs while maintaining general assistant performance. It supports JSON mode, schema adherence, function calling, and tool use, and is designed for greater steerability with reduced refusal rates.
Evaluations
13
across 11 benchmarks
Latency
555ms
Context length
131k
tokens
Cost
$0.11 · $0.38
input · output per 1M tokens
Key takeaways
Hermes 4 70B is a hybrid reasoning model built on Meta-Llama-3.1-70B, designed for advanced text generation and reasoning in mathematics, coding, STEM, and logic, supporting JSON mode and function calling.
The model utilizes a Transformer architecture and was trained with an expanded post-training corpus of approximately 60 billion tokens, emphasizing verified reasoning data to improve specialized domain performance.
The model performs well in fact retrieval for subjects like Clinical Knowledge and Anatomy (72.8%, 69.2% on MMLU), but struggles with complex multi-step reasoning, nuanced judgment, and precise calculations, as shown by its 3.33% accuracy on AIME 2024 and 3.85% on Humanity's Last Exam.
Common failure modes include incorrect factual recall, misinterpretation of ethical considerations (e.g., Moral Scenarios 44.6% on MMLU), calculation errors, and an inability to adhere strictly to specified output formats.
Despite its strengths in certain factual and 'next step' tasks, the model frequently exhibits overconfident but flawed reasoning, especially in advanced mathematics, demonstrating a need for improved precision and domain-specific reasoning.