Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either respond directly or generate explicit <think>...</think> reasoning traces before answering. This 70B variant is trained with the expanded post-training corpus (~60B tokens) emphasizing verified reasoning data, leading to improvements in mathematics, coding, STEM, logic, and structured outputs while maintaining general assistant performance. It supports JSON mode, schema adherence, function calling, and tool use, and is designed for greater steerability with reduced refusal rates.
Evaluations
13
across 11 benchmarks
Latency
555ms
Context length
131k
tokens
Cost
$0.11 · $0.38
input · output per 1M tokens
Key takeaways
Hermes 4 70B is a hybrid reasoning model built on Meta-Llama-3.1-70B, designed for advanced text generation and reasoning in mathematics, coding, STEM, and logic, supporting JSON mode and function calling.
The model utilizes a Transformer architecture and was trained with an expanded post-training corpus of approximately 60 billion tokens, emphasizing verified reasoning data to improve specialized domain performance.
Notable abilities include strong factual recall in subjects like Clinical Knowledge (72.8%) and Anatomy (69.2%). However, it struggles with complex multi-step reasoning, nuanced judgment, and precise calculations.
Common failure modes include incorrect factual recall, misinterpretation of ethical considerations (Moral Scenarios 44.6% on MMLU), calculation errors, and an inability to adhere strictly to specified output formats.
The model performs poorly on advanced mathematics benchmarks (3.33% accuracy on AIME 2024), frequently exhibiting overconfident but flawed reasoning, especially in complex geometry, number theory, and combinatorial problems.