
ModelMercury 2.5 by Inception
Text->text
Free credits included. No credit card required.
Mercury 2.5 is the most capable diffusion large language model (dLLM) from Inception and the fastest reasoning LLM in production. Unlike autoregressive models that generate tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving over 1,100 tokens per second on widely-available NVIDIA GPUs. It features a 260K token context window, tunable reasoning, parallel tool calls, and schema-aligned JSON output. Mercury 2.5 delivers intelligence comparable to cost-optimized frontier models while maintaining exceptional throughput and low cost.