Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury(opens in new tab)) here.
Modalities
Context
128K
Released
Jun 26, 2025
Knowledge Cutoff
Jan 2025
Token volume and request traffic to this model over time.
Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots.
Mercury has a 128,000 token context window.
Mercury 2 is another text model from Inception.
Mercury was released on June 26, 2025. Its knowledge cutoff is January 31, 2025.