
ModelNemotron 3.5 Lightning by Nvidia
Text->text
Free credits included. No credit card required.
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 30B total parameters and only 3B active per token, built for always-on agentic workloads. It combines a hybrid Mamba-Transformer MoE architecture with multi-token prediction to deliver up to 4x higher throughput than similarly sized open models. The model is suited for high-throughput agentic workloads and specialized tasks that benefit from domain-specific customization, including chatbots, RAG systems, and instruction-following applications. It supports a 262K-token context window and controllable reasoning.