GLM 5.2
Large reasoning model with 1M context, strong at long agent and coding tasks.
A curated selection of the most capable open models. From an 8-billion-parameter model on a single GPU to large mixture-of-experts models on a multi-GPU node. You choose what fits your domain, hardware class, and answer quality.
Large reasoning model with 1M context, strong at long agent and coding tasks.
Agentic model for long-horizon tasks, with coding variants.
Compact, efficient model — runs well on lean hardware.
Versatile MoE model for a single GPU, with tool calling and vision.
Cheap 1M context and native multimodality in one open model.
Tuned for reasoning and efficiency, optimized for NVIDIA hardware.
Mixture-of-experts with strong reasoning at low cost.
Lean EU model; Devstral specialized for coding.
Larger Mistral models — available only via cloud API.
OpenAI's open weight — a locally runnable all-round model.
Multilingual speech synthesis.
Expressive speech for long, multi-speaker text (up to 90 min).
Very natural speech with voice cloning and dialog.
Tiny (82M), fast TTS — very economical to run.
High-quality speech synthesis.
Natural, emotional speech (Llama-based).
Audio model for speech input and output.
Open audio model: transcription and speech understanding.
Very fast ASR (600M), 25 languages.
Speech recognition in the NVIDIA NeMo stack.
Efficient local transcription (Whisper in C/C++).
Speech recognition from the VibeVoice family.
Multilingual speech recognition.
Transcription service — cloud API only.
Open real-time transcription, with audio understanding.
Audio understanding and transcription.
Leading open image generation, up to 4MP, multi-reference.
Highly efficient 6B image model for fast generation.
The open classic of image generation, huge ecosystem.
Image generation and smart editing with text understanding.
Hybrid of autoregression and diffusion, strong at typography.
Photorealistic SDXL fine-tune, community favorite.
Ideogram's first open model (9.3B), strong at text in images.
Foundation image model (Raw/Turbo), tuned for aesthetics.
First open model with 4K audio+video in one pass, very fast.
One model for text-to-video, image-to-video, and editing; top photorealism.
Cinematic aesthetic, 8.3B, fast.
Video generation, geared toward longer clips.
Talking avatars / video from image and audio.
Video generation model.
Real-time video generation.
Sound design and audio textures (short samples / SFX).
Real-time music generation, open Google project.
Full songs in ~20s, with strong remix/edit tooling.
Text-to-music, stereo background music from prompts.
Text-to-audio generation.
Expressive dialogue speech synthesis.
Overview of current models by category. Unmarked: runs locally on the appliance; the ‘Cloud · not local' tag points to proprietary services shown for reference. Plus many more tasks — image and text classification, image segmentation, object detection, image masks, Gaussian splatting (3D from photos/videos), and more. Which model fits your domain and hardware we decide together.
You don't buy a frozen model. The hardware stays — when a better open model appears, the model moves up.
We select proven open models to fit your hardware class and use case — instead of a generic cloud model.
Your expertise is connected via RAG — your terms, your documents. The model itself is not retrained; your data stays in-house.
Quantized for exactly your hardware class — maximum answer quality per VRAM, not a generic cloud model.
When a better open model appears, it can be rolled out as an update — the hardware stays, the model moves up. A model change can alter how your workflows behave (the smaller the model, the more so), so we agree it with you beforehand and validate it together. Your data never leaves the building.
We do not decide that from a datasheet, but by your domain, hardware class, and answer quality — in the intro call.