OPEN WEIGHTS · CURATED

The right model is not the biggest one.

A curated selection of the most capable open models. From an 8-billion-parameter model on a single GPU to large mixture-of-experts models on a multi-GPU node. You choose what fits your domain, hardware class, and answer quality.

§01

Model catalog.

FILTER BY CATEGORY · LOCAL & CLOUD
Category
Z.ai

GLM 5.2

Text

Large reasoning model with 1M context, strong at long agent and coding tasks.

Reasoning1M contextMoEMIT
Moonshot AI

Kimi K3

Text

Agentic model for long-horizon tasks, with coding variants.

Agenticlong contextApache 2.0
Google

Gemma 4

Text

Compact, efficient model — runs well on lean hardware.

Efficientcompact
Alibaba

Qwen3.6

Text

Versatile MoE model for a single GPU, with tool calling and vision.

All-roundMoEVisionApache 2.0
MiniMax

MiniMax M2.7

Text

Cheap 1M context and native multimodality in one open model.

Multimodal1M context
NVIDIA

Nemotron 3 · Super / Ultra

Text

Tuned for reasoning and efficiency, optimized for NVIDIA hardware.

Reasoningefficient
DeepSeek

DeepSeek V4 · Flash / Pro

Text

Mixture-of-experts with strong reasoning at low cost.

ReasoningMoEMIT
Mistral AI

Mistral Small 4 · Devstral 2

Text

Lean EU model; Devstral specialized for coding.

Code / all-roundApache 2.0
Mistral AI

Mistral Medium 3.5 · Large 3

Text

Larger Mistral models — available only via cloud API.

All-roundCloud · not local
OpenAI

GPT OSS

Text

OpenAI's open weight — a locally runnable all-round model.

All-roundApache 2.0
Alibaba

Qwen3-TTS

Text → speech

Multilingual speech synthesis.

Speech synthesismultilingual
Microsoft

VibeVoice

Text → speech

Expressive speech for long, multi-speaker text (up to 90 min).

Speech synthesismulti-speaker
BosonAI

Higgs TTS

Text → speech

Very natural speech with voice cloning and dialog.

Speech synthesisVoice cloningApache 2.0
Hexgrad

Kokoro

Text → speech

Tiny (82M), fast TTS — very economical to run.

Compact82M
Fish Audio

S2 Pro

Text → speech

High-quality speech synthesis.

Speech synthesis
Canopy AI

Orpheus

Text → speech

Natural, emotional speech (Llama-based).

Speech synthesis
Moonshot AI

Kimi Audio

Text → speech

Audio model for speech input and output.

Audio
Mistral AI

Voxtral

Text → speech

Open audio model: transcription and speech understanding.

AudioApache 2.0
NVIDIA

Parakeet

Speech → text

Very fast ASR (600M), 25 languages.

Transcription600M25 languages
NVIDIA

Nemotron ASR

Speech → text

Speech recognition in the NVIDIA NeMo stack.

Transcription
Whisper (OpenAI)

Whisper.cpp

Speech → text

Efficient local transcription (Whisper in C/C++).

TranscriptionCPU-capable
Microsoft

VibeVoice ASR

Speech → text

Speech recognition from the VibeVoice family.

Transcription
Alibaba

Qwen3 ASR

Speech → text

Multilingual speech recognition.

Transcriptionmultilingual
Cohere

Cohere Transcribe

Speech → text

Transcription service — cloud API only.

TranscriptionCloud · not local
Mistral AI

Voxtral (Realtime)

Speech → text

Open real-time transcription, with audio understanding.

Transcriptionreal-timeApache 2.0
BosonAI

Higgs Audio STT

Speech → text

Audio understanding and transcription.

Transcription
Black Forest Labs

FLUX.1 (Dev / Schnell) · Flux.2

Text → image

Leading open image generation, up to 4MP, multi-reference.

Text→imageup to 4MP
Tongyi Lab

Z-Image (Turbo)

Text → image

Highly efficient 6B image model for fast generation.

Text→image6B
Stability AI

Stable Diffusion (SDXL / 2)

Text → image

The open classic of image generation, huge ecosystem.

Text→image
Alibaba

Qwen Image (Edit)

Text → image

Image generation and smart editing with text understanding.

Text→image · edit
Z.ai

GLM Image

Text → image

Hybrid of autoregression and diffusion, strong at typography.

Text→image
RunDiffusion

Juggernaut XL

Text → image

Photorealistic SDXL fine-tune, community favorite.

Text→imageSDXL
Ideogram

Ideogram 4

Text → image

Ideogram's first open model (9.3B), strong at text in images.

Text→image9.3B
Krea

Krea 2 (Turbo)

Text → image

Foundation image model (Raw/Turbo), tuned for aesthetics.

Text→image
Lightricks

LTX Video

Text → video

First open model with 4K audio+video in one pass, very fast.

Text→video4Kaudio+video
Alibaba

Wan

Text → video

One model for text-to-video, image-to-video, and editing; top photorealism.

Text→video
Tencent

Hunyuan Video

Text → video

Cinematic aesthetic, 8.3B, fast.

Text→video8.3B
Meituan

LongCat Video

Text → video

Video generation, geared toward longer clips.

Text→video

OmniAvatar

Text → video

Talking avatars / video from image and audio.

Avatar

Neodragon

Text → video

Video generation model.

Text→video
Krea

Krea (Realtime)

Text → video

Real-time video generation.

Real-time videoreal-time
Stability AI

Stable Audio 3

Text → music

Sound design and audio textures (short samples / SFX).

Audio · sound
Google

Magenta Realtime

Text → music

Real-time music generation, open Google project.

Musicreal-time
ACE-Step

Ace-Step

Text → music

Full songs in ~20s, with strong remix/edit tooling.

Musicsongs
Meta

MusicGen

Text → music

Text-to-music, stereo background music from prompts.

Musicstereo

Tango2

Text → music

Text-to-audio generation.

Text→audio
Nari Labs

Dia

Text → music

Expressive dialogue speech synthesis.

Dialogue TTS

Overview of current models by category. Unmarked: runs locally on the appliance; the ‘Cloud · not local' tag points to proprietary services shown for reference. Plus many more tasks — image and text classification, image segmentation, object detection, image masks, Gaussian splatting (3D from photos/videos), and more. Which model fits your domain and hardware we decide together.

§02

Model choice & update path.

THE MODEL MOVES UP · THE HARDWARE STAYS

You don't buy a frozen model. The hardware stays — when a better open model appears, the model moves up.

01

Curated selection

We select proven open models to fit your hardware class and use case — instead of a generic cloud model.

02

Connected to your data

Your expertise is connected via RAG — your terms, your documents. The model itself is not retrained; your data stays in-house.

03

Quantization for your GPU

Quantized for exactly your hardware class — maximum answer quality per VRAM, not a generic cloud model.

04

Update path

When a better open model appears, it can be rolled out as an update — the hardware stays, the model moves up. A model change can alter how your workflows behave (the smaller the model, the more so), so we agree it with you beforehand and validate it together. Your data never leaves the building.

Which model fits you?

We do not decide that from a datasheet, but by your domain, hardware class, and answer quality — in the intro call.

Book intro call →
Book an intro call →