LLM Hangar / Models
Models
The catalog is curated: each entry is an open-weight model paired with at least one GPU shape we have booted it on, so the deploy flow can show an hourly and monthly estimate before anything is created. Prompts and responses go straight from your client to the endpoint on your instance. If the model you want is not listed, you can also deploy an arbitrary Hugging Face repository by pasting the repo id and picking a shape yourself.
Catalog and current prices
The table is read live from the same public price feed the homepage uses. Every eight hours we check which shapes are rentable at each provider and at what rate. A plain figure comes from a configuration we checked was deployable; a figure marked with ~ is an estimate from the provider's current GPU rates. Your provider bills you for the GPU; the prices here are not a quotation.
Guides
Every model we measure gets a guide at a stable URL: requirements first, then cold boot time, throughput and observed cost from real deployments in our own account, updated in place.
- DeepSeek V4 Flash: 284B mixture of experts, 2x H200, measured cold boot.
- GLM-5.3: open weights listed for 28 August 2026; requirements from the unchanged GLM-5.2 base.
- GLM-5.3-Flash: the model previewed as Ox Alpha, 320B with 18B active, MIT, 4x H200, measured 34-minute cold boot.
- Qwen3.8-27B: dense 27B, Apache 2.0, fits one L40S quantized or one H100 at bf16.
- Qwen3-Coder-Next: 80B mixture of experts with 3B active, as a private coding endpoint for a team.
- Gemma 4 31B: measured VRAM requirements, the real context ceiling on 80 GB cards, and the w4a16 build.
Kimi K3 is in the catalog and on the homepage price index; its size (2.8T parameters, 8x B300-class hardware per RunPod's technical FAQ) puts it outside what most teams will run, and we have not published a guide for it.
How a deployment works
- Connect a cloud account: AWS, Nebius, RunPod or Verda.
- Pick a model and a shape from this catalog; choose a region, or tick EU-only.
- Set a budget cap and confirm the estimate.
- Point any OpenAI-compatible client at the endpoint, as in using your endpoint.