lean-coder-welter
Welterweight class · v1.0 · Qwen 3-Coder-Next base
Code generation - debugging, refactoring, code review
80B total with 512 experts, only 3B active per token. A 49.7 GB model that runs on 12 GB VRAM. Tuned for code generation, debugging, and software engineering.
Specifications
Total params
80B
Active per token
3B
Base model
Qwen3-Coder-Next
Architecture
MoE (512 experts)
Experts
512 experts, 10 active
Min VRAM / RAM
12 GB / 32 GB
Tool calling
Yes - non-streaming
tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.
Source weights
The exact public GGUF this model is packed from, pinned to a commit. Download it and you can reproduce our numbers with stock llama.cpp.
Repository
Quantization
UD-Q4_K_XL
Revision
ce09c67b53bc8739eef83fe67b2f5d293c270632
Source size
49.61 GB (1 file)
This .lmpack
49.72 GB - core mostly Q8_0, experts mostly Q4_K
Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.
Measured performance
Real numbers from our reference rig, not projections.
| Hardware | Prefill | Decode |
|---|---|---|
| 1x RTX 3090 (24 GB) | 14-30 tok/s | 18-20 tok/s |
| 2x RTX 3090 | 9-20 tok/s | 18-28 tok/s |
Download
$ lean pull lean-coder-welterQuick start
$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean pull lean-coder-welter
$ lean run lean-coder-welterSingle binary, 15 MB. No Python, no Docker, no cloud dependency.