← All models

lean-coder-welter

Welterweight class · v1.0 · Qwen 3-Coder-Next base

Code generation - debugging, refactoring, code review

Verified vs llama.cppApache 2.0 baseFree

80B total with 512 experts, only 3B active per token. A 49.7 GB model that runs on 12 GB VRAM. Tuned for code generation, debugging, and software engineering.

Specifications

Total params

80B

Active per token

3B

Base model

Qwen3-Coder-Next

Architecture

MoE (512 experts)

Experts

512 experts, 10 active

Min VRAM / RAM

12 GB / 32 GB

Tool calling

Yes - non-streaming

tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.

Source weights

The exact public GGUF this model is packed from, pinned to a commit. Download it and you can reproduce our numbers with stock llama.cpp.

Quantization

UD-Q4_K_XL

Revision

ce09c67b53bc8739eef83fe67b2f5d293c270632

Source size

49.61 GB (1 file)

This .lmpack

49.72 GB - core mostly Q8_0, experts mostly Q4_K

Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.

Measured performance

Real numbers from our reference rig, not projections.

HardwarePrefillDecode
1x RTX 3090 (24 GB)14-30 tok/s18-20 tok/s
2x RTX 30909-20 tok/s18-28 tok/s

Download

Download 49.72 GB - single .lmpack, from UD-Q4_K_XL
$ lean pull lean-coder-welter

Quick start

$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean pull lean-coder-welter
$ lean run lean-coder-welter

Single binary, 15 MB. No Python, no Docker, no cloud dependency.