lean-agent-light
Lightweight class · v1.1 · Qwen 3.5 base
General-purpose agent - tool calling, structured output, multi-step reasoning
The entry point. A 22 GB model that runs on 12 GB VRAM - expert offloading handles the rest. Qwen3.5 GDN hybrid architecture surpasses last-gen models many times its size.
Specifications
Total params
35B
Active per token
3B
Base model
Qwen3.5-35B-A3B
Architecture
GDN hybrid MoE
Experts
256 experts, 8 active
Min VRAM / RAM
12 GB / 16 GB
Tool calling
Yes - non-streaming
tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.
Source weights
The original model release this pack is built from, pinned to an exact commit. Our own build pipeline quantizes and tunes it for consumer GPUs, and every release is tested against the original model on standard benchmarks before it ships.
Repository
Quantization
bf16
Revision
59d61f3ce65a6d9863b86d2e96597125219dc754
Source size
71.9 GB (14 files)
This .lmpack
21.60 GB - core mostly Q6_K, experts mostly Q4_K
Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.
Measured performance
Real numbers from our reference rig, not projections.
| Hardware | Prefill | Decode |
|---|---|---|
| 1x RTX 3090 (24 GB) | 16-41 tok/s | 30-40 tok/s |
| 2x RTX 3090 | 11-27 tok/s | 24-37 tok/s |
Download
$ lean pull lean-agent-lightQuick start
$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean pull lean-agent-light
$ lean run lean-agent-lightSingle binary, 15 MB. No Python, no Docker, no cloud dependency.