CLI reference
lean <command> [options]Run lean --help or lean <command> --help for the same information from the
binary itself.
The model argument
Section titled “The model argument”Most commands take a <model> argument, which is either:
- a model name (
lean-agent-light), resolved against~/.leanmodels/models/<name>/, or - a name and version (
lean-agent-light:1.0), or - a path to a directory containing a
manifest.json.
A model whose download was interrupted is rejected with a message telling you to
resume it with lean pull.
Versions
Section titled “Versions”Models are versioned independently of each other and of the lean binary, so a
new release of one model does not move any other. Version numbers are ours, not
the base model’s: a major bump means a new base generation, a minor means a
better pack of the same base.
Leaving the version off means latest:
lean pull lean-agent-light # latestlean pull lean-agent-light:1.0 # exactly 1.0Versions install side by side under ~/.leanmodels/models/<name>/<version>/, so
pulling a new one does not remove the old, and you can pin back to it at any
time. lean models lists what you have installed; lean models --remote lists
what is published.
When a model needs a newer lean
Section titled “When a model needs a newer lean”A model version can require a minimum lean version - usually because it uses a
base architecture that older builds have no runtime support for. The requirement
is recorded in the model itself, so it is enforced whether you downloaded it
today or a year ago.
lean pull checks it before downloading any weights, and tells you both ways
forward:
lean-agent-cruiser:2.0 needs lean 0.3.0 or newer, but this is lean 0.2.1.
This model version uses runtime support that v0.2.1 does not have.
Upgrade lean: curl -fsSL https://leanmodels.ai/install.sh | sh
Or stay on this version of lean and use an earlier model version: lean models --remote # lists versions and the lean each needs lean pull lean-agent-cruiser:<version>lean models --remote flags any published version your binary cannot load, so
you can see what is coming before it blocks you.
Running models
Section titled “Running models”lean run
Section titled “lean run”Interactive chat REPL with multi-turn history.
lean run lean-agent-light --temperature 0.3| Flag | Default | Description |
|---|---|---|
--max-tokens <n> |
2048 |
Maximum tokens generated per response. |
--system <text> |
none | System prompt for the session. |
--temperature <f> |
0.7 |
Sampling temperature. 0 is greedy. |
--gpu <n> |
auto | Pin to one GPU by ordinal. |
--pipeline |
off | Split layers across all GPUs instead of using one. |
--no-think |
off | Skip the <think> phase on models that support it. |
In-REPL commands: /stats for cache and throughput numbers, /clear to reset
the conversation.
lean serve
Section titled “lean serve”Start an OpenAI-compatible HTTP server. See the API reference for endpoints and payloads.
lean serve lean-agent-light --port 8080| Flag | Default | Description |
|---|---|---|
--host <addr> |
127.0.0.1 |
Bind address. Use 0.0.0.0 to expose on your network. |
--port <n> |
8080 |
Listen port. |
--gpu <n> |
auto | Pin to one GPU by ordinal. |
--pipeline |
off | Split layers across all GPUs. |
--api-key <key> |
none | Require Authorization: Bearer <key> on /v1 routes. Also read from LEAN_API_KEY. |
--request-timeout-secs <n> |
600 |
Abort a generation after this long. 0 disables the timeout. |
Binding to a non-loopback address without an API key logs a warning: anyone who can reach the port gets unauthenticated inference.
Single-GPU vs pipeline
Section titled “Single-GPU vs pipeline”By default the runtime uses one GPU when the model fits on it, because
batch-1 decoding gains nothing from pipeline parallelism - the GPUs just run
sequentially with a hidden-state hop between them each token. --pipeline
forces the split across all cards, which helps when a model’s expert working set
is too large to cache well on a single card. --gpu <n> and --pipeline are
mutually exclusive.
Getting models
Section titled “Getting models”lean pull
Section titled “lean pull”Download a model from the registry. Resumable, checksum-verified, and it checks free disk before starting.
lean pull lean-agent-light # latestlean pull lean-agent-light:1.0 # a specific version| Flag | Default | Description |
|---|---|---|
--registry <url> |
https://registry.leanmodels.ai |
Registry to pull from. Also read from LEAN_REGISTRY_URL. |
--minimal |
off | Fetch the core plus hot experts only, skipping cold experts. |
--output <dir> |
~/.leanmodels/models/<name>/<version> |
Where to write the model. |
--minimal produces a working model sooner and on less disk. Cold experts are
the ones routing rarely selects, so quality degrades only on the inputs that
need them.
lean login / lean logout
Section titled “lean login / lean logout”Paid models require an API key.
lean login # prompts for the keylean login --api-key lm_live_...lean logout # removes stored credentials| Flag | Description |
|---|---|
--api-key <key> |
Supply the key non-interactively. |
--registry <url> |
Registry to authenticate against. |
--no-verify |
Store the key without checking it against the registry first. |
lean models
Section titled “lean models”List locally installed models, or the registry catalog with --remote.
lean modelslean models --remotelean verify
Section titled “lean verify”Walk every checksum in a model’s manifest and report damaged or missing files.
Damaged models are marked so the next lean pull repairs them.
lean verify lean-agent-middleKeeping lean current
Section titled “Keeping lean current”lean update
Section titled “lean update”Replace the binary with the latest published release.
lean updatelean update --check # report whether an update exists, install nothingIt resolves the newest release, downloads the archive matching your platform
and build kind — a CUDA install stays a CUDA install — verifies it against
the release’s SHA256SUMS, and swaps the binary in place. A missing or
mismatched checksum aborts the update; it never installs unverified bytes.
lean --version reports the release tag you are on, so v0.1.0-rc5 and
v0.1.0-rc6 are distinguishable rather than both reading as 0.1.0.
If lean lives somewhere you cannot write to, the update says so and points you
at the installer, which can be run with whatever permissions that location needs:
curl -sSf https://leanmodels.ai/install.sh | shA lean you built from source reports itself as a local build and is left alone
unless you pass --force.
Inspecting and measuring
Section titled “Inspecting and measuring”lean info
Section titled “lean info”Print a model’s manifest: architecture, parameter counts, expert layout, quant,
hardware requirements, and a one-line summary of the source weights (see
lean provenance for the full chain).
lean info lean-agent-lightlean provenance
Section titled “lean provenance”Print the source chain for a model: which public repository it was packed from - a community GGUF or the original model checkpoint - the exact commit, and the sha256 of every recorded source file.
lean provenance lean-agent-lightFor HuggingFace sources the printed sha256 is the same value the repository’s
API serves as each file’s object id, so the claim can be checked without
downloading the model. The variant named under variant is the source’s own
label (Q4_K_M, UD-Q4_K_XL, or bf16 for packs quantized directly from the
original checkpoint); the core/experts lines below it describe this pack’s
own tensors, which are not the same thing.
Models packed before provenance was recorded print not recorded rather than a
guess.
lean tune
Section titled “lean tune”Detect hardware and recommend models and settings. Takes no arguments. Reports the detected GPUs and RAM, which locally installed models fit in VRAM, RAM, or need NVMe, and the expert cache budget available per GPU.
lean tunelean bench
Section titled “lean bench”Measure prefill and decode throughput, cache hit rates, and memory use.
lean bench lean-agent-light| Flag | Default | Description |
|---|---|---|
--gpu <n> |
auto | Pin to one GPU by ordinal. |
--pipeline |
off | Split layers across all GPUs. |
--ram-budget <mb> |
auto | Cap the host-side expert staging budget. |
--timed |
off | Print a per-layer timing breakdown. |
--runs <n> |
3 |
Repeat each phase this many times and report median plus min/max. |
Lowering --ram-budget simulates a machine with less RAM than the one running
the benchmark.
lean smoke
Section titled “lean smoke”Run a fixed set of diverse prompts to sanity-check output quality after an install, a pull, or a hardware change.
lean smoke lean-agent-light --verbose| Flag | Description |
|---|---|
--gpu <n> |
Pin to one GPU by ordinal. |
--pipeline |
Split layers across all GPUs. |
--verbose |
Print the generated output for each prompt. |
Environment variables
Section titled “Environment variables”| Variable | Effect |
|---|---|
LEAN_REGISTRY_URL |
Default registry for pull, login, and models --remote. |
LEAN_API_KEY |
API key required by lean serve on /v1 routes. |
RUST_LOG |
Log filter. Defaults to info; RUST_LOG=debug for more detail. |