Skip to content

CLI reference

Terminal window
lean <command> [options]

Run lean --help or lean <command> --help for the same information from the binary itself.

Most commands take a <model> argument, which is either:

  • a model name (lean-agent-light), resolved against ~/.leanmodels/models/<name>/, or
  • a name and version (lean-agent-light:1.0), or
  • a path to a directory containing a manifest.json.

A model whose download was interrupted is rejected with a message telling you to resume it with lean pull.

Models are versioned independently of each other and of the lean binary, so a new release of one model does not move any other. Version numbers are ours, not the base model’s: a major bump means a new base generation, a minor means a better pack of the same base.

Leaving the version off means latest:

Terminal window
lean pull lean-agent-light # latest
lean pull lean-agent-light:1.0 # exactly 1.0

Versions install side by side under ~/.leanmodels/models/<name>/<version>/, so pulling a new one does not remove the old, and you can pin back to it at any time. lean models lists what you have installed; lean models --remote lists what is published.

A model version can require a minimum lean version - usually because it uses a base architecture that older builds have no runtime support for. The requirement is recorded in the model itself, so it is enforced whether you downloaded it today or a year ago.

lean pull checks it before downloading any weights, and tells you both ways forward:

lean-agent-cruiser:2.0 needs lean 0.3.0 or newer, but this is lean 0.2.1.
This model version uses runtime support that v0.2.1 does not have.
Upgrade lean:
curl -fsSL https://leanmodels.ai/install.sh | sh
Or stay on this version of lean and use an earlier model version:
lean models --remote # lists versions and the lean each needs
lean pull lean-agent-cruiser:<version>

lean models --remote flags any published version your binary cannot load, so you can see what is coming before it blocks you.

Interactive chat REPL with multi-turn history.

Terminal window
lean run lean-agent-light --temperature 0.3
Flag Default Description
--max-tokens <n> 2048 Maximum tokens generated per response.
--system <text> none System prompt for the session.
--temperature <f> 0.7 Sampling temperature. 0 is greedy.
--gpu <n> auto Pin to one GPU by ordinal.
--pipeline off Split layers across all GPUs instead of using one.
--no-think off Skip the <think> phase on models that support it.

In-REPL commands: /stats for cache and throughput numbers, /clear to reset the conversation.

Start an OpenAI-compatible HTTP server. See the API reference for endpoints and payloads.

Terminal window
lean serve lean-agent-light --port 8080
Flag Default Description
--host <addr> 127.0.0.1 Bind address. Use 0.0.0.0 to expose on your network.
--port <n> 8080 Listen port.
--gpu <n> auto Pin to one GPU by ordinal.
--pipeline off Split layers across all GPUs.
--api-key <key> none Require Authorization: Bearer <key> on /v1 routes. Also read from LEAN_API_KEY.
--request-timeout-secs <n> 600 Abort a generation after this long. 0 disables the timeout.

Binding to a non-loopback address without an API key logs a warning: anyone who can reach the port gets unauthenticated inference.

By default the runtime uses one GPU when the model fits on it, because batch-1 decoding gains nothing from pipeline parallelism - the GPUs just run sequentially with a hidden-state hop between them each token. --pipeline forces the split across all cards, which helps when a model’s expert working set is too large to cache well on a single card. --gpu <n> and --pipeline are mutually exclusive.

Download a model from the registry. Resumable, checksum-verified, and it checks free disk before starting.

Terminal window
lean pull lean-agent-light # latest
lean pull lean-agent-light:1.0 # a specific version
Flag Default Description
--registry <url> https://registry.leanmodels.ai Registry to pull from. Also read from LEAN_REGISTRY_URL.
--minimal off Fetch the core plus hot experts only, skipping cold experts.
--output <dir> ~/.leanmodels/models/<name>/<version> Where to write the model.

--minimal produces a working model sooner and on less disk. Cold experts are the ones routing rarely selects, so quality degrades only on the inputs that need them.

Paid models require an API key.

Terminal window
lean login # prompts for the key
lean login --api-key lm_live_...
lean logout # removes stored credentials
Flag Description
--api-key <key> Supply the key non-interactively.
--registry <url> Registry to authenticate against.
--no-verify Store the key without checking it against the registry first.

List locally installed models, or the registry catalog with --remote.

Terminal window
lean models
lean models --remote

Walk every checksum in a model’s manifest and report damaged or missing files. Damaged models are marked so the next lean pull repairs them.

Terminal window
lean verify lean-agent-middle

Replace the binary with the latest published release.

Terminal window
lean update
lean update --check # report whether an update exists, install nothing

It resolves the newest release, downloads the archive matching your platform and build kind — a CUDA install stays a CUDA install — verifies it against the release’s SHA256SUMS, and swaps the binary in place. A missing or mismatched checksum aborts the update; it never installs unverified bytes.

lean --version reports the release tag you are on, so v0.1.0-rc5 and v0.1.0-rc6 are distinguishable rather than both reading as 0.1.0.

If lean lives somewhere you cannot write to, the update says so and points you at the installer, which can be run with whatever permissions that location needs:

Terminal window
curl -sSf https://leanmodels.ai/install.sh | sh

A lean you built from source reports itself as a local build and is left alone unless you pass --force.

Print a model’s manifest: architecture, parameter counts, expert layout, quant, hardware requirements, and a one-line summary of the source weights (see lean provenance for the full chain).

Terminal window
lean info lean-agent-light

Print the source chain for a model: which public repository it was packed from - a community GGUF or the original model checkpoint - the exact commit, and the sha256 of every recorded source file.

Terminal window
lean provenance lean-agent-light

For HuggingFace sources the printed sha256 is the same value the repository’s API serves as each file’s object id, so the claim can be checked without downloading the model. The variant named under variant is the source’s own label (Q4_K_M, UD-Q4_K_XL, or bf16 for packs quantized directly from the original checkpoint); the core/experts lines below it describe this pack’s own tensors, which are not the same thing.

Models packed before provenance was recorded print not recorded rather than a guess.

Detect hardware and recommend models and settings. Takes no arguments. Reports the detected GPUs and RAM, which locally installed models fit in VRAM, RAM, or need NVMe, and the expert cache budget available per GPU.

Terminal window
lean tune

Measure prefill and decode throughput, cache hit rates, and memory use.

Terminal window
lean bench lean-agent-light
Flag Default Description
--gpu <n> auto Pin to one GPU by ordinal.
--pipeline off Split layers across all GPUs.
--ram-budget <mb> auto Cap the host-side expert staging budget.
--timed off Print a per-layer timing breakdown.
--runs <n> 3 Repeat each phase this many times and report median plus min/max.

Lowering --ram-budget simulates a machine with less RAM than the one running the benchmark.

Run a fixed set of diverse prompts to sanity-check output quality after an install, a pull, or a hardware change.

Terminal window
lean smoke lean-agent-light --verbose
Flag Description
--gpu <n> Pin to one GPU by ordinal.
--pipeline Split layers across all GPUs.
--verbose Print the generated output for each prompt.
Variable Effect
LEAN_REGISTRY_URL Default registry for pull, login, and models --remote.
LEAN_API_KEY API key required by lean serve on /v1 routes.
RUST_LOG Log filter. Defaults to info; RUST_LOG=debug for more detail.