Documentation / CLI reference
CLI — global flags & core commands
Run pyrusllm --help to list every command, or
pyrusllm <command> --help for one command's full detail.
Global flags
pyrusllm [flags] [command]
| Flag | Description |
|---|---|
--version, -v | Print the current version |
--storage <dir> | Custom storage directory |
--no-updates | Disable OTA updates for this run |
--port <n> | App HTTP port (default 8787) |
--no-serve | Only the OTA updater, without starting the app |
--no-open | Do not open the browser on start |
--update-delay <ms> | OTA jitter window in ms (default 10000) |
One process per storage directory. The Corestore takes a
lock on it. A second node against the same
--storage fails with
"File descriptor could not be locked" — use a different directory, or
--no-store.pyrusllm prompt
Answer a prompt with 100% local inference. Loads an LLM with QVAC and answers without leaving this machine.
pyrusllm prompt "<prompt>" [flags]
| Argument / flag | Description |
|---|---|
<prompt> | The text to answer, or - to read it from stdin |
--model <alias> | Model alias or exact registry name (default llama1b) |
--ctx <n> | Context size |
--gpu-layers <n> | Layers to send to the GPU. 0 = all CPU |
--no-download | Fail instead of downloading the weights when not cached |
--quiet, -q | Print only the answer, no diagnostics or measurements |
Use
- on Windows for non-ASCII prompts. The
standalone binary receives argv in the ANSI codepage, which mangles accented
characters. stdin is a byte stream and does not go through that conversion:
echo "Respondé…" | pyrusllm prompt -On a weak iGPU (Intel UHD 620),
--gpu-layers 0 measured
5× faster than letting the SDK decide. See NOTES.md.pyrusllm serve
Start the OpenAI-compatible gateway and the browser panels. This is the command that turns on the HTTP API.
pyrusllm serve [flags]
| Flag | Description |
|---|---|
--port <n> | Gateway HTTP port (default 8787) |
--swarm | Join the P2P topic and populate the registry with verified peers |
--operator <name> | Operator name announced in the manifest |
--demo | Populate the registry with simulated nodes (1 real + 3 mocks) |
--no-store | Do not open the Hyperbee/Hyperdrive: no persistence, no files |
--model <alias> | Model this node serves — the one advertised in the manifest AND the one the engine loads, a single source |
--ctx <n> | Model's context window. Prompt + reasoning + answer together: a "thinking" model with too small a window runs out of room before answering |
--gpu-layers <n> | Layers to send to the real node's GPU. 0 = all CPU |
--log-inference | Log TTFT, bytes and chunks every 5s per generation. Without it, an answer that takes minutes looks exactly like a hung process |
Without
--swarm the gateway serves the panels and the
API but stays off the network: the registry starts empty and requests
for a model nobody serves return a clear "no nodes serving that model" error. That is
the honest state, not a bug.pyrusllm peers
Join the P2P topic and list the peers with a verified manifest, without starting the gateway. Reports join → first peer and join → first verified manifest timings.
pyrusllm peers [flags]
| Flag | Description |
|---|---|
--operator <name> | Operator name announced in the manifest |
--timeout <s> | Exit after N seconds (default: never exits, Ctrl+C) |
--expect <n> | Exit code 1 if fewer than N verified peers on exit |
Discovery takes ~17 s. Run this on both machines with
overlapping windows — a short
--timeout on one side while
the other has not joined yet reports 0 peers and looks like a network fault when it is
not.