Documentation / CLI reference

CLI — global flags & core commands

Run pyrusllm --help to list every command, or pyrusllm <command> --help for one command's full detail.

Global flags

pyrusllm [flags] [command]
FlagDescription
--version, -vPrint the current version
--storage <dir>Custom storage directory
--no-updatesDisable OTA updates for this run
--port <n>App HTTP port (default 8787)
--no-serveOnly the OTA updater, without starting the app
--no-openDo not open the browser on start
--update-delay <ms>OTA jitter window in ms (default 10000)
One process per storage directory. The Corestore takes a lock on it. A second node against the same --storage fails with "File descriptor could not be locked" — use a different directory, or --no-store.

pyrusllm prompt

Answer a prompt with 100% local inference. Loads an LLM with QVAC and answers without leaving this machine.

pyrusllm prompt "<prompt>" [flags]
Argument / flagDescription
<prompt>The text to answer, or - to read it from stdin
--model <alias>Model alias or exact registry name (default llama1b)
--ctx <n>Context size
--gpu-layers <n>Layers to send to the GPU. 0 = all CPU
--no-downloadFail instead of downloading the weights when not cached
--quiet, -qPrint only the answer, no diagnostics or measurements
Use - on Windows for non-ASCII prompts. The standalone binary receives argv in the ANSI codepage, which mangles accented characters. stdin is a byte stream and does not go through that conversion: echo "Respondé…" | pyrusllm prompt -
On a weak iGPU (Intel UHD 620), --gpu-layers 0 measured 5× faster than letting the SDK decide. See NOTES.md.

pyrusllm serve

Start the OpenAI-compatible gateway and the browser panels. This is the command that turns on the HTTP API.

pyrusllm serve [flags]
FlagDescription
--port <n>Gateway HTTP port (default 8787)
--swarmJoin the P2P topic and populate the registry with verified peers
--operator <name>Operator name announced in the manifest
--demoPopulate the registry with simulated nodes (1 real + 3 mocks)
--no-storeDo not open the Hyperbee/Hyperdrive: no persistence, no files
--model <alias>Model this node serves — the one advertised in the manifest AND the one the engine loads, a single source
--ctx <n>Model's context window. Prompt + reasoning + answer together: a "thinking" model with too small a window runs out of room before answering
--gpu-layers <n>Layers to send to the real node's GPU. 0 = all CPU
--log-inferenceLog TTFT, bytes and chunks every 5s per generation. Without it, an answer that takes minutes looks exactly like a hung process
Without --swarm the gateway serves the panels and the API but stays off the network: the registry starts empty and requests for a model nobody serves return a clear "no nodes serving that model" error. That is the honest state, not a bug.

pyrusllm peers

Join the P2P topic and list the peers with a verified manifest, without starting the gateway. Reports join → first peer and join → first verified manifest timings.

pyrusllm peers [flags]
FlagDescription
--operator <name>Operator name announced in the manifest
--timeout <s>Exit after N seconds (default: never exits, Ctrl+C)
--expect <n>Exit code 1 if fewer than N verified peers on exit
Discovery takes ~17 s. Run this on both machines with overlapping windows — a short --timeout on one side while the other has not joined yet reports 0 peers and looks like a network fault when it is not.