Pular para o conteúdo principal

CLI Reference

All commands support --context, --output and --token (JWT bearer token, or NIMBUS_TOKEN env).

--context picks which Nimbus installation to talk to: a name resolves to ~/.nimbus/<name>.json, a path is used as is, and without it the CLI uses NIMBUS_CONFIG (same rules) or else ~/.nimbus/config.json. It used to be called --config, renamed because nimbus config is now a command (see nimbus secret / nimbus config).

nimbus --context prod configure --from prod-config.json # save a context
nimbus --context prod node list # use it
export NIMBUS_CONFIG=prod # or for the whole shell

--output (-o) sets how read commands print: table (the default) for people, or json for scripts. It applies to every list command and to node events, gateway status, manifest status and compute instance-types; describe and task show print JSON already. The JSON is what the API returns, with its field names, so a script reads a field by name instead of by column position, and keeps working when a column is added to the table. An empty list prints [].

nimbus service list -o json | jq -r '.[] | select(.name == "web") | .status'
nimbus node list -o json | jq -r '.[] | select(.ip_address == "10.0.0.5") | .id'

nimbus configure​

Configure CLI credentials. Import from a JSON config file (downloaded from the GUI or bootstrap) or set individual flags.

# Import from JSON (recommended)
nimbus configure --from nimbus-config.json

# Override the API URL (e.g. when connecting from outside the WireGuard mesh)
nimbus configure --from nimbus-config.json --api-url https://<PUBLIC_IP>:8443

# Manual configuration
nimbus configure --api-url URL --access-key KEY --secret-key SECRET

nimbus bootstrap​

Initialize the control plane and create the admin user.

nimbus bootstrap --api-url URL [--insecure]

nimbus version​

Show client and server versions.

nimbus node​

SubcommandDescriptionKey Flags
addAdd a node via SSH or locally--ip (required), --user, --port, --key, --password, --profile, --name, --local, --gpu-driver
agent updateUpdate the agent on a node[node-id-or-name], --all, --ssh, --local, and with --ssh: --profile, --user, --port, --key, --password
os updateUpgrade OS packages on a node[node-id-or-name], --all
os checkRe-check available OS updates[node-id-or-name], --all
stopStop the agent on a node via its task queue[node-id-or-name], --all
startStart the agent on a node over SSH from the control plane[node-id-or-name], --profile, --user, --port, --key, --password, --all
listList all nodes
eventsShow a node's events (deploy steps, join failures, SSH updates)[node-id-or-ip]
describeShow node details[id]
cordonStop scheduling new workloads on a node[id]
drainCordon a node and move its workloads elsewhere[id]
uncordonPut a drained node back into service[id]
deleteDelete a node[id]
set-ipChange a node's IP / pinned interface[node-id], --ip, --interface, --regenerate-cert
gpu-overcommitSet GPU overcommit factor[node-id] [factor] (1, 2, 4, or 8)

There are two ways to get a new agent onto a node, and they differ only in transport, so they are one command with a flag. agent update queues the work on the agent's own task queue: no SSH credentials, but the agent has to be reachable and the node ready. agent update --ssh has the control plane connect over SSH, which also works when the agent is down, and reissues the node's client certificate while it is there. Reach for the default first and fall back to --ssh when an agent is unreachable. An SSH-only flag passed without --ssh is refused rather than ignored.

os update and os check are the operating system's own packages, which have nothing to do with the agent's version — they sat beside it under similar names (os-update, check-updates) and were easy to confuse.

Like add, update runs its SSH connection from the control plane, not from your machine. It can use a stored SSH profile (--profile), and it reaches nodes your machine cannot. A --key file is read on your machine and sent to the control plane. The control plane copies the agent binary and a freshly issued client certificate to a private staging directory on the node, installs them owned by root (the certificate goes to /var/lib/nimbus, where the agent reads it), and restarts the agent. A non-root SSH user needs passwordless sudo. Progress is recorded as node events (ssh_update_started, ssh_update_completed or ssh_update_failed), which update prints as they arrive; it exits non-zero if any node failed. update --local updates the agent on the machine you run it on, without SSH.

os update runs a full OS package upgrade (apt-get upgrade / dnf upgrade) on the node through its agent — it does not use SSH. If the upgrade requires a reboot (for example a new kernel), the node transitions to the needs_reboot state and stops running tasks until an administrator reboots it manually; the agent never reboots on its own. Use --all to queue the upgrade on every ready node.

Because the upgrade can take several minutes, the task reports a running status while it is in progress (visible in the node's task list in the UI and via nimbus task list). When it finishes, the package-manager output is stored on the task result — inspect it with nimbus task show <task-id>.

The upgrade does not run inside the agent. The agent writes a short shell script, starts it as a transient systemd unit (nimbus-os-update) and goes back to its other work. The script runs the upgrade, captures the log, and calls the agent binary back with --finish-os-update to report the result, request a reboot if one is needed, and re-count what is still outstanding.

This matters because the upgrade includes nimbus's own packages: installing them restarts nimbus-agent, and an upgrade running inside the agent's control group would be killed with it, leaving dpkg half-configured and the node offline. Detached, the upgrade survives the restart, and because it reports for itself the task still completes — the agent that started it does not have to be alive at the end. On a control plane the upgrade also restarts nimbus-api, so the callback retries for a few minutes and, failing that, leaves the result in /var/lib/nimbus/pending-results for the agent to deliver on its next start.

Before starting, the agent records which nimbus services are running, and the callback restarts any the upgrade stopped — naming them on the task result. Releases up to 0.7.0 stopped and disabled these services on every upgrade, so a node upgrading from one of them would otherwise be left offline, and off at the next boot too; the callback and the new package's post-install script both repair that. A service an administrator had already stopped is not started, because it was not in the pre-upgrade snapshot.

stop and start are a pair. stop goes through the node's own task queue: the agent reports the node offline and then stops itself, so the control plane shows the stop immediately instead of discovering a stale heartbeat. The machine keeps running and so do its workloads — containers carry on serving, and Nimbus simply stops managing the node. Because the workloads are no longer reported, they show as unreachable while the node is stopped.

start cannot use the task queue: there is no agent left to take a task. The control plane opens an SSH connection itself, like node add and node update, so it can use a stored SSH profile (--profile) and reach nodes your machine cannot; a non-root SSH user needs passwordless sudo. It runs systemctl enable --now nimbus-agent, which also repairs a node whose unit was left disabled and keeps the agent across reboots. Progress is recorded as node events (ssh_start_started, ssh_start_completed, ssh_start_failed), which start prints as they arrive; it exits non-zero if any node failed. The node returns to ready on its next heartbeat.

Taking a node out of service​

Four commands sound similar and do different things:

CommandWhat happens to the machineWhat happens to its workloadsGetting back
cordonKeeps running, agent keeps reportingLeft alone; nothing new is placed herenode uncordon
drainKeeps running, agent keeps reportingMoved to other nodes by the orchestrator that owns themnode uncordon
stopKeeps running; only the agent stopsKeep running, but are reported unreachable because nothing is reporting themnode start (over SSH)
rebootRestartsCome back as the runtime restarts themAutomatic — the node reports ready when it is back
deleteUntouchedLeft as they areRe-add the node

cordon is the gentler of the two: the node stops receiving new work and keeps everything it already has. Use it when you want a node to wind down naturally, or ahead of a window where you would rather it did not pick anything up.

drain is the one for planned maintenance. It marks the node cordoned, so Nimbus places nothing new on it, and asks every orchestrator the node belongs to to drain it: Docker Swarm sets the node's availability to drain and Kubernetes evicts its pods, so the workloads move to other nodes rather than stopping. The order goes to the swarm manager or the cluster's control node, not to the node being drained.

Cluster membership is untouched — the node stays in the swarm or cluster — which is what makes node uncordon enough to undo it. Workloads that moved are not moved back; the orchestrators simply start placing work there again. A node whose workloads belong to no orchestrator keeps running them: there is nowhere for them to move.

Cordoning is independent of a node's status, so a cordoned node still reports ready and can be rebooted, updated or stopped like any other.

stop is the one to reach for when a machine needs to be left alone for a while (hardware work, a noisy neighbour, a node you want out of the cluster without deleting it). Nimbus lets go of the node entirely, and nothing is scheduled on it because it is offline.

set-ip records a node's new address and pins the interface the agent watches for future changes. For the control-plane node, add --regenerate-cert: the API's TLS certificate lists the addresses it was issued for, so after the control plane moves, clients cannot verify it on the new one. The flag reissues that certificate from the existing CA and swaps it in without a restart — agents and already-downloaded connection configs keep working, since the CA is unchanged. See Handling Dynamic Node IPs.

The agent also checks for available OS package updates automatically once per day and reports the counts back to the control plane. nimbus node list shows an UPDATES column (total, with the security subset in parentheses; - means the node has not reported a check yet), and the node detail view shows the same along with when it was last checked. Nothing is installed automatically — use os update when you want to apply them.

nimbus swarm​

SubcommandDescriptionKey Flags
createCreate a swarm group--name (required), --lb, --cloudflare
add-nodeAdd node to swarm--swarm (required), --node (required)
listList swarms
deleteDelete a swarm[id], --force

nimbus compute​

SubcommandDescriptionKey Flags
createCreate an instance--name, --image (required), --swarm or --k8s, --type, --replicas, --port, --domain, --volume, --env, --command, --platform, --gpu
scaleScale replicas--id (required), --replicas (required)
cloneClone an instance--id (required), --name, --target-swarm, --target-k8s
listList instances
describeShow instance details[id]
deleteDelete an instance[id]
stopStop an instance[id]
startStart an instance[id]
logsFetch instance logs[id], -n (lines), -w (follow)
instance-typesList instance types

nimbus k8s​

SubcommandDescriptionKey Flags
createCreate a K3s cluster--name (required), --nodes (required, comma-separated), --ha, --lb
kubeconfigGet kubeconfig--name (required)
add-nodeAdd worker node--cluster (required), --node (required)
remove-nodeRemove worker node--cluster (required), --node (required)
promotePromote a worker to control-plane--cluster (required), --node (required)
demoteDemote a control-plane node back to worker--cluster (required), --node (required)
listList clusters
deleteDelete a cluster[id]

The verbs match nimbus swarm, so the same operation is spelled the same way whichever orchestrator it is on.

--lb deploys EasyHAProxy on the cluster, exactly as swarm create --lb does for a swarm. Without it the cluster has no ingress controller, and a manifest's Ingress objects are not served by anything — a valid choice when ingress is handled elsewhere. One can be added or removed later with nimbus gateway lb set --cluster CLUSTER_ID and nimbus gateway lb delete --cluster CLUSTER_ID.

nimbus s3​

SubcommandDescriptionKey Flags
createDeploy RustFS instance--name (required), --swarm (required), --volume (required), --password, --oidc-cert
listList S3 instances
deleteDelete S3 instance[id]

nimbus volume​

Every volume operation lives here, including attaching. The volume is the noun; where it goes is a flag.

SubcommandDescriptionKey Flags
createCreate a volume (NFS, or local with --local)--name (required), --node (required), --folder (required), --local
listList volumes, or what a target has--cluster, --instance
attachAttach to an instance or a Kubernetes cluster[volume-id], --instance or --cluster (required), --path, --size
detachDetach from an instance or a cluster[volume-id], --instance or --cluster (required)
deleteDelete volume[id]

attach takes exactly one target. --path is the mount path inside an instance and applies only to --instance; --size is the claim size and applies only to --cluster, where it defaults to 10Gi — a claim has to carry a figure, and NFS ignores it.

There is no --swarm, for either verb. A volume is created as a Docker volume on every swarm node when it is created, so a Compose stack can reference nfs-<volume> as external with nothing to attach first. Kubernetes is the exception: the claim is per-cluster and exists only once attached. See Volumes.

list with no flags shows every volume. --cluster shows what is attached to a cluster and the claim each provides; --instance shows what is mounted in an instance and where. Detaching from a cluster waits for the cluster's answer and reports it: Kubernetes will not complete the deletion while a pod still mounts the claim.

nimbus secret / nimbus config​

Values services read by name instead of carrying them in a compose file or manifest. The two commands are the same; a secret's value is stored encrypted and never shown, a config's can be read back. See Secrets and Configs.

SubcommandDescriptionKey Flags
createCreate one on a swarm, or on a cluster in a namespace--name (required), --swarm or --cluster (required), --namespace, --from-file, --from-literal
listList them, with the services using each
describeShow one (never a secret's value)[id-or-name], --value (config only: print the value)
updateReplace the value and redeploy the services using it[id-or-name], --from-file, --from-literal
deleteDelete one no service uses[id-or-name]

Without --from-file or --from-literal, the value is read from stdin.

nimbus service​

A service is a deployed workload bundle: a Docker Compose stack on a swarm, or a Kubernetes manifest on a cluster. Both are managed with the same commands.

SubcommandDescriptionKey Flags
deployDeploy a compose stack, or apply a k8s manifest--file (required), --swarm or --cluster (required), --name, --env, --volume
listList services
describeShow service details[id]
stopStop a service[id]
startStart a service[id]
deleteDelete a service[id]
logsFetch logs from a service[id], -n (lines), -w (follow), --container (one workload)

Service logs​

A service is many workloads — a stack becomes one Docker service per Compose entry, a manifest becomes whatever it declares — and neither docker nor kubectl tails all of them at once. service logs reads each one and tags every line with the workload it came from:

$ nimbus service logs svc-1a2b -n 20
[web] 172.17.0.1 - GET / 200
[web] 172.17.0.1 - GET /api 200
[db] LOG: database system is ready

--container narrows the read to one workload, and its output is passed through untagged, exactly as docker service logs or kubectl logs would print it:

nimbus service logs svc-1a2b --container web

The name to pass is the one in the Compose file (web, not mystack_web) or the workload's name in the manifest; on Kubernetes namespace/name and kind/name also work when a manifest declares the same name twice. Only objects that can have logs are read — a Service, ConfigMap or Ingress is skipped. A workload that cannot be read reports its error in place, so one crash-looping pod does not hide the rest of the stack.

Compose stacks and Kubernetes manifests​

--swarm makes the file a Docker Compose file, deployed as a stack on that swarm. --cluster makes it a Kubernetes manifest, applied with kubectl on the cluster's control node; it may hold several documents separated by ---. A service targets one or the other, never both.

nimbus service list names the difference in the KIND column — compose or manifest — with the swarm or cluster in TARGET:

ID NAME KIND TARGET REPLICAS STATUS
svc-2878de16 vllm-lite compose swarm-3f1d5a37 2 running
svc-91a0c4e2 api manifest cl-7f2b18d0 3 running

In the web UI the same distinction is a badge on each row, and the deploy screen asks which target you mean before it asks for the file.

What Nimbus records. A manifest can declare anything — Deployments, Ingresses, CRDs — so there is no label or name to watch it by. Nimbus records the objects kubectl reports applying, and that record is what everything else works from:

  • Monitoring checks those objects by name. Kinds that run pods (Deployment, StatefulSet, DaemonSet, Job) report ready against desired, so a service is running only once its workloads are ready and pending while they roll out (api has 1/3 ready). Anything else counts as present or missing; an object that has gone missing makes the service error, because Nimbus applied it and something else removed it.
  • Re-applying the same service — deploy again with the same name, or editing it in the UI — applies the new manifest and deletes the objects it no longer declares. Only objects Nimbus recorded applying are ever deleted, which is the safe form of kubectl apply --prune: that one selects by label and is well known for removing more than intended.
  • stop deletes the objects but keeps the manifest, so start re-applies it. remove deletes the objects and the service record.

Using an NFS volume. The two backends reference a Nimbus volume differently, and the web UI's deploy screen can insert either snippet at the cursor for any existing volume.

On a swarm, Nimbus creates a Docker volume on the node named nfs-<volume>, so a compose file refers to it as external. A compose file is a single YAML document, and Docker ignores top-level x- keys, so an anchor can live there and be aliased by every service that wants the mount:

volumes:
nfs-media:
external: true

x-nfs-media: &nfs-media
- nfs-media:/data/media

services:
web:
image: nginx
volumes: *nfs-media

On a cluster, attaching a volume (POST /v1/kubernetes/clusters/{id}/volumes) applies a PersistentVolume and a claim named nfs-<volume>. Reuse there is by claim name: every workload naming the same claim gets the same share, across documents and across services. YAML anchors cannot help across documents — each document has its own anchor namespace, so an anchor defined before a --- is not visible after it — but they still save repetition within one document:

spec:
template:
spec:
containers:
- name: app
volumeMounts:
- &mnt-media
name: media
mountPath: /data/media
- name: sidecar
volumeMounts:
- *mnt-media
volumes:
- name: media
persistentVolumeClaim:
claimName: nfs-media

A volume gains its claim when it is attached to the cluster — nimbus volume attach <volume-id> --cluster <id>, or the NFS Volumes card on the cluster page. nimbus volume list --cluster <id> shows what is attached.

Mounting a claim that was never attached is the quiet failure worth knowing about: kubectl apply succeeds, and the pods then sit unschedulable waiting for a claim that does not exist. Nimbus reports Kubernetes' own reason on the service (Deployment/api in default has 0/2 ready: persistentvolumeclaim "nfs-test" not found), and the web UI offers Attach and add so the claim is created before the manifest names it. On a swarm the equivalent failure is loud: the deploy itself fails with Docker's error.

describe prints the manifest and the recorded objects, and the UI lists them under Applied objects — which is what is really in the cluster, as opposed to what the manifest asks for.

nimbus manifest​

SubcommandDescriptionKey Flags
applyProvision from manifest--file (required), --env, --prune
removeRemove the resources a manifest declares[name] or --file, --env, --purge
listList saved manifests
statusShow a saved manifest's status[name]

remove keeps the manifest record, so the same manifest can be applied again. --purge deletes the record as well, once its resources are gone. There used to be a separate manifest delete that did the second thing, which left two verbs for the same act where one was quietly more destructive.

nimbus ssh-profile​

Manage SSH profiles — named, reusable SSH credential sets stored encrypted in the database. Sensitive data (private keys, passwords) is encrypted using the control plane's CA key.

SubcommandDescriptionKey Flags
createCreate an SSH profile--name (required), --user, --port, --key, --password
listList SSH profiles
deleteDelete an SSH profile[id]

At least one of --key (path to private key file) or --password is required. The --key flag reads the file and stores its content encrypted — the original file is not needed afterward.

# Create a profile with an SSH key
nimbus ssh-profile create --name prod-servers --user deploy --key ~/.ssh/id_ed25519

# Create a profile with password auth
nimbus ssh-profile create --name staging --user root --password secret123

# List profiles (sensitive data is never shown)
nimbus ssh-profile list

# Delete a profile
nimbus ssh-profile delete sshp-a1b2c3d4

SSH profiles are referenced in manifests via ssh.profile:

nodes:
web1:
ip: 192.168.1.10
ssh:
profile: prod-servers

nimbus certificate​

Manage TLS certificates for SNI-based multi-domain serving. Certificates are stored in the database and served immediately — no restart required. The default WireGuard IP cert (10.106.103.1) is built-in and cannot be deleted.

SubcommandDescriptionKey Flags
listList certificates
createCreate a TLS certificate--domain (required), --cert (required), --key (required), --ca
deleteDelete a certificate[id]
# Add a Let's Encrypt certificate for a public domain
nimbus certificate create \
--domain nimbus.example.com \
--cert /etc/letsencrypt/live/nimbus.example.com/fullchain.pem \
--key /etc/letsencrypt/live/nimbus.example.com/privkey.pem

# Add a certificate with a custom CA (for internal PKI)
nimbus certificate create \
--domain internal.example.com \
--cert /path/to/cert.pem \
--key /path/to/key.pem \
--ca /path/to/ca.pem

# List all certificates
nimbus certificate list

# Delete a certificate
nimbus certificate delete cert-a1b2c3d4

See Exposing OIDC/OAuth2 Publicly for a full setup guide.

nimbus dns​

SubcommandDescriptionKey Flags
createConfigure local DNS resolution--swarm-id
deleteRemove DNS configuration--swarm-id

nimbus gateway​

SubcommandDescriptionKey Flags
statusShow the routing table
lb setDeploy the load balancer--swarm or --cluster
lb deleteDelete the load balancer--swarm or --cluster
lb listList load balancers

The load balancer (EasyHAProxy) is what the gateway routes through, so it is managed here rather than under nimbus swarm — nimbus swarm create --lb still deploys one at creation time.

lb list shows every load balancer, on swarms and clusters alike, with the one it serves:

$ nimbus gateway lb list
ID NAME SERVES NODE PORT STATUS
lb-edge-9f2a easyhaproxy-sw-1 swarm-1a2b node-1 80 active
lb-k8s-3c71 easyhaproxy-k8s-p1 cls-ccf2d1c4 node-4 31080 active

Both backends work the same way: a load balancer exists when it was asked for — swarm create --lb, k8s create --lb, or lb set later — and lb delete takes it away. Removing one leaves the workloads running; what stops is the ingress in front of them.

nimbus iam​

SubcommandDescriptionKey Flags
get-tokenGet JWT token (HMAC auth)
loginLogin with password--email (required), --password (required)

login and get-token act on whoever runs the command rather than on a resource, so they stay at the top. Everything else is grouped by the thing it manages.

nimbus iam user​

SubcommandDescriptionKey Flags
createCreate a user--email (required), --name, --admin
listList users
updateUpdate email or name<user-id> (arg), --email, --name
deleteDelete a user<user-id> (arg)
set-passwordSet user password--user-id (required), --password (required)

nimbus iam key​

SubcommandDescriptionKey Flags
createGenerate API key pair--user-id (required)
listList a user's keys--user-id (required)
revokeRevoke an API key--key-id (required)

nimbus iam group​

SubcommandDescriptionKey Flags
listList all groups
createCreate a group--name (required), --description, --scope (repeatable)
deleteDelete a custom group<group-id> (arg)
set-scopesReplace a group's scope list--group (required), --scope (repeatable)

nimbus iam scope​

SubcommandDescriptionKey Flags
listList all registered ARN scopes
registerRegister a new scope--scope (required), --description
unregisterRemove a registered scope<scope> (arg)

nimbus iam client​

Manage OAuth2 clients used by external services (Grafana, ArgoCD, Nextcloud, etc.) to authenticate via Nimbus SSO.

SubcommandDescriptionKey Flags
listList all OAuth2 clients
createCreate an OAuth2 client--name (required), --redirect-uri (required, repeatable)
deleteDelete an OAuth2 client<id> (arg — use the ID column from list, not Client ID)

The create command prints the Client ID and Client Secret. The secret is shown only once — save it immediately.

# Register a client for Grafana
nimbus iam client create \
--name grafana \
--redirect-uri "https://grafana.example.com/login/generic_oauth"

# Register a client with multiple redirect URIs
nimbus iam client create \
--name argocd \
--redirect-uri "https://argocd.example.com/auth/callback" \
--redirect-uri "https://argocd-staging.example.com/auth/callback"

# List all clients
nimbus iam client list

# Delete a client (use the ID column, not the client_id)
nimbus iam client delete oac-abc123

nimbus iam user-group​

SubcommandDescriptionKey Flags
listList a user's groups--user (required)
addAdd user to a group--user (required), --group (required)
removeRemove user from a group--user (required), --group (required)

nimbus cleanup​

Force-clean a resource stuck in error state.

nimbus cleanup [resource-type] [resource-id]

Valid resource types: instance, cluster, s3, loadbalancer, swarm, volume.