Skip to content

AI models ​

A model is a first-class resource: an inference engine running as a container on a GPU node, a persistent volume holding the downloaded weights, and an OpenAI-compatible endpoint protected by an API key. It is available from the dashboard (AI models), the CLI (levelrail models), the REST API (/api/v1/models) and the MCP server.

Version 1 supports NVIDIA GPUs on Linux nodes (Ubuntu is the documented path). AMD, Intel and Apple GPUs are not supported.

GPU nodes ​

Every node reports its NVIDIA GPUs to the control plane: driver version, per-GPU VRAM total and used, utilization, and whether Docker has the nvidia container runtime registered. Detection runs nvidia-smi on the host and asks the Docker Engine API for its runtimes. The control plane detects its own host every minute (APP_GPU_COLLECT_INTERVAL), agents report on connect and every minute after.

The AI models page shows a card per GPU node with VRAM used and total. levelrail models gpus prints the same data.

Prerequisites on the node (Ubuntu) ​

  1. Install the NVIDIA driver: sudo ubuntu-drivers install, then reboot. nvidia-smi must work.
  2. Add NVIDIA's apt repository for the container toolkit, then install and register it:
sh
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

When a node has a GPU but Docker has no nvidia runtime, the dashboard card shows "nvidia runtime missing" with this fix, levelrail doctor and the Status page raise a warning, and models placed there stay in the GPURuntimeMissing state until it is fixed. Run levelrail models restart <name> afterwards.

If the agent runs inside a container, nvidia-smi must be reachable from it, or run the agent directly on the host.

GPU attach: CDI and legacy ​

Containers get GPUs one of two ways. Levelrail prefers CDI (the Container Device Interface) when the Docker daemon on the node lists NVIDIA CDI devices (nvidia.com/gpu=0, nvidia.com/gpu=all), and otherwise uses the legacy nvidia runtime with a device request. The choice is made per container at create time on the node that runs it, so mixed fleets work. A node with CDI devices needs no nvidia runtime registered and is treated as usable.

To use CDI, generate a spec on the node and make sure Docker discovers it (Docker 28.3 or newer does by default; older Docker needs {"features": {"cdi": true}} in /etc/docker/daemon.json):

sh
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
sudo systemctl restart docker
docker info | grep -A3 "Discovered Devices"

A request that CDI cannot serve exactly (for example more GPUs than the spec lists) falls back to the legacy runtime instead of failing. With a count, CDI attaches the first N GPUs by index. Set APP_GPU_ATTACH=legacy on a node to never use CDI (the default is auto). Regenerate the spec after a driver update or GPU change.

levelrail doctor (and the Doctor page) checks the control plane host in detail: the driver, per-GPU memory (a warning above APP_GPU_DOCTOR_VRAM_WARN_PERCENT, default 90), the NVIDIA container toolkit version, whether a CDI spec exists and Docker lists its devices, whether the nvidia runtime is registered, and the resulting attach mode, with the exact fix command for each problem. Remote nodes report only driver, runtime or CDI availability and per-GPU memory; the toolkit and spec files of a remote node are not inspected. None of this was exercised against a real GPU: the detection and the CDI request are covered with fakes.

Deploying a model ​

sh
levelrail models deploy --name chat --engine ollama --model llama3.1:8b
levelrail models deploy --name llama --engine vllm --model meta-llama/Llama-3.1-8B-Instruct \
  --gpus 2 --context 8192 --hf-token-from-env --node <node-id> --domain llm.example.com
OptionMeaning
--engineollama, vllm or llamacpp.
--modelAn Ollama tag (llama3.1:8b), a HuggingFace repo (org/name) for vLLM, or a GGUF repo for llama.cpp, optionally with a quant suffix (org/name-GGUF:Q4_K_M).
--nodeNode ID. Default is the control plane's own host.
--gpus, --gpu-devicesA count, all, or specific GPU indexes or UUIDs.
--contextMaximum context length in tokens.
--quantizationvLLM only (for example awq). Ollama encodes it in the tag, llama.cpp in the :quant suffix.
--domainHostname for the endpoint. Without one, a zero-config hostname is used when APP_PUBLIC_HOST is a public IP.
--hf-token-from-envReads HF_TOKEN from your environment for gated models.

The engine image can be overridden per engine with APP_MODEL_IMAGE_OLLAMA, APP_MODEL_IMAGE_VLLM and APP_MODEL_IMAGE_LLAMACPP.

HuggingFace token ​

The token is stored through the envelope-encryption path (internal/secrets) under the model's own namespace and injected into the engine container at create time. It is never returned by the API and never logged. It requires a master key (APP_MASTER_KEY). Rotate it with PUT /api/v1/models/{name}/hf-token, which restarts the model.

Hugging Face preflight ​

Before deploying a Hugging Face repository, check it:

sh
levelrail models preflight bartowski/Llama-3.2-3B-Instruct-GGUF --engine llamacpp --node <node-id>
HF_TOKEN=hf_... levelrail models preflight meta-llama/Llama-3.1-8B-Instruct --engine vllm --hf-token-from-env

The check queries the public Hub API (through the same outbound guard as webhooks, with a timeout and a response size cap) and reports:

  • whether the repository exists, and whether it is gated. A gated repository says what to do next: accept the license on huggingface.co and provide an access token.
  • the license, total download size and per-file sizes.
  • the GGUF quantizations it contains, with sizes, and a recommended one for the target node.
  • whether the node has enough free disk for the download, and which engine can load the weights (GGUF for llama.cpp and Ollama, safetensors for vLLM).

In the dashboard the same panel appears under the model field of the deploy dialog and updates as you type. Picking a quantization there fills in the :quant suffix.

Fit figures are estimates. A quantization is judged against the node's free VRAM as its file size plus a flat overhead for the KV cache and buffers. Real usage depends on context length, batch size, engine and driver. Fit is shown only when the node has reported its GPU, and disk only when it has reported its disk; otherwise the answer is "unknown", never a guess. System RAM is not considered.

The token you supply is used for that one request. It is not stored or logged, and it is never part of a cache key.

VariableDefaultMeaning
APP_HF_CACHE_TTL5mHow long Hub answers are reused. 0 disables the cache.
APP_HF_TIMEOUT10sPer-request timeout to the Hub.
APP_HF_MAX_RESPONSE_BYTES16777216Cap on the repository listing.
APP_HF_MAX_LISTED_FILES100Files returned per repository (totals always cover all files).
APP_HF_BASE_URLhttps://huggingface.coHub base URL (a mirror).
APP_MODEL_FIT_OVERHEAD_PERCENT20Added to weight size for KV cache and buffers.
APP_MODEL_FIT_PERCENT90Share of free VRAM a model may use to count as "fits". Above it, up to 100%, is "tight".
APP_MODEL_DISK_HEADROOM_PERCENT10Free disk required beyond the download.

The disk check measures the filesystem holding Docker's volumes on the control plane host. If that path is not visible from the control plane, or the node is remote, disk is reported as unknown.

The API is POST /api/v1/models/preflight (repo, optional engine, quant, file, gpu_count, node_id, hf_token) and the MCP tool is preflight_model. Hub problems (missing, gated, rate limited, unreachable) come back as a status in a normal response so a client can always show the next step.

VRAM fit check ​

Before you deploy, and on the model page, Levelrail estimates whether a model fits in each GPU node's free VRAM and rates every node: fits, tight, wont_fit or unknown. The pill has an info tip with the arithmetic, for example weights 4.7 GiB + KV 1.2 GiB + overhead 0.8 GiB = 6.7 GiB of 8.0 GiB free.

This is an estimate, not a guarantee. The pieces:

  • Weights. Exact when the Hugging Face file sizes are known (pass weights_bytes from a preflight), otherwise estimated from the parameter count in the model name (8b, 8x7b) and the quantization (a tag suffix such as q4_K_M, a :quant suffix, or vLLM's --quantization). A name with no size, such as mistral, is unknown, never guessed. Ollama and llama.cpp default to about 4.85 bits per weight, vLLM to 16 (bf16).
  • KV cache. Approximated from the parameter count and the context length. With no context length set, APP_MODEL_FIT_DEFAULT_CONTEXT tokens (default 8192) are assumed and the tip says so.
  • Overhead. A flat APP_MODEL_FIT_GPU_OVERHEAD_MIB per GPU (default 512), plus APP_MODEL_FIT_VLLM_EXTRA_MIB (default 1024) for vLLM.
  • Free VRAM. From the node's GPU report (refreshed every minute). VRAM estimated for other models placed on the node that are not loaded yet is subtracted, so two deploys in the same minute do not both see the same free memory. On the model page, a loaded model's own VRAM is added back.
  • Verdict. fits when the total is at most APP_MODEL_FIT_PERCENT of free VRAM (default 90), tight up to 100%, otherwise wont_fit. vLLM also needs each GPU to have APP_MODEL_VLLM_GPU_UTILIZATION (default 0.9) of its memory free, because it pre-allocates that much.

Nodes are ranked best first, with non-schedulable nodes last. The check is advisory: it never blocks a deploy. APP_MODEL_FIT_KV_KIB_PER_TOKEN overrides the KV heuristic if your model family differs.

  • Dashboard. The deploy dialog shows a verdict per node and highlights the selected one. The model page Overview shows the same for the deployed model, with its current node marked.
  • CLI. levelrail models fit --engine ollama --model llama3.1:8b --context 8192, or levelrail models fit --name chat for a deployed model.
  • API. POST /api/v1/models/fit and GET /api/v1/models/{name}/fit (read ability).
  • MCP. check_model_fit (read-only).

No GPU was involved in testing this; the numbers come from the GPU facts nodes report and have not been compared against real engine memory use.

Model cache ​

Each model keeps its downloaded weights in its own Docker volume. Deleting a model keeps the volume, so those volumes are what accumulates.

sh
levelrail models cache list
levelrail models cache prune --dry-run
levelrail models cache prune

The list shows each volume with its size, its owning model, when it was last used, and whether it is safe to prune. Last use comes from gateway traffic when the model has any, otherwise from the model's last change (or, for a volume no model owns, the volume's creation time). Two models that point at the same weights are flagged and counted once in the unique total.

Prune removes only volumes that no model owns, that no container mounts, and that have been unused for APP_MODEL_CACHE_UNUSED_DAYS days (default 30). Weights of any deployed model, running or not, are never removed. The server rechecks each volume right before removing it. In the dashboard, the Model cache section shows a dry run in a confirmation dialog first. Pruning needs the root ability.

Only the control plane's own host is inspected today; remote nodes are listed as not yet supported. The MCP tool list_model_cache is read-only; the API is GET /api/v1/model-cache and POST /api/v1/model-cache/prune.

Lifecycle and status ​

A model has one Ready condition whose reason tells you where it is:

ReasonMeaning
StartingThe container is being created or started.
DownloadingThe weights are downloading. Ollama reports percent and bytes; vLLM and llama.cpp download on start, follow levelrail models logs <name> --follow.
LoadingThe model is being loaded into VRAM.
ModelLoadedThe model is loaded and serving. This is the readiness signal.
IdleAn on-demand model whose engine was stopped after its idle time. It starts on the first request (see On-demand residency).
WakingUpAn on-demand model's engine is starting or loading after a wake.
WaitingForGPUAn on-demand model would fit the GPUs but other workloads hold too much VRAM right now; it starts when they free it.
NoGPUOnNode, GPURuntimeMissing, InsufficientGPUsThe node cannot serve the request. See GPU nodes above.
HFTokenUnavailableA token is set but cannot be decrypted (no master key).
EndpointUnreachableThe node is remote and not on the WireGuard mesh, so the control plane cannot reach the engine.
DownloadFailedThe engine reported a download error (retried after 30 seconds).

The downloaded weights live in a persistent Docker volume, so a restart or redeploy does not download again. Deleting a model removes its container but keeps the volume.

The model page ​

Each model has its own page at /models/<name> (click its name in the list). Tabs:

  • Overview. Status, engine, node, GPU, context length, quantization, the endpoint URL with a copy button, and a ready-made curl example.
  • Keys. The named keys with their limits, expiry and rotation (see Virtual keys below).
  • Usage. Requests, tokens, errors and first-byte latency per key.
  • Logs. The engine's live log tail.

The tab is kept in the URL (?tab=keys), so a link opens straight to it. From the CLI, levelrail models get <name> prints the same overview facts.

Engine metrics ​

The Overview tab shows the inference engine's own metrics with a health summary. The control plane scrapes each running engine every APP_MODEL_ENGINE_METRICS_INTERVAL (default 15s) over the same path its readiness probe uses, and stores the readings in the telemetry database under model:<name>.

MetricvLLMllama.cppOllama
KV cache usageyesyesnot available
Queued and running requestsyesyesnot available
Prefix cache hit rateyesnot availablenot available
Tokens per secondyesyesnot available
Time to first tokenyesnot availablenot available
VRAM in use, share running on CPUnot availablenot availableyes (from /api/ps)

Rates (tokens per second, prefix hit rate, time to first token) are computed between two scrapes, so they appear one interval after the model starts and skip an engine restart. llama.cpp serves /metrics only when started with --metrics, which new containers now get; a model deployed earlier needs levelrail models restart <name>. Ollama exposes no Prometheus metrics, so the page says "Not available for ollama" instead of guessing; per-key tokens and first-byte latency are on the Usage tab.

The health pill warns when the KV cache is at or above APP_MODEL_KV_WARN_PERCENT (default 90), when more than APP_MODEL_QUEUE_WARN requests are queued (default 4), or when an Ollama model runs partly on CPU.

  • CLI. levelrail models metrics <name> --since 6h.
  • API. GET /api/v1/models/{name}/engine-metrics?since=1h (read ability).
  • MCP. get_model_engine_metrics (read-only).

The scrape is done by the control plane, not by a node agent, so a remote node's engine must be reachable over the mesh (the same requirement as serving it). Numbers depend on the engine version's metric names; vLLM's older gpu_cache_usage_perc and current kv_cache_usage_perc are both read.

On-demand residency ​

By default a model is always loaded. On a shared GPU you can make it on demand: after an idle period the engine container is stopped, freeing its VRAM, and the first request starts it again.

sh
levelrail models deploy --name chat --engine ollama --model llama3.1:8b --residency on_demand --idle-ttl 20m
levelrail models residency chat --mode on_demand --idle-ttl 30m
levelrail models wake chat      # start now
levelrail models sleep chat     # stop now

How it behaves:

  • Idle. Every authenticated gateway request marks the model used (a long streaming response keeps refreshing it). When nothing has used it for the idle TTL, the reconciler stops the container and the status reason becomes Idle. The TTL is per model (--idle-ttl), defaulting to APP_MODEL_IDLE_TTL (15m). It cannot be set below APP_MODEL_MIN_IDLE_TTL (1m), so a busy model is never judged idle between activity writes (APP_MODEL_ACTIVITY_TOUCH_INTERVAL, default 10s).
  • Wake. The first authenticated request to an idle model asks the reconciler to start it and is held while the engine loads, for at most APP_MODEL_WAKE_WAIT (default 90s). If the engine is not serving by then, or the wait is 0, the client gets 503 with Retry-After (APP_MODEL_WAKE_RETRY_AFTER, default 15s) and error code model_waking; the wake keeps going, so a retry usually succeeds. Unauthenticated or invalid requests never wake a model.
  • Ollama. Ollama's own keep-alive is set to forever so weights stay in VRAM; unloading it means stopping the container, which is what on-demand does.
  • GPU contention. If a wake would fit an empty GPU set but other workloads hold too much VRAM (per the fit estimate above), the model reports WaitingForGPU and starts when memory frees. A model too large for the GPUs is not held back, since waiting would not help. The GPU snapshot refreshes every minute, so this can lag.
  • State. The reconciler owns the asleep, waking and awake states. Stop and start are idempotent: a stop that succeeded but failed to record is repaired on the next pass, and a request that lands while the reconciler is deciding to stop wins (the row is re-read right before stopping).
  • Dashboard. The model page Overview has a Residency card (mode, idle minutes, Wake now, Sleep now) and the deploy dialog has a Residency field. API. PUT /api/v1/models/{name}/residency, POST /api/v1/models/{name}/wake, POST /api/v1/models/{name}/sleep (write ability); the model resource carries residency, idle_ttl_seconds, effective_idle_ttl_seconds, residency_state and last_active_at. MCP. get_model returns the same fields and deploy_model accepts residency and idle_ttl_seconds.

Swap groups (several models sharing one GPU, only one loaded at a time) are not built; give each model on a shared GPU an idle TTL instead. Cold start includes the engine's load time (minutes for a large vLLM model), so on-demand suits models used in bursts, not latency-critical ones. No GPU was involved in testing this; the engine start, stop and wake paths are covered with fakes.

The endpoint and API key ​

The model's hostname is routed through the built-in ingress to the control plane, whose gateway checks the API key and proxies to the engine. It speaks the OpenAI API:

sh
curl https://llm.example.com/v1/chat/completions \
  -H "Authorization: Bearer lr-..." -H "Content-Type: application/json" \
  -d '{"model": "llama3.1:8b", "messages": [{"role": "user", "content": "hello"}]}'
  • The key is generated at deploy time and shown once. Only its SHA-256 hash is stored. Lost it? levelrail models rotate-key <name> replaces the default key at once. For more than one key, see Virtual keys.
  • Only an allowlist of OpenAI-compatible routes is served (see below). Engine admin APIs (Ollama's pull and delete, vLLM's runtime LoRA load and unload, llama.cpp's /props and /lora-adapters) are never exposed, even ones that live under /v1/.
  • The engine's port is published on loopback (local node) or the WireGuard mesh address (remote node), never on a public interface.

Served routes ​

Matching is exact and case sensitive: trailing slashes, //, .., backslashes and encoded slashes (%2F) are rejected. Anything else returns a 404 OpenAI-style error before the API key is checked, and a wrong method returns 405 with an Allow header.

RouteMethodOllamavLLMllama.cpp
/v1/models, /v1/models/{id}GETyesyesyes
/v1/chat/completionsPOSTyesyesyes
/v1/completionsPOSTyesyesyes
/v1/embeddingsPOSTyesyesyes
/v1/responsesPOSTyesyesyes
/v1/audio/transcriptions, /v1/audio/translationsPOSTnoyesno

Not served on purpose: vLLM's /v1/load_lora_adapter, /v1/unload_lora_adapter, /v1/chat/completions/batch and the /render routes; llama.cpp's /v1/chat/completions/control and its Anthropic-style /v1/messages; stateful GET and cancel on /v1/responses/{id}.

Limits ​

Limits are global (every model) and set with environment variables on the control plane. 0 or a negative value disables a limit. The current values show on GET /api/v1/models/{name} (limits) and in levelrail models get.

VariableDefaultEffect
APP_MODEL_GATEWAY_MAX_BODY_BYTES33554432 (32 MiB)Larger request bodies get 413.
APP_MODEL_GATEWAY_MAX_N16n or best_of above this gets 400.
APP_MODEL_GATEWAY_MAX_GEN_LEN32768max_tokens, max_completion_tokens or max_output_tokens above this, or negative (unlimited on llama.cpp), gets 400. A request that sets none is passed through; the engine's context length bounds it.
APP_MODEL_GATEWAY_MAX_INFLIGHT32Concurrent requests per model; more get 429 with Retry-After. Streams hold a slot until they end.
APP_MODEL_GATEWAY_RETRY_AFTER5sThe Retry-After value.
APP_MODEL_GATEWAY_DIAL_TIMEOUT5sConnecting to the engine.
APP_MODEL_GATEWAY_HEADER_TIMEOUT5mWaiting for the engine's response headers. A non-streaming completion sends none until it finishes, so keep this above your longest generation. Exceeded: 504.
APP_MODEL_GATEWAY_IDLE_TIMEOUT2mLongest gap between engine output on a response. It is a gap, not a total, so long streams that keep producing tokens are never cut.

JSON bodies are read up to the size cap so n and the token limits can be checked; the body is then forwarded unchanged. Audio uploads are size-capped but not parsed. Request and response bodies are never logged: each request logs only method, model, status, duration and bytes.

Virtual keys and usage ​

A model can have several named keys, so each client gets its own identity, limits and usage. The key made at deploy time is the key named default; existing models were migrated to it, and it keeps working unchanged.

bash
levelrail models keys create chat --name ci --rpm 60 --tpm 100000 --tpd 2000000 --max-parallel 4 \
  --allow-paths /v1/chat/completions --expires-in 720h
levelrail models keys list chat
levelrail models keys rotate chat <key-id> --grace 30m
levelrail models keys revoke chat <key-id>
levelrail models usage chat --since 168h

In the dashboard, the gauge button on a model row opens the keys panel (create, rotate, revoke, last used, limits) and the usage card (requests, tokens, errors and time to first byte over time, plus a per key table).

  • Only the SHA-256 of a key is stored. The key is shown once, with its first 8 characters kept as a prefix for identification. last used is updated on each flush interval.
  • Rotation issues a replacement with the same name, limits and expiry. The old key keeps working for a grace window (--grace, default APP_MODEL_KEY_ROTATION_GRACE, 1h; at most APP_MODEL_KEY_MAX_GRACE, 168h; 0 retires it at once), then stops. Revoking stops a key immediately.
  • Limits are per key. rpm and max parallel are enforced at the gateway: a request over either gets 429 with Retry-After and the same generic OpenAI-style error, so the response never says which limit tripped. tpm and tpd (tokens per day, rolling 24 hours) are soft: tokens are counted from responses after they finish, so the request that crosses the limit still completes and later ones get 429 until the window rolls over. An unknown, revoked or expired key always gets 401.
  • Allow lists: allow_paths are exact gateway paths (a trailing / allows everything below it) and must be routes the engine's allowlist already serves, so a key can never reach an engine admin route. allow_models are compared with the model field of the request body; a request without one is refused when the list is set.
  • Each key records who created it (created_by). The MCP server can list keys (list_model_keys) and revoke one (revoke_model_key).
  • At most APP_MODEL_MAX_KEYS (50) live keys per model.

What is metered ​

Per key and model, per hour: requests, 2xx, 4xx and 5xx counts, requests refused with 429, input and output tokens, response bytes, total duration and time to first byte (the first byte the engine sends, not the first token).

Tokens are read from the response's usage object: non-streaming responses, and streams opened with stream_options.include_usage. A stream without it, or a compressed response, is counted as a request only. Nothing is estimated, and the report says so (usage_requests shows how many requests carried usage). Prompts and bodies are never stored or logged; only the last APP_MODEL_USAGE_SCAN_BYTES of a JSON or event-stream body are held in memory to find usage.

Counts are aggregated in memory and written in bounded batches, so the request path never waits on the database. If the control plane stops between flushes, up to one interval of counts is lost.

VariableDefaultEffect
APP_MODEL_USAGE_FLUSH_INTERVAL30sHow often aggregates are written.
APP_MODEL_USAGE_BATCH_SIZE200Rows per write.
APP_MODEL_USAGE_MAX_BUFFERED10000Hourly aggregates held between flushes; more are dropped and logged.
APP_MODEL_USAGE_RETENTION720hHourly rows older than this are deleted. 0 keeps them.
APP_MODEL_USAGE_SCAN_BYTES16384Tail of a response searched for usage. 0 turns token counting off.
APP_MODEL_KEY_MAX_LIMIT10000000Largest rpm, tpm or parallel value a key may be given.

The control plane has no Prometheus endpoint yet, so usage is exposed through GET /api/v1/models/{name}/usage?since=24h, the CLI and the get_model_usage MCP tool.

GPU apps ​

Any app can request a GPU with resources.gpu in app.yaml (see the app spec reference). Such an app only runs on a node with a usable GPU: the reconciler reports NoGPUOnNode or GPURuntimeMissing instead of starting it, and moving it to a node without one is rejected.

GPU scheduling ​

Placement counts GPUs, not just workloads. Each node has a ledger built from the desired state: every app with resources.gpu and every model on the node reserves GPUs.

RequestReserves
a count, gpus: 2that many GPUs, anonymous
device IDs (index or UUID)those exact devices; a second workload asking for the same device does not fit
all or no countevery GPU, so the node must be completely free

A workload fits a node when the node reports a GPU, Docker has the nvidia runtime, and enough GPUs are free (total minus reserved, never below zero). A node that is exactly full fits nothing more. A workload's own reservation is ignored when re-checking it against its own node.

Where the ledger is consulted:

  • Create. A new GPU app without node_id is auto-placed on the least loaded node that fits (the local host is the fallback), or refused with 409 and a per-node reason. An explicit node_id is checked the same way once the node has reported.
  • Move. PUT /api/v1/apps/{name}/node rejects a target that lacks the runtime or the free GPUs.
  • Drain. Each GPU app is placed on a node that fits, and GPUs it takes are reserved for the apps after it in the same drain. An app no node can host stays put and is reported under blocked (dashboard drain dialog, levelrail nodes drain, API). Models are never moved, so they are always listed as blocked.

Docker does not isolate GPUs by count: two containers can share a device, so the ledger is scheduling accounting, not enforcement. A running app is never stopped because the ledger says the node is oversubscribed.

Seeing reservations ​

  • API. GET /api/v1/gpus adds reserved_gpus, free_gpus and reservations (app:<name>, model:<name>) per node. GET /api/v1/nodes and GET /api/v1/nodes/{id} carry a gpu summary.
  • CLI. levelrail nodes list has a GPU column (free/total), levelrail nodes get prints a GPU block, levelrail models gpus has a RESERVED column.
  • Dashboard. The AI models page GPU cards and the node detail GPU card show reserved vs total GPUs and VRAM used vs total; the node list shows a GPU free/total badge.
  • MCP. list_gpu_nodes returns the same fields.

Attention and doctor ​

When a GPU app or model cannot run on its own node and no eligible GPU node has enough free GPUs, levelrail doctor and the attention list (Status page, levelrail attention, MCP get_attention) raise a warning GPU app <name> cannot be placed. Fix it by freeing a GPU (stop or shrink another GPU workload), adding a GPU node, or installing the nvidia container toolkit on the node that has GPUs.

Access control ​

Listing and reading models and GPUs needs the read ability. Deleting and restarting needs write. Listing keys and reading usage need read; revoking a key needs write. Deploying, creating or rotating a key and setting a HuggingFace token need write:sensitive. Model resources can be targeted by IAM policies as model:<name>. The AI assistant asks for confirmation before any of the mutating model tools.

Not in version 1 ​

AMD and Apple GPUs, MIG partitioning, automatic model-to-node scheduling (you pick the node; apps are spread, see GPU scheduling), moving a model between nodes, per-model limit overrides, swap groups, and a Prometheus endpoint for usage.

GPU in Compose templates ​

A Compose file can request NVIDIA GPUs with the standard deploy.resources.reservations.devices block. Levelrail reads driver (only nvidia or unset), count (a number or all, unset means all), device_ids and capabilities: [gpu], and maps it onto resources.gpu. The service then only starts on a node with a working NVIDIA runtime. Catalogue entries that need a GPU carry a "Needs NVIDIA GPU" badge.

Released under the Apache 2.0 License.