# Nodeau — please read this first Nodeau is young software. It is useful, it is honest about what it does, and it is not finished. This page says exactly where the edges are, because finding them yourself at 1am is a worse experience. ## What is actually supported | | | |---|---| | Operating system | **Ubuntu 24.04 LTS**, native install · **macOS 14+** on Apple Silicon, standalone | | Architecture | **x86_64** (Linux) · **arm64** (Apple Silicon) | | GPU | **NVIDIA** on Linux — one or more per machine · **Apple Silicon** integrated GPU via Metal | | Machines | **one** on Home · **up to three** on Home Pro | | Runtime | **llama.cpp** | | Kinds of work | **chat · embeddings · reranking · structured output · tool calling · image understanding** — measured on Linux/NVIDIA | | Curated models | **8 recommended**, spanning 8 GB / 12 GB / 16 GB / 24 GB-or-multi-GPU cards, plus 8 retired ones that still install and run | | Your own models | **GGUF only**, on Linux/NVIDIA — imported, inspected and qualified on your own hardware | Anything outside that list is untested, not "probably fine". **Beyond chat, on Linux/NVIDIA.** Nodeau serves embeddings at `/v1/embeddings`, reranking at `/v1/rerank`, and structured output, tool calling and image input on the ordinary chat endpoint. Each was measured end to end on real hardware, and each has a control proving the result could not have come from anywhere else. **On a Mac those five are not measured yet**, so the support matrix shows a dash rather than a tick. The runtime installed there carries the routes; carrying a route is not the same as having served a request, and this project does not publish the difference as if it were. **Speech to text is not offered.** It is built and deliberately withheld: through the ordinary customer path the model returned a usable transcript in ten runs of twelve, and llama.cpp describes its own audio support as experimental. One request in six coming back with a non-answer is not something to put beside the rest of this list. **Two things about reranking and structured output, said precisely.** There is no OpenAI reranking API, so `/v1/rerank` follows the Jina/Cohere convention rather than claiming a compatibility that does not exist. And structured output *constrains generation* to your schema with a compiled grammar — it is much stronger than asking nicely in a prompt, and it is not the same as validating the response afterwards. ## Bring your own model You can give Nodeau a **GGUF** file of your own. It reads the file's header to work out what the model is — architecture, quantisation, the shape its KV cache depends on, whether it carries a chat template — hashes the bytes, and registers it under a name you choose: ``` nodeau model import ./my-model.gguf --alias finance-model nodeau model qualify finance-model nodeau run finance-model ``` **What the bytes are is the identity.** A model is identified by the SHA-256 of its contents, never by its filename or the name you gave it. Importing the same bytes twice costs no extra disk. Importing *different* bytes under a name you have used before creates a new artifact and tells you so — it never silently repoints the old one, and the old qualification evidence stops applying because it was recorded against a different digest. **Nothing in your file is executed.** A GGUF is data. Nodeau reads the header and stops, and your model runs on Nodeau's own pinned runtime. There is no `trust_remote_code`, no publisher script and no download hook. **Importing is not qualifying, and loading is not qualified.** After an import Nodeau knows what your file says about itself. `nodeau model qualify` is what finds out the rest: it starts the model on a real card, records how much memory it actually took, and exercises each capability with a probe designed to fail if the model cannot do it — a tool call whose arguments parse and name the city you asked about; JSON checked against the schema it was asked for, rather than merely checked for being JSON; an answer that is actually there, because a reasoning model can spend its whole budget thinking and return an empty string with a perfectly successful response code. An HTTP 200 is not a capability. **Partial results are kept as partial.** If chat passes and tool calling fails, your model stays usable for chat and Nodeau does not claim the rest. **A model too big for one card can use two, in the same machine.** Qualified on an 18.9 GB model across an RTX 3080 and an RTX 5060 Ti, neither of which can hold it alone: Nodeau refuses each card by name with the arithmetic, then splits the layers across the pair. Memory does not pool — each card holds its own share plus its own overhead — so a pair is only as usable as its tighter card. ### What it will not do - It is **GGUF only**. Not Safetensors, not PyTorch, not ONNX, and there is no conversion step. A model published in another format is not supported, and Nodeau will say so rather than try. - It does not run **any** model. A valid GGUF whose architecture the pinned runtime does not implement will import cleanly and fail to serve — that is reported as unsupported, not as a damaged file, because your file is fine. - It will not guess. If a GGUF does not declare the shape a memory estimate needs, Nodeau refuses to schedule it rather than invent a VRAM figure. A figure that is too low is an out-of-memory kill on your hardware. - **Image input is not available for a model you import.** A vision model is two files — the weights and a multimodal projector — and this version imports one. Your model is registered and serves text; image requests are refused up front rather than started and broken later. That is a limit of Nodeau's import, not a judgement about your file. - **Qualification never quietly lowers Nodeau's safety margin to make your model fit.** If it does not fit, the run says so and stops. Accepting that risk is your decision — `--accept-estimate-risk` or `--spend-safety-reserve` — and a model qualified with either records the fact permanently, so a result that came from margin you chose to spend never reads like one Nodeau measured. - Custom models are **not** part of the curated set. Nodeau did not choose your model, review its licence, or test it before release. `nodeau model list` shows the two separately and always will. - Nodeau does not fetch a custom model for you, and does not copy it between your machines. Import it on each machine that should run it; the digest makes the second import free if the bytes are already there. **Your model's weights stay on your machines.** Nodeau Cloud is never sent them and never sees them. What leaves your fleet is the same thing that always did: your machines calling out to report that they are alive. **On Apple Silicon this is untested.** The code builds and the Mac plane is wired, but no Mac was available to qualify it, so the support table above says Linux/NVIDIA and means it. **On a Mac, Nodeau runs standalone**: it does not install a cluster, cannot join a fleet, and batch inference is Linux-only. That is a deliberate boundary, not a gap waiting to close — see the support matrix on nodeau.ai. ## What "profiled" means, and why it matters Nodeau refuses to deploy a model it cannot predict the memory use of. That is the point of the product — it would rather say no than let you discover the problem as a crash loop. Nodeau recommends **8 curated models**, and every one of them carries a hardware recommendation — 8 GB, 12 GB, 16 GB, or 24 GB and multi-GPU — so you can choose without doing VRAM arithmetic from a filename. Eight more models are retired: still listed under `nodeau model list --all`, still installable, still runnable, and not recommended for new installs. **A hardware recommendation is guidance, not the decision.** Admission still works from your actual card, your actual context and what is actually running on it, so it may well prove that a model labelled "12 GB" fits your 10 GB card — and it may refuse one on a 12 GB card that is already busy. The label is what to buy; admission is what will run. **9 configurations are fully measured** on real hardware. Everything else is **estimated** from the model's architecture and artifact size, carries an extra margin, and says so in the decision it produces. Nodeau will not quietly present an estimate as a measurement. That is a real limitation, and it is deliberate. ## What Nodeau will never do to your machine - It will **not** install, upgrade, downgrade or remove your NVIDIA driver. - It will **not** disable Secure Boot or enrol keys. - It will **not** modify your bootloader or kernel command line. - It will **not** touch Windows partitions, or mount NTFS. - It will **not** delete a Kubernetes cluster it did not create. - It will **not** expose inference to your network. The endpoint binds `127.0.0.1` only, and there is no option to change that. - It will **not** send us anything. There is no telemetry. ## Known limitations - **Multi-GPU is qualified narrowly.** Nodeau schedules every qualified NVIDIA GPU in a Linux machine independently, and since v0.4.0 a single model workload can span several of them. That was qualified on **one heterogeneous pair** (RTX 3080 + RTX 5060 Ti) with two models. Other combinations are untested. - **Multi-GPU is not a speedup.** It exists so a model too large for one card can run. On the qualifying pair, a model that fitted on one card was **21% slower** split across two than on the faster card alone. Memory is also not pooled: each card must independently hold its own share. - **Batch jobs can use more than one GPU, narrowly.** A job may run several independent workers, each holding its own card and taking records from the same input, or one worker whose model spans several cards. Both were qualified on the same single heterogeneous pair as everything else here, with one model. A job still does not span machines, and its workers are not restarted or moved if a card is lost. - **A worker count is a request, not a guarantee.** Each worker is a complete model instance and reserves a card of its own, so asking for more workers than there are free cards runs the ones that fit and leaves the rest waiting. The job is not refused and no records are lost — workers take the next unclaimed record, so one worker finishes the job and two finish it sooner. `nodeau batch status` reports how many are running and how many are waiting. - **No scale-to-zero, queueing or cloud overflow.** - **WSL is not supported.** The installer detects it and refuses cleanly. - **AMD and Intel GPUs are not supported.** - **The local endpoint lives in your login session.** It survives closing the terminal, but stops when you log out unless you enable systemd lingering yourself. See the troubleshooting guide. - **Day-to-day commands currently use your kubeconfig.** Narrowing that is tracked and is not finished. ## What we are not claiming Nodeau is not production infrastructure. It has no uptime guarantee, no failover, no support SLA, and it is not hardened for hostile multi-tenancy. If you are putting something important on it, don't — yet. ## Reporting problems `nodeau support bundle` produces a sanitised archive: no API keys, no kubeconfig credentials, no prompts, no model output, no weights. Read it before you send it — it is a plain tar.gz and that is a reasonable thing to want to check.