Skip to main content

Quickstart: web

Time to first result on a two-core Linux laptop: about 25 minutes, 15 of them waiting for the run. The tutorial explains each step; this page is the short form.

  1. Sign in. From a signup code or an invitation (Sign up with a signup code). The verification link lands you on the Default Project dashboard.
  2. Register a device (2 minutes). Devices, Register Device, the tab for the machine's operating system, copy the command, run it on the machine. Leave the Reusable, many devices switch alone for now: turning it on asks whether to keep the command you copied or make a new one, and a new one revokes the one you copied. The dialog turns into the device list when the device arrives, within a minute. Give the device one more minute before step 4.
  3. Add a model (1 minute). Models, Add Model, paste a Hugging Face URL or owner/repo; HuggingFaceTB/SmolLM2-135M-Instruct is a good first model. A model somebody registered before you comes back as a shared record under your models; your results on it are yours.
  4. Run a benchmark (15 minutes). Benchmarks, New Benchmark, pick the device and the model, Continue, then LLM Performance with Quick on, Run Benchmarks. The log shows the engine push for the first five minutes and then one line per measured cell; the progress bar stays at 0 of 1 until the leg completes.
  5. Read the result (1 minute). The run's page: TTFT, Tokens/sec, Peak decode, TPOT and Peak Memory, each with median, minimum and maximum, and a table of cells (128:128 is prompt tokens : output tokens, c8 is eight parallel streams). Share Results makes a link.
  6. Deploy it (2 to 3 minutes after a benchmark on the same device; 5 to 9 on a device that has never run the engine). Deploy on the result, keep the name, Deploy. The page reads "Starting" and "Waiting for logs" while the engine starts on the device, then Running with the platform address, a hidden key (Reveal key) and the device's own address. Send one request:
curl -s -X POST "https://<platform>/api/v1/servings/<serving>/devices/<device>/proxy/v1/chat/completions" -H "Authorization: Bearer <deployment key>" -H "Content-Type: application/json" -d '{"model":"<id from /v1/models>","messages":[{"role":"user","content":"Hello"}]}'
  1. Invite your team. Settings, then Organization, Add member: the organization role is Admin or Member; give project roles (Developer, Viewer) on the project's Members tab. Only an owner or admin can open a device's terminal. Roles has the table.

What spends credits, where your plan meters them: every benchmark leg and every deployment start sends the engine to the device, and the Subscription card under Settings shows the balance.