Skip to main content

clika_runtime.modelverse.serve

Serve a text-generation, a text-to-image or a text-to-video model over the OpenAI-compatible HTTP API.

::

import clika_runtime.modelverse as mv

server = mv.serve("Qwen/Qwen2.5-0.5B-Instruct", device="cuda:0", port=8000)
print(server.base_url) # POST /v1/chat/completions and /v1/messages, GET /v1/models, /health
server.serve_forever() # until Ctrl-C; or server.stop() from another thread

images = mv.serve("Tongyi-MAI/Z-Image-Turbo", task="text-to-image", device="cuda:0", port=8001)
# POST /v1/images/generations (b64_json), GET / renders the image page

videos = mv.serve("Wan-AI/Wan2.1-T2V-1.3B-Diffusers", task="text-to-video", device="cuda:0", port=8002)
# POST /v1/videos/generations (the clip as base64 MP4; this model reads the prompt alone,
# with no soundtrack and no keyframes), GET / renders the video page

A :class:Server wraps the library's own HTTP server: it listens on a background thread the library owns, wait_until_ready probes /health with a deadline, and stop cancels in-flight requests at their next decode step and joins every server thread. No Python runs on the request path.

NameKind
Serverclass
functionsmodule functions