clika_runtime.modelverse.serve
Serve a text-generation, a text-to-image or a text-to-video model over the OpenAI-compatible HTTP API.
::
import clika_runtime.modelverse as mv
server = mv.serve("Qwen/Qwen2.5-0.5B-Instruct", device="cuda:0", port=8000)
print(server.base_url) # POST /v1/chat/completions and /v1/messages, GET /v1/models, /health
server.serve_forever() # until Ctrl-C; or server.stop() from another thread
images = mv.serve("Tongyi-MAI/Z-Image-Turbo", task="text-to-image", device="cuda:0", port=8001)
# POST /v1/images/generations (b64_json), GET / renders the image page
videos = mv.serve("Wan-AI/Wan2.1-T2V-1.3B-Diffusers", task="text-to-video", device="cuda:0", port=8002)
# POST /v1/videos/generations (the clip as base64 MP4; this model reads the prompt alone,
# with no soundtrack and no keyframes), GET / renders the video page
A :class:Server wraps the library's own HTTP server: it listens on a
background thread the library owns, wait_until_ready probes /health
with a deadline, and stop cancels in-flight requests at their next decode
step and joins every server thread. No Python runs on the request path.
| Name | Kind |
|---|---|
Server | class |
| functions | module functions |