---
title: "clika_runtime.modelverse.serve"
sidebar_label: "clika_runtime.modelverse.serve"
description: "The clika_runtime.modelverse.serve module, introspected from the installed wheel."
---

<!-- Generated by tools/api_reference/generate_api_docs.py. Do not edit. -->

Serve a text-generation, a text-to-image or a text-to-video model over the OpenAI-compatible HTTP API.

::

    import clika_runtime.modelverse as mv

    server = mv.serve("Qwen/Qwen2.5-0.5B-Instruct", device="cuda:0", port=8000)
    print(server.base_url)          # POST /v1/chat/completions and /v1/messages, GET /v1/models, /health
    server.serve_forever()          # until Ctrl-C; or server.stop() from another thread

    images = mv.serve("Tongyi-MAI/Z-Image-Turbo", task="text-to-image", device="cuda:0", port=8001)
    # POST /v1/images/generations (b64_json), GET / renders the image page

    videos = mv.serve("Wan-AI/Wan2.1-T2V-1.3B-Diffusers", task="text-to-video", device="cuda:0", port=8002)
    # POST /v1/videos/generations (the clip as base64 MP4; this model reads the prompt alone,
    # with no soundtrack and no keyframes), GET / renders the video page

A :class:`Server` wraps the library's own HTTP server: it listens on a
background thread the library owns, ``wait_until_ready`` probes ``/health``
with a deadline, and ``stop`` cancels in-flight requests at their next decode
step and joins every server thread. No Python runs on the request path.

| Name | Kind |
| --- | --- |
| [`Server`](./Server.md) | class |
| [functions](./functions.md) | module functions |