Translate text
A translation family provides translate, and the workflow is the transcription guide's sibling: resolve the model, run its command, get the payload on stdout. This guide translates sentences locally, then serves translation over HTTP. The model is facebook/m2m100_418M, a sequence-to-sequence translation model covering a hundred languages under two-letter codes; its model card is where its license and its language inventory are read, and the runtime prints the license at fetch, at load and in info. A chat model that translates (Hy-MT2 among them) answers through its prompt command instead, since its family provides no translate.
From the command line
clikart-cli facebook/m2m100_418M translate --to ko \
"The installation is complete." "Pick a model from the catalog."
설치가 완료되었습니다.
카탈로그에서 모델을 선택합니다.
Each source text prints as one line, in input order; the payload is only the translations, so redirecting into a file gives one translation per line. Decoding progress and timing ride stderr. The flags:
--to LANG(required): the target language code from the model's own inventory; M2M-100 uses two-letter codes (en,ko,fr,de,ja,zh, ...), and the server's/propslists them underlanguages.--from LANG: the source language code; unset means the family's default source, which read the English above.--max-new-tokens N: the decode budget in new tokens (0, the default, keeps the family default, clamped to the model's window).- The load knobs apply unchanged:
--device,--cache-dir,--offline.
When the budget binds first
A translation that stops at its decode budget rather than at the model's end token still exits 0 and still prints its payload; what marks it is one warning on stderr naming the limit that bound it:
clikart-cli facebook/m2m100_418M translate --to ko --max-new-tokens 8 \
"The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon."
회의는 10시에 시작되며
[2026-10-09 15:05:25.113] [m2m-100] [warning] translation stopped at the 8-token budget (--max-new-tokens); output may be incomplete
The output above is cut on purpose: an 8-token budget cannot hold the sentence, and the warning is the demonstration. A script that must catch truncation without parsing stderr uses the HTTP route below, where each row carries finish_reason.
Over HTTP
The same model serves POST /v1/translations (the route table):
clikart-cli facebook/m2m100_418M serve --port 8000
curl -s http://127.0.0.1:8000/v1/translations \
-H "Content-Type: application/json" \
-d '{"texts": ["Hello.", "The meeting starts at ten, the slides are on the shared drive, and the notes will follow by email in the afternoon."],
"source_lang": "en", "target_lang": "ko", "max_new_tokens": 8}'
{"model":"m2m-100","source_lang":"en","target_lang":"ko","translations":[{"text":"안녕하세요","finish_reason":"stop"},{"text":"회의는 10시에 시작되며","finish_reason":"length"}],"usage":{"prompt_tokens":34,"completion_tokens":10,"total_tokens":44}}
Batch through the texts array; the response carries one translation per input, in order. Each row's finish_reason is the machine-readable form of the budget story above: stop means the model produced its end token, length means the decode budget cut the translation short (the small max_new_tokens above forces one of each). Omit max_new_tokens to keep the family default. source_lang is optional, as --from is.
The server-side flags and operations story (binding, health, admission control) is the common one in Serve a model over the OpenAI and Anthropic APIs. Sizing and quantization work as everywhere else (Model requirements); for speech instead of text as the input, Transcribe audio chains into this guide.