5.0 KiB
Pipeline Platform Model API Docs
Unified Model Governance Inference API
Base url: https://pipeline.opencomputing.cn/pipeline-llm/api/v1
All endpoints require Bearer Token authentication (short-lived token, NOT the real model key).
Tenant isolation: requests are routed through the governance chain by the tenant bound to the token; real model keys never leave the server process.
POST /v1/chat/completions
Unified model invocation entry (OpenAI-compatible). Synchronous models return in one round-trip; async models (image-to-video, etc.) are handled automatically by the server: submit task → poll → fetch result, returning a persisted local URL.
Required Parameters
| Param | Type | Description |
|---|---|---|
messages |
array | Conversation messages, [{"role": "user", "content": "..."}] |
Optional Parameters
| Param | Type | Description |
|---|---|---|
model |
string | Model name (e.g. happyhorse-1.1-i2v). Omit = tenant policy default model (primary→backup chain) |
temperature |
float | Sampling temperature |
tools |
array | Tool definitions (t2t models, OpenAI-compatible) |
_purpose |
string | Purpose marker: utility = auxiliary task (classify/select/summarize), uses utility model chain |
_timeout |
int | Per-call timeout in seconds (0=endpoint default; max 900) |
Extra Parameters for Async Models (i2v / t2v / t2i ...)
| Param | Type | Description |
|---|---|---|
image_file |
string | Uploaded media (base64/data URL auto-converted to local public URL; http(s) URL used as-is). Same-capability params uniformly use xxx_file naming |
resolution |
string | Resolution, e.g. 480P / 720P / 1080P |
duration |
int | Video duration (seconds) |
Generated artifacts: upstream artifact URLs expire very quickly (video only 24h); the server persists them locally before returning.
Request Example (chat)
curl -X POST 'https://pipeline.opencomputing.cn/pipeline-llm/api/v1/chat/completions' \
-H 'Authorization: Bearer ***' \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3.8-max",
"messages": [{"role": "user", "content": "Hello"}]
}'
Request Example (image-to-video, async)
curl -X POST 'https://pipeline.opencomputing.cn/pipeline-llm/api/v1/chat/completions' \
-H 'Authorization: Bearer ***' \
-H 'Content-Type: application/json' \
-d '{
"model": "happyhorse-1.1-i2v",
"messages": [{"role": "user", "content": "The character starts rapping"}],
"image_file": "https://example.com/first-frame.png",
"resolution": "720P",
"duration": 5,
"_timeout": 600
}'
Response Example (sync model)
{
"choices": [
{"message": {"content": "Hello!", "role": "assistant"}, "finish_reason": "stop"}
],
"usage": {"prompt_tokens": 10, "completion_tokens": 5}
}
Response Example (async model)
{
"choices": [
{"message": {"content": "https://pipeline.opencomputing.cn/idfile/tmp/.../xxx.mp4", "role": "assistant"},
"finish_reason": "stop"}
],
"usage": {"duration": 5, "SR": 720, "video_count": 1},
"task_id": "e53dee87-...",
"media": {
"status": "SUCCEEDED",
"video": "https://pipeline.opencomputing.cn/idfile/tmp/.../xxx.mp4",
"usage": {"duration": 5, "SR": 720}
}
}
media is the unified output field (consistent across same-capability models): video → video, image → image, 3D model → glb.
Error Response
Failures return OpenAI error structure with actionable messages (gate failures / timeout / missing template config all give specific reasons):
{"error": {"message": "All models in chain unavailable. Last reason: ...", "type": "invalid_request_error", "code": "govern_error"}}
GET /v1/models
List models available to the current tenant, grouped by capability (empty list if tenant has no governance policy).
Optional Parameters
| Param | Type | Description |
|---|---|---|
catelogid |
string | Capability: t2t / t2i / i2v / t2v / embedding / rerank / tts / asr / i2t. Omit = all |
Request Example
curl 'https://pipeline.opencomputing.cn/pipeline-llm/api/v1/models?catelogid=i2v' \
-H 'Authorization: Bearer ***'
Response Example
{
"object": "list",
"data": [
{"id": "happyhorse-1.1-i2v", "object": "model", "name": "happyhorse-1.1-i2v",
"catelogid": "i2v", "vendor": "Alibaba Bailian"}
]
}
Authentication & Short-lived Token
- Tokens are issued by the platform (bound to tenant/project/task, with expiry and call limits, revocable)
- Usage: `Authorization: Bearer ***
- Tokens are NOT real model keys; real keys are resolved server-side only and never leave the process
Billing
- Each call automatically creates a usage record (
llm_usage), billed asynchronously per the pricing program (ppid) attached to the model - Non-token pricing factors (video duration/resolution, etc.) are stored in the record's
usagesfield