15 KiB

name description trigger_conditions
gpu-async-service-pattern Pattern for deploying GPU services with longtasks (submit+status dual endpoints) and integrating them into Sage llmage as async models.
Deploying or fixing a GPU service that needs async submit/status endpoints
Converting a sync Sage model to async (submit+poll)
Setting up longtasks-based workers with Redis
Debugging asyncinference KeyError/collation/status issues

GPU Async Service Pattern

Architecture

Client → Sage llmage → upapp → GPU service (ahserver + longtasks)
                ↓ asyncinference polls every N seconds
                ↓ GET /api/status?task_id=xxx
                ↓ Redis: {taskname}:task:{taskid}

GPU Service Requirements

1. Submit Endpoint (/api/{service}-submit)

  • Use longtasks to submit tasks
  • Return the longtasks-generated task_id (not a custom UUID)
  • Return {"task_id": "...", "status": "queued"}
# CORRECT — use longtasks task_id
result = await longtasks.submit_task(payload)
task_id = result['task_id']
return json.dumps({'task_id': task_id, 'status': 'queued'})

# WRONG — uses custom task_id, loses longtasks tracking
task_id = str(uuid.uuid4()).replace("-", "")[:12]
await longtasks.submit_task(payload)
return json.dumps({'task_id': task_id, 'status': 'queued'})

2. Status Query Endpoint (/api/{service}-status)

  • Read from Redis using longtasks.get_redis_task(task_id)
  • Return status field (uppercase: SUCCEEDED/FAILED/PENDING)
  • Include usage in SUCCEEDED responses
task = await longtasks.get_redis_task(task_id)
status = task.get('status', 'unknown')
result = {'task_id': task_id, 'status': status}
if status == 'SUCCEEDED':
    data = task.get('result', {})
    result['usage'] = data.get('usage', {})
    # Include business-specific output fields
    result['output_url'] = data.get('output_path', '')
elif status == 'FAILED':
    result['error'] = str(task.get('result', ''))
return json.dumps(result, ensure_ascii=False)

3. Worker Requirements

  • Return status: "SUCCEEDED" (uppercase, not "success" or "ok")
  • Include usage field with appropriate units
  • Match task_type between submit and worker
async def process_task(self, payload, workid=None):
    task_type = payload.get('task_type', '')
    if task_type == 'separate_full':
        return {
            'status': 'SUCCEEDED',
            'usage': {'audio_seconds': round(duration, 2)},
            'output_path': '/tmp/output.wav'
        }
    raise ValueError(f'Unknown task_type: {task_type}')

Sage llmage Integration

4. Submit UAPI — Async Template

UPDATE uapi SET stream='async',
  data='{"audio_path":"{{audio_file}}"}',
  response='{"taskid":"{{task_id}}","taskstatus":"PENDING","status":"PENDING"}'
WHERE id='uapi_xxx';

5. Status Query UAPI

INSERT INTO uapi (id,name,upappid,stream,path,httpmethod,data,response)
VALUES ('uapi_xxx_status','xxx-status','ktv-gateway','sync',
  '/api/status?task_id={{taskid}}','GET',NULL,
  '{"status":"{{status}}","output_url":"{{output_url}}","usage":{{json.dumps(usage,ensure_ascii=False)}}}');

6. llm_api_map

UPDATE llm_api_map SET query_apiname='xxx-status',query_period=5
WHERE llmid='llm_xxx';

Common Pitfalls

getID() not defined in dspy

When installing packages from pipeline-app into Sage's venv, the ahserver module may be overwritten. The pipeline-app ahserver uses different dspy globals. Fix: replace getID() with str(__import__("uuid").uuid4()).replace("-","") in dspy files.

Collation mismatch

Production DB may have utf8mb4_general_ci columns. Fix with:

ALTER TABLE pricing_program MODIFY ownerid VARCHAR(32) COLLATE utf8mb4_unicode_ci;

This matches xls2ddl standard: CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci.

KeyError('status') in asyncinference

The query endpoint response must include status key (not just taskstatus). Sage's asyncinference.py line 179 checks new_output.get('status').

Language code for ASR

Faster-whisper does NOT accept auto as language code. Use zh, en, etc. explicitly.

Nested Result in longtasks response

GPU longtasks wrap output in result key. RUNNING state has NO result key — only SUCCEEDED does:

// RUNNING — no result key
{"status":"RUNNING","task_id":"...","started_at":...}
// SUCCEEDED — result contains usage
{"status":"SUCCEEDED","result":{"segments":[...],"usage":{...}}}

Template guard: Use top-level {%if status == "SUCCEEDED"%}, never {%if result.status == "SUCCEEDED"%} (result undefined during RUNNING → UndefinedError).

-- CORRECT uapi response template:
UPDATE uapi SET response = 
  '{"status":"{{status}}"{%if status == "SUCCEEDED"%},"usage":{{json.dumps(result.usage,ensure_ascii=False)}}{%endif%}}';

MySQL eats backslash-quotes in uapi templates

UPDATE uapi SET response = '{\"status\"...}' — MySQL strips \" → stored as {status:...} (invalid JSON). Use heredoc SQL file:

cat > /tmp/fix.sql << 'SQLEOF'
UPDATE uapi SET response = '{"status":"{{status}}"}' WHERE name = 'x';
SQLEOF
mysql < /tmp/fix.sql

VibeVoice-ASR Deployment (No Docker, No vLLM)

When Docker is unavailable, deploy via transformers directly — same ahserver + LongTasks pattern as fastwhisper. Transformers ≥ 5.14.1 has built-in VibeVoice ASR — prefer AutoModel.from_pretrained() over source-code imports. The source __init__.py triggers AutoModel.register() at module level that conflicts with Transformers' pre-registered configs.

Model Download (HF Blocked → ModelScope)

from modelscope import snapshot_download
snapshot_download('microsoft/VibeVoice-ASR', local_dir='/share/models/VibeVoice-ASR-7B')
# Tokenizer separately (HF unreachable for processor auto-download)
snapshot_download('Qwen/Qwen2.5-7B', local_dir='/share/models/Qwen2.5-7B',
    allow_patterns=['tokenizer*', 'vocab*', '*.json', '*.txt'])

Model Loading — Hybrid Approach

Model: Transformers built-in (avoids source-code registration conflicts). Processor: VibeVoice source (has ffmpeg/soundfile audio loading that Transformers' VibeVoiceAsrProcessor alone lacks).

import sys; sys.path.insert(0, "/share/ymq/VibeVoice")
from transformers import AutoModel
from vibevoice.processor.vibevoice_asr_processor import VibeVoiceASRProcessor

model = AutoModel.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    device_map="cuda:0",           # SINGLE GPU — auto causes cross-device indexing errors
    ignore_mismatched_sizes=True,   # checkpoint architecture differs from HF class
)
processor = VibeVoiceASRProcessor.from_pretrained(
    model_path,
    language_model_pretrained_name="/share/models/Qwen2.5-7B",  # LOCAL path
)

Why device_map="cuda:0" not "auto"? With multi-GPU sharding, the model's custom encode_speech() produces tensors on mixed devices, causing RuntimeError: indices should be on cpu or same device (cuda:N). 7B BF16 (~14GB) fits in 24GB with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.

Python Environment: Mix venv + system transformers

The vllm venv (/share/vllm-0.8.5) has Transformers 4.57.6 (too old). System python3 has 5.14.1. Use venv python with system transformers prepended:

PYTHONPATH=/data/ymq/.local/lib/python3.10/site-packages:/share/ymq/VibeVoice:$PWD:$PYTHONPATH \
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
/share/vllm-0.8.5/bin/python app/vibevoice_asr_app.py -p 9926 -w $PWD

This keeps appPublic, ahserver, longtasks from the venv while getting newer transformers from user site-packages.

Audio Inference

def _transcribe(self, fpath):
    inputs = self.processor(
        audio=fpath,              # file path — NOT (array, sr) tuple
        return_tensors="pt",
        padding=True,             # required: single-sample needs batch dim
        add_generation_prompt=True,
    )
    inputs = {k: v.to("cuda:0") if isinstance(v, torch.Tensor) else v
              for k, v in inputs.items()}
    with torch.no_grad():
        gen = self.model.generate(**inputs, max_new_tokens=4096,  # 512 truncates >30s audio
                                   temperature=0.0, do_sample=False)
    text = self.processor.decode(gen[0], skip_special_tokens=True)
    segments = self.processor.post_process_transcription(text)
    # Output: [{start_time, end_time, speaker_id, text}, ...]

Full Pitfall List

# Symptom Cause Fix
1 model_type vibevoice not recognized ModelScope checkpoint uses vibevoice; HF expects vibevoice_asr Edit config.json model_type
2 HTTPSConnectionPool HF unreachable Processor tries downloading Qwen tokenizer from HF Use local ModelScope copy
3 ValueError: already used by a Transformers model Source AutoModel.register() without exist_ok=True Patch all modular/*.py: add exist_ok=True
4 No module named 'tokenization_qwen2_fast' Transformers 5.14.1 renamed fast tokenizer Import tokenization_qwen2 as Qwen2TokenizerFast
5 tie_weights() unexpected kwarg 'recompute_mapping' TF 5.x changed signature; source model overrides it Use AutoModel (TF built-in), not source model class
6 Can't load feature extractor: no preprocessor_config.json Checkpoint lacks this file Create minimal {"feature_extractor_type":"vibevoice_asr"}
7 setting an array element with a sequence Passing (array, sr) tuple to processor Pass file path directly, add padding=True
8 indices on cuda:N cross-device error device_map="auto" shards model across GPUs Use device_map="cuda:0" (single GPU)
9 405 Method Not Allowed on /api/asr ahserver startswiths RegisterFunction only handles GET requests — POST returns 405 Use GET with query params: curl 'http://localhost:9926/api/asr?audio_file=xxx'. If POST is required, ahserver config needs explicit HTTP method handling (not supported by default RegisterFunction).
10 TypeError: can only concatenate list (not "BatchEncoding") to list at full_tokens = system_tokens + user_tokens Transformers 5.x tokenizer.apply_chat_template(tokenize=True) returns BatchEncoding, not a plain list (4.x behavior) In vibevoice_asr_processor.py: full_tokens = list(system_tokens) + list(user_tokens)
11 Import hangs indefinitely during from vibevoice.processor import ... vibevoice/modular/__init__.py triggers AutoModel.register() chain at module level — must patch ALL files, not just modular_vibevoice_tokenizer.py Files needing exist_ok=True: modular_vibevoice_tokenizer.py, modular_vibevoice_diffusion_head.py, modeling_vibevoice.py, modeling_vibevoice_asr.py, modeling_vibevoice_streaming.py, modeling_vibevoice_streaming_inference.py. Also fix tokenization_qwen2_fasttokenization_qwen2 as Qwen2TokenizerFast in modular_vibevoice_text_tokenizer.py.
12 max_new_tokens=512 truncates long-audio JSON: parser fails, 0 segments 224s audio needs ~820 tokens for full JSON output Set max_new_tokens=4096 for audio > 30s. Add fallback parser: if JSON truncated at last "}, recover by closing array bracket: json_str[:last_good+2] + ']'
13 post_process_transcription returns empty list despite valid model output Model outputs [Lyric]/[Silence] prefixes in Content field that post_process_transcription can't parse Fallback: regex-extract JSON from `assistant\n[...]<

API note: ahserver startswiths with RegisterFunction only handles GET. POST returns 405. Use GET: curl 'http://localhost:9926/api/asr?audio_file=xxx'.

See references/vibevoice-pitfalls.md for full error transcripts.

Output Format Mapping (→ fastwhisper-compatible)

# VibeVoice output: [{start_time, end_time, speaker_id, text}]
# fastwhisper expects: {language, content, segments: [[start, end, text, [word_timestamps]]]}
whisper_segments = []
for seg in segments:
    whisper_segments.append([seg['start_time'], seg['end_time'], seg['text'], []])
return {
    'task_status': 'SUCCEEDED',
    'language': 'auto',
    'content': ' '.join(s['text'] for s in segments),
    'segments': whisper_segments,
}

Note: VibeVoice produces segment-level timestamps (utterance/sentence), not word-level like Whisper's word_timestamps=True. For lyrics where each line is a segment, this is typically sufficient.

Verified Working Services

Service Port Taskname Status
Demucs 9083 demucs async
ASR (faster-whisper) 9925 fastwhisper async
VibeVoice-ASR-7B 9926 vibevoice-asr async (GET-only)
RealESRGAN 9082 realesrgan async
ECAPA-TDNN Voiceprint 9087 voiceprint async (aiohttp)

Lightweight Service Pattern (aiohttp, no ahserver)

When ahserver's dependency chain is unavailable (missing sqlor, checkedHash, etc.), deploy a standalone aiohttp service with in-process task queue:

import asyncio, json, uuid
from aiohttp import web

PENDING = {}

async def handle_submit(request):
    data = dict(request.query)
    task_id = str(uuid.uuid4()).replace('-', '')[:16]
    PENDING[task_id] = {'status': 'queued'}
    asyncio.create_task(_process(task_id, data))
    return web.json_response({'task_id': task_id, 'status': 'queued'})

async def _process(task_id, data):
    try:
        PENDING[task_id]['status'] = 'running'
        loop = asyncio.get_event_loop()
        result = await loop.run_in_executor(None, do_work, data)
        PENDING[task_id] = {'status': 'SUCCEEDED', **result}
    except Exception as e:
        PENDING[task_id] = {'status': 'FAILED', 'error': str(e)}

async def handle_status(request):
    d = PENDING.get(request.query.get('task_id', ''), {})
    return web.json_response({'status': d.get('status', 'unknown'), **d})

app = web.Application()
app.router.add_post('/api/submit', handle_submit)
app.router.add_get('/api/submit', handle_submit)
app.router.add_get('/api/status', handle_status)
web.run_app(app, host='0.0.0.0', port=9087)

GPU offload: use loop.run_in_executor(None, fn, args) to run inference off the event loop. Load the model at module level before web.run_app.

HF Blocked → ModelScope Download

Pre-download models via ModelScope when HF is unreachable:

from modelscope import snapshot_download
snapshot_download('iic/speech_ecapa-tdnn_sv_en_voxceleb_16k', local_dir='/share/models/ecapa-tdnn')

Then load from local path: SpeakerRecognition.from_hparams(..., savedir='/share/models/ecapa-tdnn'). If SpeechBrain still tries HF on first run, edit hyperparams.yaml to set pretrained_path to the local directory. See references/speechbrain-hf-blocked-fix.md.