10 KiB
| name | version | description | trigger_conditions | ||||
|---|---|---|---|---|---|---|---|
| multimodal-ai-inference-spec | 1.0.0 | Standardized interface specification for multimodal AI inference functions across different model types (text, image, audio, video) to enable consistent integration with agent systems. |
|
Multimodal AI Inference Interface Specification
Overview
This specification defines standardized function interfaces for various multimodal AI capabilities to ensure consistent integration patterns across text, image, audio, and video models. The goal is to provide a uniform calling convention while accommodating the unique requirements of each modality.
Core Principles
1. Unified Error Handling
All functions must return consistent error structures regardless of modality.
2. Standardized Metadata
Common metadata fields (request_id, timestamp, usage) should be present across all modalities.
3. Async-First Design
All inference functions must be asynchronous to support concurrent operations.
4. User Context Propagation
Support user_id parameter for multi-tenant isolation and quota management.
5. Flexible Input/Output
Accommodate modality-specific inputs (images, audio files) while maintaining consistent structure.
Function Interface Templates
Text-to-Text (local_llm_inference)
Already defined in previous specification
Vision/Image Understanding (local_vision_inference)
async def local_vision_inference(
model: str,
image: Union[str, bytes], # URL, file path, or base64 encoded image
prompt: str = None, # Optional text prompt for VQA
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "gpt-4-vision", "llava-1.6", "claude-3-opus")image: Image input (URL string, file path string, or base64 bytes)prompt: Optional text prompt for visual question answeringkwargs:max_tokens: int (default: 1024)temperature: float (default: 0.7)detail: str ("low", "high", "auto") for vision detail level- Additional modality-specific parameters
Returns:
{
"success": bool,
"content": str, # Generated text description/answer
"usage": {
"input_tokens": int, # Image tokens + text tokens
"output_tokens": int,
"total_tokens": int
},
"metadata": {
"image_format": str, # "jpeg", "png", etc.
"image_size": [int, int], # [width, height]
"model": str,
"provider": str
},
# ... standard fields (request_id, timestamp, error)
}
Text-to-Image (local_image_generation)
async def local_image_generation(
model: str,
prompt: str,
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "stable-diffusion-xl", "dalle-3", "midjourney")prompt: Text prompt for image generationkwargs:negative_prompt: str (optional)width: int (default: 1024)height: int (default: 1024)num_images: int (default: 1)guidance_scale: float (default: 7.5)steps: int (default: 30)
Returns:
{
"success": bool,
"images": List[str], # List of image URLs or base64 encoded images
"usage": {
"prompt_tokens": int,
"generation_steps": int,
"estimated_cost": float
},
"metadata": {
"model": str,
"provider": str,
"dimensions": [int, int],
"format": str # "jpeg", "png"
},
# ... standard fields
}
Text-to-Speech (local_tts_inference)
async def local_tts_inference(
model: str,
text: str,
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "elevenlabs", "google-tts", "azure-tts")text: Input text to synthesizekwargs:voice: str (voice identifier)speed: float (0.5-2.0, default: 1.0)pitch: float (-1.0 to 1.0, default: 0.0)format: str ("mp3", "wav", "ogg", default: "mp3")
Returns:
{
"success": bool,
"audio": str, # Audio URL or base64 encoded audio
"usage": {
"character_count": int,
"duration_seconds": float,
"estimated_cost": float
},
"metadata": {
"model": str,
"provider": str,
"format": str,
"sample_rate": int,
"voice": str
},
# ... standard fields
}
Automatic Speech Recognition (local_asr_inference)
async def local_asr_inference(
model: str,
audio: Union[str, bytes],
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "whisper-large", "google-speech", "azure-stt")audio: Audio input (URL, file path, or base64)kwargs:language: str (ISO 639-1 code, e.g., "en", "zh")task: str ("transcribe", "translate", default: "transcribe")temperature: float (for whisper models)
Returns:
{
"success": bool,
"text": str, # Transcribed/translated text
"usage": {
"audio_duration_seconds": float,
"word_count": int,
"estimated_cost": float
},
"metadata": {
"model": str,
"provider": str,
"language": str,
"confidence": float # 0.0-1.0
},
# ... standard fields
}
Image-to-Image (local_image_editing)
async def local_image_editing(
model: str,
image: Union[str, bytes],
prompt: str = None,
mask: Union[str, bytes] = None,
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "stable-diffusion-inpaint", "dalle-2-edit")image: Original image inputprompt: Edit instruction promptmask: Optional mask for inpainting (same dimensions as image)kwargs:strength: float (0.0-1.0, default: 0.8)num_images: int (default: 1)
Returns: Similar to text-to-image but includes original image metadata
Text-to-Video (local_video_generation)
async def local_video_generation(
model: str,
prompt: str,
**kwargs
) -> Dict[str, Any]:
Parameters:
model: Model identifier (e.g., "sora", "runway-gen2", "pika")prompt: Video generation promptkwargs:duration: float (seconds, default: 4.0)fps: int (default: 24)resolution: str ("720p", "1080p", default: "720p")
Returns:
{
"success": bool,
"video": str, # Video URL or base64 encoded video
"usage": {
"prompt_tokens": int,
"duration_seconds": float,
"estimated_cost": float # Typically higher for video
},
"metadata": {
"model": str,
"provider": str,
"format": str, # "mp4", "webm"
"resolution": str,
"fps": int
},
# ... standard fields
}
Image-to-Video (local_image_to_video)
async def local_image_to_video(
model: str,
image: Union[str, bytes],
prompt: str = None,
**kwargs
) -> Dict[str, Any]:
Parameters: Similar to text-to-video but with image input
Standard Return Structure (All Modalities)
Success Response
{
"success": True,
# Modality-specific content field(s)
"usage": { /* modality-specific usage metrics */ },
"metadata": { /* modality-specific metadata */ },
"request_id": str, # Unique request identifier
"timestamp": str, # ISO 8601 timestamp
"user_id": str, # User context (if provided)
"model": str, # Actual model used
"provider": str # Actual provider used
}
Error Response (Universal)
{
"success": False,
"error": {
"code": str, # Standard error codes (see below)
"message": str, # User-friendly message
"details": str # Technical details (optional)
},
"request_id": str,
"timestamp": str,
"model": str, # Requested model (if available)
"user_id": str # User context (if provided)
}
Standard Error Codes
MODEL_NOT_FOUND: Model identifier not recognizedINVALID_INPUT: Input format/type invalid for modalityUNSUPPORTED_OPERATION: Model doesn't support requested operationRATE_LIMIT_EXCEEDED: Too many requestsQUOTA_EXHAUSTED: User quota exceededPROVIDER_UNAVAILABLE: Underlying provider service downINPUT_TOO_LARGE: Input exceeds size limitsTIMEOUT: Request timed outINTERNAL_ERROR: Unexpected internal error
Integration Guidelines
For Platform Developers
- Implement all functions as async to support concurrent operations
- Include comprehensive input validation with clear error messages
- Support both URL and base64 inputs for media files where applicable
- Implement proper timeout handling (30-300 seconds based on modality)
- Include usage tracking for billing and monitoring
- Support user context propagation via user_id parameter
For Agent System Integration (Hermes Agent)
- Use consistent calling patterns across all modalities
- Handle errors uniformly regardless of underlying model type
- Cache results appropriately based on modality characteristics
- Implement retry logic for transient failures
- Respect rate limits and quota constraints per user
Example Usage Patterns
Vision Analysis
result = await local_vision_inference(
model="gpt-4-vision",
image="https://example.com/image.jpg",
prompt="Describe this image in detail",
temperature=0.2
)
Generate and Speak
# Generate response
text_result = await local_llm_inference(
model="qwen3-max",
messages=[{"role": "user", "content": "Tell me a joke"}]
)
# Convert to speech
if text_result["success"]:
audio_result = await local_tts_inference(
model="elevenlabs",
text=text_result["content"],
voice="Rachel"
)
This specification enables consistent integration of diverse multimodal AI capabilities while accommodating the unique requirements of each modality.