--- name: multimodal-ai-inference-spec version: 1.0.0 description: Standardized interface specification for multimodal AI inference functions across different model types (text, image, audio, video) to enable consistent integration with agent systems. trigger_conditions: - Need to define standardized interfaces for multimodal AI models (vision, TTS, ASR, image/video generation) - Building an AI platform that aggregates diverse model types requiring unified calling conventions - Integrating multimodal capabilities into agent systems like Hermes Agent - Ensuring consistent error handling and metadata across different AI modalities --- # Multimodal AI Inference Interface Specification ## Overview This specification defines standardized function interfaces for various multimodal AI capabilities to ensure consistent integration patterns across text, image, audio, and video models. The goal is to provide a uniform calling convention while accommodating the unique requirements of each modality. ## Core Principles ### 1. Unified Error Handling All functions must return consistent error structures regardless of modality. ### 2. Standardized Metadata Common metadata fields (request_id, timestamp, usage) should be present across all modalities. ### 3. Async-First Design All inference functions must be asynchronous to support concurrent operations. ### 4. User Context Propagation Support user_id parameter for multi-tenant isolation and quota management. ### 5. Flexible Input/Output Accommodate modality-specific inputs (images, audio files) while maintaining consistent structure. ## Function Interface Templates ### Text-to-Text (local_llm_inference) **Already defined in previous specification** ### Vision/Image Understanding (local_vision_inference) ```python async def local_vision_inference( model: str, image: Union[str, bytes], # URL, file path, or base64 encoded image prompt: str = None, # Optional text prompt for VQA **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "gpt-4-vision", "llava-1.6", "claude-3-opus") - `image`: Image input (URL string, file path string, or base64 bytes) - `prompt`: Optional text prompt for visual question answering - `kwargs`: - `max_tokens`: int (default: 1024) - `temperature`: float (default: 0.7) - `detail`: str ("low", "high", "auto") for vision detail level - Additional modality-specific parameters **Returns:** ```python { "success": bool, "content": str, # Generated text description/answer "usage": { "input_tokens": int, # Image tokens + text tokens "output_tokens": int, "total_tokens": int }, "metadata": { "image_format": str, # "jpeg", "png", etc. "image_size": [int, int], # [width, height] "model": str, "provider": str }, # ... standard fields (request_id, timestamp, error) } ``` ### Text-to-Image (local_image_generation) ```python async def local_image_generation( model: str, prompt: str, **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "stable-diffusion-xl", "dalle-3", "midjourney") - `prompt`: Text prompt for image generation - `kwargs`: - `negative_prompt`: str (optional) - `width`: int (default: 1024) - `height`: int (default: 1024) - `num_images`: int (default: 1) - `guidance_scale`: float (default: 7.5) - `steps`: int (default: 30) **Returns:** ```python { "success": bool, "images": List[str], # List of image URLs or base64 encoded images "usage": { "prompt_tokens": int, "generation_steps": int, "estimated_cost": float }, "metadata": { "model": str, "provider": str, "dimensions": [int, int], "format": str # "jpeg", "png" }, # ... standard fields } ``` ### Text-to-Speech (local_tts_inference) ```python async def local_tts_inference( model: str, text: str, **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "elevenlabs", "google-tts", "azure-tts") - `text`: Input text to synthesize - `kwargs`: - `voice`: str (voice identifier) - `speed`: float (0.5-2.0, default: 1.0) - `pitch`: float (-1.0 to 1.0, default: 0.0) - `format`: str ("mp3", "wav", "ogg", default: "mp3") **Returns:** ```python { "success": bool, "audio": str, # Audio URL or base64 encoded audio "usage": { "character_count": int, "duration_seconds": float, "estimated_cost": float }, "metadata": { "model": str, "provider": str, "format": str, "sample_rate": int, "voice": str }, # ... standard fields } ``` ### Automatic Speech Recognition (local_asr_inference) ```python async def local_asr_inference( model: str, audio: Union[str, bytes], **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "whisper-large", "google-speech", "azure-stt") - `audio`: Audio input (URL, file path, or base64) - `kwargs`: - `language`: str (ISO 639-1 code, e.g., "en", "zh") - `task`: str ("transcribe", "translate", default: "transcribe") - `temperature`: float (for whisper models) **Returns:** ```python { "success": bool, "text": str, # Transcribed/translated text "usage": { "audio_duration_seconds": float, "word_count": int, "estimated_cost": float }, "metadata": { "model": str, "provider": str, "language": str, "confidence": float # 0.0-1.0 }, # ... standard fields } ``` ### Image-to-Image (local_image_editing) ```python async def local_image_editing( model: str, image: Union[str, bytes], prompt: str = None, mask: Union[str, bytes] = None, **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "stable-diffusion-inpaint", "dalle-2-edit") - `image`: Original image input - `prompt`: Edit instruction prompt - `mask`: Optional mask for inpainting (same dimensions as image) - `kwargs`: - `strength`: float (0.0-1.0, default: 0.8) - `num_images`: int (default: 1) **Returns:** Similar to text-to-image but includes original image metadata ### Text-to-Video (local_video_generation) ```python async def local_video_generation( model: str, prompt: str, **kwargs ) -> Dict[str, Any]: ``` **Parameters:** - `model`: Model identifier (e.g., "sora", "runway-gen2", "pika") - `prompt`: Video generation prompt - `kwargs`: - `duration`: float (seconds, default: 4.0) - `fps`: int (default: 24) - `resolution`: str ("720p", "1080p", default: "720p") **Returns:** ```python { "success": bool, "video": str, # Video URL or base64 encoded video "usage": { "prompt_tokens": int, "duration_seconds": float, "estimated_cost": float # Typically higher for video }, "metadata": { "model": str, "provider": str, "format": str, # "mp4", "webm" "resolution": str, "fps": int }, # ... standard fields } ``` ### Image-to-Video (local_image_to_video) ```python async def local_image_to_video( model: str, image: Union[str, bytes], prompt: str = None, **kwargs ) -> Dict[str, Any]: ``` **Parameters:** Similar to text-to-video but with image input ## Standard Return Structure (All Modalities) ### Success Response ```python { "success": True, # Modality-specific content field(s) "usage": { /* modality-specific usage metrics */ }, "metadata": { /* modality-specific metadata */ }, "request_id": str, # Unique request identifier "timestamp": str, # ISO 8601 timestamp "user_id": str, # User context (if provided) "model": str, # Actual model used "provider": str # Actual provider used } ``` ### Error Response (Universal) ```python { "success": False, "error": { "code": str, # Standard error codes (see below) "message": str, # User-friendly message "details": str # Technical details (optional) }, "request_id": str, "timestamp": str, "model": str, # Requested model (if available) "user_id": str # User context (if provided) } ``` ## Standard Error Codes - `MODEL_NOT_FOUND`: Model identifier not recognized - `INVALID_INPUT`: Input format/type invalid for modality - `UNSUPPORTED_OPERATION`: Model doesn't support requested operation - `RATE_LIMIT_EXCEEDED`: Too many requests - `QUOTA_EXHAUSTED`: User quota exceeded - `PROVIDER_UNAVAILABLE`: Underlying provider service down - `INPUT_TOO_LARGE`: Input exceeds size limits - `TIMEOUT`: Request timed out - `INTERNAL_ERROR`: Unexpected internal error ## Integration Guidelines ### For Platform Developers 1. **Implement all functions as async** to support concurrent operations 2. **Include comprehensive input validation** with clear error messages 3. **Support both URL and base64 inputs** for media files where applicable 4. **Implement proper timeout handling** (30-300 seconds based on modality) 5. **Include usage tracking** for billing and monitoring 6. **Support user context propagation** via user_id parameter ### For Agent System Integration (Hermes Agent) 1. **Use consistent calling patterns** across all modalities 2. **Handle errors uniformly** regardless of underlying model type 3. **Cache results appropriately** based on modality characteristics 4. **Implement retry logic** for transient failures 5. **Respect rate limits** and quota constraints per user ## Example Usage Patterns ### Vision Analysis ```python result = await local_vision_inference( model="gpt-4-vision", image="https://example.com/image.jpg", prompt="Describe this image in detail", temperature=0.2 ) ``` ### Generate and Speak ```python # Generate response text_result = await local_llm_inference( model="qwen3-max", messages=[{"role": "user", "content": "Tell me a joke"}] ) # Convert to speech if text_result["success"]: audio_result = await local_tts_inference( model="elevenlabs", text=text_result["content"], voice="Rachel" ) ``` This specification enables consistent integration of diverse multimodal AI capabilities while accommodating the unique requirements of each modality.