10 KiB

name version description trigger_conditions
multimodal-ai-inference-spec 1.0.0 Standardized interface specification for multimodal AI inference functions across different model types (text, image, audio, video) to enable consistent integration with agent systems.
Need to define standardized interfaces for multimodal AI models (vision, TTS, ASR, image/video generation)
Building an AI platform that aggregates diverse model types requiring unified calling conventions
Integrating multimodal capabilities into agent systems like Hermes Agent
Ensuring consistent error handling and metadata across different AI modalities

Multimodal AI Inference Interface Specification

Overview

This specification defines standardized function interfaces for various multimodal AI capabilities to ensure consistent integration patterns across text, image, audio, and video models. The goal is to provide a uniform calling convention while accommodating the unique requirements of each modality.

Core Principles

1. Unified Error Handling

All functions must return consistent error structures regardless of modality.

2. Standardized Metadata

Common metadata fields (request_id, timestamp, usage) should be present across all modalities.

3. Async-First Design

All inference functions must be asynchronous to support concurrent operations.

4. User Context Propagation

Support user_id parameter for multi-tenant isolation and quota management.

5. Flexible Input/Output

Accommodate modality-specific inputs (images, audio files) while maintaining consistent structure.

Function Interface Templates

Text-to-Text (local_llm_inference)

Already defined in previous specification

Vision/Image Understanding (local_vision_inference)

async def local_vision_inference(
    model: str,
    image: Union[str, bytes],  # URL, file path, or base64 encoded image
    prompt: str = None,        # Optional text prompt for VQA
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "gpt-4-vision", "llava-1.6", "claude-3-opus")
  • image: Image input (URL string, file path string, or base64 bytes)
  • prompt: Optional text prompt for visual question answering
  • kwargs:
    • max_tokens: int (default: 1024)
    • temperature: float (default: 0.7)
    • detail: str ("low", "high", "auto") for vision detail level
    • Additional modality-specific parameters

Returns:

{
    "success": bool,
    "content": str,  # Generated text description/answer
    "usage": {
        "input_tokens": int,      # Image tokens + text tokens
        "output_tokens": int,
        "total_tokens": int
    },
    "metadata": {
        "image_format": str,      # "jpeg", "png", etc.
        "image_size": [int, int], # [width, height]
        "model": str,
        "provider": str
    },
    # ... standard fields (request_id, timestamp, error)
}

Text-to-Image (local_image_generation)

async def local_image_generation(
    model: str,
    prompt: str,
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "stable-diffusion-xl", "dalle-3", "midjourney")
  • prompt: Text prompt for image generation
  • kwargs:
    • negative_prompt: str (optional)
    • width: int (default: 1024)
    • height: int (default: 1024)
    • num_images: int (default: 1)
    • guidance_scale: float (default: 7.5)
    • steps: int (default: 30)

Returns:

{
    "success": bool,
    "images": List[str],  # List of image URLs or base64 encoded images
    "usage": {
        "prompt_tokens": int,
        "generation_steps": int,
        "estimated_cost": float
    },
    "metadata": {
        "model": str,
        "provider": str,
        "dimensions": [int, int],
        "format": str  # "jpeg", "png"
    },
    # ... standard fields
}

Text-to-Speech (local_tts_inference)

async def local_tts_inference(
    model: str,
    text: str,
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "elevenlabs", "google-tts", "azure-tts")
  • text: Input text to synthesize
  • kwargs:
    • voice: str (voice identifier)
    • speed: float (0.5-2.0, default: 1.0)
    • pitch: float (-1.0 to 1.0, default: 0.0)
    • format: str ("mp3", "wav", "ogg", default: "mp3")

Returns:

{
    "success": bool,
    "audio": str,  # Audio URL or base64 encoded audio
    "usage": {
        "character_count": int,
        "duration_seconds": float,
        "estimated_cost": float
    },
    "metadata": {
        "model": str,
        "provider": str,
        "format": str,
        "sample_rate": int,
        "voice": str
    },
    # ... standard fields
}

Automatic Speech Recognition (local_asr_inference)

async def local_asr_inference(
    model: str,
    audio: Union[str, bytes],
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "whisper-large", "google-speech", "azure-stt")
  • audio: Audio input (URL, file path, or base64)
  • kwargs:
    • language: str (ISO 639-1 code, e.g., "en", "zh")
    • task: str ("transcribe", "translate", default: "transcribe")
    • temperature: float (for whisper models)

Returns:

{
    "success": bool,
    "text": str,  # Transcribed/translated text
    "usage": {
        "audio_duration_seconds": float,
        "word_count": int,
        "estimated_cost": float
    },
    "metadata": {
        "model": str,
        "provider": str,
        "language": str,
        "confidence": float  # 0.0-1.0
    },
    # ... standard fields
}

Image-to-Image (local_image_editing)

async def local_image_editing(
    model: str,
    image: Union[str, bytes],
    prompt: str = None,
    mask: Union[str, bytes] = None,
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "stable-diffusion-inpaint", "dalle-2-edit")
  • image: Original image input
  • prompt: Edit instruction prompt
  • mask: Optional mask for inpainting (same dimensions as image)
  • kwargs:
    • strength: float (0.0-1.0, default: 0.8)
    • num_images: int (default: 1)

Returns: Similar to text-to-image but includes original image metadata

Text-to-Video (local_video_generation)

async def local_video_generation(
    model: str,
    prompt: str,
    **kwargs
) -> Dict[str, Any]:

Parameters:

  • model: Model identifier (e.g., "sora", "runway-gen2", "pika")
  • prompt: Video generation prompt
  • kwargs:
    • duration: float (seconds, default: 4.0)
    • fps: int (default: 24)
    • resolution: str ("720p", "1080p", default: "720p")

Returns:

{
    "success": bool,
    "video": str,  # Video URL or base64 encoded video
    "usage": {
        "prompt_tokens": int,
        "duration_seconds": float,
        "estimated_cost": float  # Typically higher for video
    },
    "metadata": {
        "model": str,
        "provider": str,
        "format": str,  # "mp4", "webm"
        "resolution": str,
        "fps": int
    },
    # ... standard fields
}

Image-to-Video (local_image_to_video)

async def local_image_to_video(
    model: str,
    image: Union[str, bytes],
    prompt: str = None,
    **kwargs
) -> Dict[str, Any]:

Parameters: Similar to text-to-video but with image input

Standard Return Structure (All Modalities)

Success Response

{
    "success": True,
    # Modality-specific content field(s)
    "usage": { /* modality-specific usage metrics */ },
    "metadata": { /* modality-specific metadata */ },
    "request_id": str,           # Unique request identifier
    "timestamp": str,            # ISO 8601 timestamp
    "user_id": str,              # User context (if provided)
    "model": str,                # Actual model used
    "provider": str              # Actual provider used
}

Error Response (Universal)

{
    "success": False,
    "error": {
        "code": str,             # Standard error codes (see below)
        "message": str,          # User-friendly message
        "details": str           # Technical details (optional)
    },
    "request_id": str,
    "timestamp": str,
    "model": str,                # Requested model (if available)
    "user_id": str               # User context (if provided)
}

Standard Error Codes

  • MODEL_NOT_FOUND: Model identifier not recognized
  • INVALID_INPUT: Input format/type invalid for modality
  • UNSUPPORTED_OPERATION: Model doesn't support requested operation
  • RATE_LIMIT_EXCEEDED: Too many requests
  • QUOTA_EXHAUSTED: User quota exceeded
  • PROVIDER_UNAVAILABLE: Underlying provider service down
  • INPUT_TOO_LARGE: Input exceeds size limits
  • TIMEOUT: Request timed out
  • INTERNAL_ERROR: Unexpected internal error

Integration Guidelines

For Platform Developers

  1. Implement all functions as async to support concurrent operations
  2. Include comprehensive input validation with clear error messages
  3. Support both URL and base64 inputs for media files where applicable
  4. Implement proper timeout handling (30-300 seconds based on modality)
  5. Include usage tracking for billing and monitoring
  6. Support user context propagation via user_id parameter

For Agent System Integration (Hermes Agent)

  1. Use consistent calling patterns across all modalities
  2. Handle errors uniformly regardless of underlying model type
  3. Cache results appropriately based on modality characteristics
  4. Implement retry logic for transient failures
  5. Respect rate limits and quota constraints per user

Example Usage Patterns

Vision Analysis

result = await local_vision_inference(
    model="gpt-4-vision",
    image="https://example.com/image.jpg",
    prompt="Describe this image in detail",
    temperature=0.2
)

Generate and Speak

# Generate response
text_result = await local_llm_inference(
    model="qwen3-max",
    messages=[{"role": "user", "content": "Tell me a joke"}]
)

# Convert to speech
if text_result["success"]:
    audio_result = await local_tts_inference(
        model="elevenlabs",
        text=text_result["content"],
        voice="Rachel"
    )

This specification enables consistent integration of diverse multimodal AI capabilities while accommodating the unique requirements of each modality.