Skip to content

Image and Audio APIs ​

CoreRouter supports OpenAI-compatible image and audio interfaces. Actual capabilities, sizes, voices, file sizes, and format limits depend on models and channels; refer to console descriptions.

API Overview ​

CapabilityEndpointRequest FormatKey Fields
Image Generation/v1/images/generationsJSONmodel, prompt, size, n
Image Editing/v1/images/editsmultipart/form-data or JSONmodel, image, prompt
Legacy Image Edit Compatibility Path/v1/editsJSONLegacy client compatibility; prefer /v1/images/edits when uploading files
Audio Transcription/v1/audio/transcriptionsmultipart/form-datamodel, file
Audio Translation/v1/audio/translationsmultipart/form-datamodel, file
Text-to-Speech/v1/audio/speechJSONmodel, input, voice

Image Generation ​

bash
export COREROUTER_API_KEY="sk-xxxxxxxxxxxxxxxx"

curl https://api.corerouter.cloud/v1/images/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COREROUTER_API_KEY" \
  -d '{
    "model": "image-model-id",
    "prompt": "A clean tech product documentation cover, dark background, blue-green light effects",
    "size": "1024x1024",
    "n": 1
  }'

Image Editing ​

Image editing typically uses multipart/form-data:

bash
curl https://api.corerouter.cloud/v1/images/edits \
  -H "Authorization: Bearer $COREROUTER_API_KEY" \
  -F "model=image-model-id" \
  -F "[email protected]" \
  -F "prompt=Change background to dark tech style, keep subject"

/v1/images/edits can also use JSON per client protocol, but images must be passed as URLs or Base64 data supported by that protocol. Don't convert the same multipart request to /v1/edits; the legacy path won't process it as a new multipart image edit flow.

Audio Transcription ​

bash
curl https://api.corerouter.cloud/v1/audio/transcriptions \
  -H "Authorization: Bearer $COREROUTER_API_KEY" \
  -F "model=audio-model-id" \
  -F "[email protected]"

Audio Translation ​

bash
curl https://api.corerouter.cloud/v1/audio/translations \
  -H "Authorization: Bearer $COREROUTER_API_KEY" \
  -F "model=audio-model-id" \
  -F "[email protected]"

Text-to-Speech ​

bash
curl https://api.corerouter.cloud/v1/audio/speech \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $COREROUTER_API_KEY" \
  -d '{
    "model": "tts-model-id",
    "voice": "alloy",
    "input": "Hello, welcome to CoreRouter."
  }' \
  --output speech.mp3

Usage Recommendations ​

  • Image count, size, quality, and audio duration all affect costs; test on a small scale first.
  • n in image generation indicates the number of images. Don't allow users to pass unlimited n in your application.
  • Image sizes use the letter x, e.g., 1024x1024, not the multiplication sign ×.
  • When providing upload interfaces, limit file size, format, and duration.
  • Don't reuse chat model IDs for image and audio models unless the console explicitly indicates support.
  • If SDK wrapping of media interfaces is incomplete, directly use curl or HTTP clients.
  • TTS returns are typically audio bytes; curl needs --output to save files.
  • Voice recognition and translation typically require file upload, not just passing file path strings to JSON.

Common Issues ​

  • 400: Check if multipart/form-data field names are correct and file paths exist.
  • 415: Check if file format is supported by the model.
  • Empty file returned: Confirm command has --output and check if response Header is audio content.
  • Abnormal costs: Check if n, size, audio duration, and quality parameters meet expectations.

Released under the MIT License.