API Reference
Cumulus provides an OpenAI-compatible REST API with access to language, vision, image, video, and audio models — all at high throughput on dedicated GPU infrastructure.
https://api.ionrouter.io/v1https://kimi.ionrouter.io/v1https://minimax.ionrouter.io/v1Getting Started
From zero to first API call in four steps.
curl https://api.ionrouter.io/v1/chat/completions \
-H "Authorization: Bearer sk-your-key-here" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-30b-a3b",
"messages": [{"role": "user", "content": "Hello!"}]
}'Authentication
All API requests require a Bearer token in the Authorization header.
Authorization: Bearer sk-your-api-keyModels
Pass the model ID in the model field of your request. Prices are per million tokens unless noted otherwise.
https://api.ionrouter.io/v1https://kimi.ionrouter.io/v1https://minimax.ionrouter.io/v1Kimi
Frontier reasoning model from MoonShot AI. Endpoint: https://kimi.ionrouter.io/v1
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Kimi-K2.5featured | kimi-k2.5 | $0.20 | $1.60 | ~120 tok/s |
MiniMax
MiniMax-M2.5 — 1M context language model. Endpoint: https://minimax.ionrouter.io/v1
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| MiniMax-M2.5featured | minimax-m2.5 | $0.40 | $1.50 | ~120 tok/s |
Language
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Qwen3.5-122B-A10Bfeatured | qwen3.5-122b-a10b | $0.20 | $1.60 | ~120 tok/s |
| GPT-OSS-120Bfeatured | gpt-oss-120b | $0.020 | $0.095 | ~100 tok/s |
| Qwen3.5-35B-A3Bfeatured | qwen3.5-35b-a3b | $0.125 | $1.00 | ~90 tok/s |
| Qwen3.5-27Bfeatured | qwen3.5-27b | $0.15 | $1.20 | ~90 tok/s |
| Qwen3-30B-A3B | qwen3-30b-a3b | $0.040 | $0.14 | ~90 tok/s |
| Qwen2.5-32B | qwen2.5-32b | $0.075 | $0.30 | ~45 tok/s |
| DeepSeek-R1-14B | deepseek-r1-14b | $0.075 | $0.075 | ~75 tok/s |
| Qwen3-14B | qwen3-14b | $0.030 | $0.12 | ~80 tok/s |
| Qwen2.5-14B | qwen2.5-14b | $0.030 | $0.12 | ~80 tok/s |
| Qwen3-8B | qwen3-8b | $0.025 | $0.20 | ~150 tok/s |
| Qwen2.5-7B | qwen2.5-7b | $0.025 | $0.10 | ~160 tok/s |
Vision
Dedicated vision-language models. The Qwen3.5 language models also accept image inputs — pass them via the same messages array format.
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Molmo2-8Bfeatured | molmo2-8b | $0.040 | $0.25 | ~120 tok/s |
| InternVL3.5-8Bfeatured | internvl3.5-8b | $0.040 | $0.25 | ~120 tok/s |
| Qwen3-VL-8Bfeatured | qwen3-vl-8b | $0.040 | $0.25 | ~129 tok/s |
| Qwen2.5-VL-7B | qwen2.5-vl-7b | $0.040 | $0.25 | ~130 tok/s |
Image Generation
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Flux Schnellfeatured | flux-schnell | ~$0.005 | per image | ~3s/image |
| Flux Dev | flux-dev | ~$0.025 | per image | ~15s/image |
Video Generation
Video models are billed per GPU·second of compute time.
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Wan2.2 Text-to-Videofeatured | wan2.2-t2v-general | $0.00194 | $0.00194 / GPU·sec | ~8s/clip |
| Wan2.2 Image-to-Videofeatured | wan2.2-i2v-general | $0.00194 | $0.00194 / GPU·sec | ~8s/clip |
| HunyuanVideo | hunyuanvideo | $0.00194 | $0.00194 / GPU·sec | ~60s/clip |
| FastWan Text-to-Video 14B | fastwan-t2v-14b | $0.00194 | $0.00194 / GPU·sec | ~10s/clip |
| FastWan Text+Image-to-Video 5B | fastwan-ti2v-5b | $0.00194 | $0.00194 / GPU·sec | ~8s/clip |
| FastWan Text-to-Video 1.3B | fastwan-t2v-1.3b | $0.00194 | $0.00194 / GPU·sec | ~5s/clip |
| FramePack Image-to-Video | framepack-i2v | $0.00194 | $0.00194 / GPU·sec | ~30s/clip |
| LTX Video | ltx-video | $0.00194 | $0.00194 / GPU·sec | ~25s/clip |
Audio / TTS
Audio models are billed per minute of generated audio.
| Model | ID | Input | Output | Speed |
|---|---|---|---|---|
| Orpheus 3Bfeatured | orpheus-3b | ~$0.006 | per min | realtime |
| Dia 1.6B | dia-1.6b | ~$0.004 | per min | realtime |
| F5-TTS | f5-tts | ~$0.003 | per min | realtime |
Endpoints
Chat / LLM
OpenAI-compatible chat completions. Supports streaming via SSE.
POST https://api.ionrouter.io/v1/chat/completions| Parameter | Type | Description |
|---|---|---|
| modelreq | string | Model ID (e.g. qwen3-30b-a3b) |
| messagesreq | array | Array of message objects with role and content |
| stream | boolean | Enable SSE streaming. Default: false |
| max_tokens | integer | Maximum tokens to generate |
| temperature | number | Sampling temperature 0–2. Default: 1 |
| top_p | number | Nucleus sampling probability. Default: 1 |
| system_prompt | string | System prompt (alternative to messages[0].role=system) |
from openai import OpenAI
client = OpenAI(
api_key="sk-your-key-here",
base_url="https://api.ionrouter.io/v1",
)
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[{"role": "user", "content": "Explain transformers in one paragraph."}],
stream=True,
)
for chunk in response:
print(chunk.choices[0].delta.content or "", end="", flush=True)Vision
Use the same /v1/chat/completions endpoint with image content in the messages array. Use a vision-capable model ID.
from openai import OpenAI
client = OpenAI(
api_key="sk-your-key-here",
base_url="https://api.ionrouter.io/v1",
)
response = client.chat.completions.create(
model="qwen3-vl-8b",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
],
}],
)
print(response.choices[0].message.content)Image Generation
Generate images from text prompts using Flux models.
POST https://api.ionrouter.io/v1/images/generations| Parameter | Type | Description |
|---|---|---|
| modelreq | string | flux-schnell or flux-dev |
| promptreq | string | Text description of the image to generate |
| width | integer | Image width in pixels. Default: 1024 |
| height | integer | Image height in pixels. Default: 1024 |
| num_inference_steps | integer | Denoising steps. Default: 20 |
| guidance_scale | number | Classifier-free guidance scale. Default: 7 |
| seed | integer | Random seed for reproducibility |
| negative_prompt | string | Things to exclude from the generated image |
import requests, base64
response = requests.post(
"https://api.ionrouter.io/v1/images/generations",
headers={"Authorization": "Bearer sk-your-key-here"},
json={
"model": "flux-schnell",
"prompt": "A photorealistic mountain landscape at golden hour",
"width": 1024,
"height": 1024,
},
)
# Response contains a base64-encoded PNG
data = response.json()["data"][0]["b64_json"]
with open("output.png", "wb") as f:
f.write(base64.b64decode(data))Video Generation
Video generation is asynchronous. Submit a job, then poll for the result. Billed per GPU·second of compute time at $0.00194 / GPU·sec.
POST https://api.ionrouter.io/v1/video/generations
GET https://api.ionrouter.io/v1/video/generations/{job_id}| Parameter | Type | Description |
|---|---|---|
| modelreq | string | Video model ID (e.g. wan2.2-t2v-general) |
| promptreq | string | Text description of the video to generate |
| image_url | string | Input image URL for image-to-video models |
| height | integer | Video height in pixels |
| width | integer | Video width in pixels |
| num_frames | integer | Number of frames to generate |
| num_inference_steps | integer | Denoising steps |
| guidance_scale | number | Guidance scale |
import requests, time
# Submit job
resp = requests.post(
"https://api.ionrouter.io/v1/video/generations",
headers={"Authorization": "Bearer sk-your-key-here"},
json={
"model": "wan2.2-t2v-general",
"prompt": "A serene lake at sunset with gentle ripples",
},
)
job_id = resp.json()["id"]
# Poll for result
while True:
status = requests.get(
f"https://api.ionrouter.io/v1/video/generations/{job_id}",
headers={"Authorization": "Bearer sk-your-key-here"},
).json()
if status["status"] == "succeeded":
print("Video URL:", status["output"]["video_url"])
break
elif status["status"] == "failed":
print("Error:", status.get("error"))
break
time.sleep(3)Audio / TTS
Convert text to speech. Returns an audio file. Billed per minute of generated audio.
POST https://api.ionrouter.io/v1/audio/speech| Parameter | Type | Description |
|---|---|---|
| modelreq | string | TTS model ID (orpheus-3b, dia-1.6b, f5-tts) |
| inputreq | string | Text to synthesize |
| voice | string | Voice preset (model-dependent) |
| ref_audio | string | Base64 audio for voice cloning (f5-tts) |
| ref_text | string | Transcript of ref_audio (f5-tts) |
import requests
response = requests.post(
"https://api.ionrouter.io/v1/audio/speech",
headers={"Authorization": "Bearer sk-your-key-here"},
json={
"model": "orpheus-3b",
"input": "Welcome to Cumulus. Fast inference for everyone.",
"voice": "tara",
},
)
with open("output.wav", "wb") as f:
f.write(response.content)SDKs
Cumulus is fully compatible with the OpenAI SDK. Set base_url to the endpoint for the model you're using (api.ionrouter.io/v1, kimi.ionrouter.io/v1, or minimax.ionrouter.io/v1) and pass your Cumulus API key.
pip install openainpm install openaiRate Limits
| Limit | Value | Scope |
|---|---|---|
| Requests | 100 / minute | Per API key |
| Concurrent streams | Unlimited | Per API key |
| Max tokens (chat) | 128,000 | Per request |
| Max image size | 20 MB | Per request |
Rate limit errors return HTTP 429. If you need higher limits, email founders@cumuluslabs.io.
Error Codes
| Status | Meaning | Common cause |
|---|---|---|
| 400 | Bad Request | Invalid or missing model, malformed JSON |
| 401 | Unauthorized | Missing or invalid API key |
| 402 | Payment Required | Insufficient credits — add funds at /billing |
| 422 | Unprocessable Entity | Invalid parameter values |
| 429 | Too Many Requests | Rate limit exceeded |
| 500 | Internal Server Error | Upstream model error — retry with backoff |
| 503 | Service Unavailable | Model temporarily offline |