Documentation

API Reference

Cumulus provides an OpenAI-compatible REST API with access to language, vision, image, video, and audio models — all at high throughput on dedicated GPU infrastructure.

APIhttps://api.ionrouter.io/v1
Kimihttps://kimi.ionrouter.io/v1
MiniMaxhttps://minimax.ionrouter.io/v1

Getting Started

From zero to first API call in four steps.

01
Create an account
Sign up at ionrouter.io
02
Add credits
Go to Billing and add funds to your account. Credits are consumed as you use the API.
03
Generate an API key
Visit Keys, create a key, and copy it — it's only shown once.
04
Make your first request
Pass your key as a Bearer token in the Authorization header.
bash
curl https://api.ionrouter.io/v1/chat/completions \
  -H "Authorization: Bearer sk-your-key-here" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-30b-a3b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Authentication

All API requests require a Bearer token in the Authorization header.

http
Authorization: Bearer sk-your-api-key
FormatStarts with sk-, minimum 10 characters
ScopeKeys are scoped to your account and its credit balance
RotationGenerate a new key from the Keys page at any time
SecurityKeys are hashed on our servers — we never store the raw value

Models

Pass the model ID in the model field of your request. Prices are per million tokens unless noted otherwise.

Most models https://api.ionrouter.io/v1
Kimi K2.5 https://kimi.ionrouter.io/v1
MiniMax-M2.5 https://minimax.ionrouter.io/v1

Kimi

Frontier reasoning model from MoonShot AI. Endpoint: https://kimi.ionrouter.io/v1

ModelIDInputOutputSpeed
Kimi-K2.5featuredkimi-k2.5$0.20$1.60~120 tok/s

MiniMax

MiniMax-M2.5 — 1M context language model. Endpoint: https://minimax.ionrouter.io/v1

ModelIDInputOutputSpeed
MiniMax-M2.5featuredminimax-m2.5$0.40$1.50~120 tok/s

Language

ModelIDInputOutputSpeed
Qwen3.5-122B-A10Bfeaturedqwen3.5-122b-a10b$0.20$1.60~120 tok/s
GPT-OSS-120Bfeaturedgpt-oss-120b$0.020$0.095~100 tok/s
Qwen3.5-35B-A3Bfeaturedqwen3.5-35b-a3b$0.125$1.00~90 tok/s
Qwen3.5-27Bfeaturedqwen3.5-27b$0.15$1.20~90 tok/s
Qwen3-30B-A3Bqwen3-30b-a3b$0.040$0.14~90 tok/s
Qwen2.5-32Bqwen2.5-32b$0.075$0.30~45 tok/s
DeepSeek-R1-14Bdeepseek-r1-14b$0.075$0.075~75 tok/s
Qwen3-14Bqwen3-14b$0.030$0.12~80 tok/s
Qwen2.5-14Bqwen2.5-14b$0.030$0.12~80 tok/s
Qwen3-8Bqwen3-8b$0.025$0.20~150 tok/s
Qwen2.5-7Bqwen2.5-7b$0.025$0.10~160 tok/s

Vision

Dedicated vision-language models. The Qwen3.5 language models also accept image inputs — pass them via the same messages array format.

ModelIDInputOutputSpeed
Molmo2-8Bfeaturedmolmo2-8b$0.040$0.25~120 tok/s
InternVL3.5-8Bfeaturedinternvl3.5-8b$0.040$0.25~120 tok/s
Qwen3-VL-8Bfeaturedqwen3-vl-8b$0.040$0.25~129 tok/s
Qwen2.5-VL-7Bqwen2.5-vl-7b$0.040$0.25~130 tok/s

Image Generation

ModelIDInputOutputSpeed
Flux Schnellfeaturedflux-schnell~$0.005per image~3s/image
Flux Devflux-dev~$0.025per image~15s/image

Video Generation

Video models are billed per GPU·second of compute time.

ModelIDInputOutputSpeed
Wan2.2 Text-to-Videofeaturedwan2.2-t2v-general$0.00194$0.00194 / GPU·sec~8s/clip
Wan2.2 Image-to-Videofeaturedwan2.2-i2v-general$0.00194$0.00194 / GPU·sec~8s/clip
HunyuanVideohunyuanvideo$0.00194$0.00194 / GPU·sec~60s/clip
FastWan Text-to-Video 14Bfastwan-t2v-14b$0.00194$0.00194 / GPU·sec~10s/clip
FastWan Text+Image-to-Video 5Bfastwan-ti2v-5b$0.00194$0.00194 / GPU·sec~8s/clip
FastWan Text-to-Video 1.3Bfastwan-t2v-1.3b$0.00194$0.00194 / GPU·sec~5s/clip
FramePack Image-to-Videoframepack-i2v$0.00194$0.00194 / GPU·sec~30s/clip
LTX Videoltx-video$0.00194$0.00194 / GPU·sec~25s/clip

Audio / TTS

Audio models are billed per minute of generated audio.

ModelIDInputOutputSpeed
Orpheus 3Bfeaturedorpheus-3b~$0.006per minrealtime
Dia 1.6Bdia-1.6b~$0.004per minrealtime
F5-TTSf5-tts~$0.003per minrealtime

Endpoints

Chat / LLM

OpenAI-compatible chat completions. Supports streaming via SSE.

http
POST https://api.ionrouter.io/v1/chat/completions
ParameterTypeDescription
modelreqstringModel ID (e.g. qwen3-30b-a3b)
messagesreqarrayArray of message objects with role and content
streambooleanEnable SSE streaming. Default: false
max_tokensintegerMaximum tokens to generate
temperaturenumberSampling temperature 0–2. Default: 1
top_pnumberNucleus sampling probability. Default: 1
system_promptstringSystem prompt (alternative to messages[0].role=system)
python
from openai import OpenAI

client = OpenAI(
    api_key="sk-your-key-here",
    base_url="https://api.ionrouter.io/v1",
)

response = client.chat.completions.create(
    model="qwen3-30b-a3b",
    messages=[{"role": "user", "content": "Explain transformers in one paragraph."}],
    stream=True,
)

for chunk in response:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Vision

Use the same /v1/chat/completions endpoint with image content in the messages array. Use a vision-capable model ID.

python
from openai import OpenAI

client = OpenAI(
    api_key="sk-your-key-here",
    base_url="https://api.ionrouter.io/v1",
)

response = client.chat.completions.create(
    model="qwen3-vl-8b",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
        ],
    }],
)

print(response.choices[0].message.content)

Image Generation

Generate images from text prompts using Flux models.

http
POST https://api.ionrouter.io/v1/images/generations
ParameterTypeDescription
modelreqstringflux-schnell or flux-dev
promptreqstringText description of the image to generate
widthintegerImage width in pixels. Default: 1024
heightintegerImage height in pixels. Default: 1024
num_inference_stepsintegerDenoising steps. Default: 20
guidance_scalenumberClassifier-free guidance scale. Default: 7
seedintegerRandom seed for reproducibility
negative_promptstringThings to exclude from the generated image
python
import requests, base64

response = requests.post(
    "https://api.ionrouter.io/v1/images/generations",
    headers={"Authorization": "Bearer sk-your-key-here"},
    json={
        "model": "flux-schnell",
        "prompt": "A photorealistic mountain landscape at golden hour",
        "width": 1024,
        "height": 1024,
    },
)

# Response contains a base64-encoded PNG
data = response.json()["data"][0]["b64_json"]
with open("output.png", "wb") as f:
    f.write(base64.b64decode(data))

Video Generation

Video generation is asynchronous. Submit a job, then poll for the result. Billed per GPU·second of compute time at $0.00194 / GPU·sec.

http
POST https://api.ionrouter.io/v1/video/generations
GET  https://api.ionrouter.io/v1/video/generations/{job_id}
ParameterTypeDescription
modelreqstringVideo model ID (e.g. wan2.2-t2v-general)
promptreqstringText description of the video to generate
image_urlstringInput image URL for image-to-video models
heightintegerVideo height in pixels
widthintegerVideo width in pixels
num_framesintegerNumber of frames to generate
num_inference_stepsintegerDenoising steps
guidance_scalenumberGuidance scale
python
import requests, time

# Submit job
resp = requests.post(
    "https://api.ionrouter.io/v1/video/generations",
    headers={"Authorization": "Bearer sk-your-key-here"},
    json={
        "model": "wan2.2-t2v-general",
        "prompt": "A serene lake at sunset with gentle ripples",
    },
)
job_id = resp.json()["id"]

# Poll for result
while True:
    status = requests.get(
        f"https://api.ionrouter.io/v1/video/generations/{job_id}",
        headers={"Authorization": "Bearer sk-your-key-here"},
    ).json()
    if status["status"] == "succeeded":
        print("Video URL:", status["output"]["video_url"])
        break
    elif status["status"] == "failed":
        print("Error:", status.get("error"))
        break
    time.sleep(3)

Audio / TTS

Convert text to speech. Returns an audio file. Billed per minute of generated audio.

http
POST https://api.ionrouter.io/v1/audio/speech
ParameterTypeDescription
modelreqstringTTS model ID (orpheus-3b, dia-1.6b, f5-tts)
inputreqstringText to synthesize
voicestringVoice preset (model-dependent)
ref_audiostringBase64 audio for voice cloning (f5-tts)
ref_textstringTranscript of ref_audio (f5-tts)
python
import requests

response = requests.post(
    "https://api.ionrouter.io/v1/audio/speech",
    headers={"Authorization": "Bearer sk-your-key-here"},
    json={
        "model": "orpheus-3b",
        "input": "Welcome to Cumulus. Fast inference for everyone.",
        "voice": "tara",
    },
)

with open("output.wav", "wb") as f:
    f.write(response.content)

SDKs

Cumulus is fully compatible with the OpenAI SDK. Set base_url to the endpoint for the model you're using (api.ionrouter.io/v1, kimi.ionrouter.io/v1, or minimax.ionrouter.io/v1) and pass your Cumulus API key.

Python
bash
pip install openai
Node.js
bash
npm install openai

Rate Limits

LimitValueScope
Requests100 / minutePer API key
Concurrent streamsUnlimitedPer API key
Max tokens (chat)128,000Per request
Max image size20 MBPer request

Rate limit errors return HTTP 429. If you need higher limits, email founders@cumuluslabs.io.

Error Codes

StatusMeaningCommon cause
400Bad RequestInvalid or missing model, malformed JSON
401UnauthorizedMissing or invalid API key
402Payment RequiredInsufficient credits — add funds at /billing
422Unprocessable EntityInvalid parameter values
429Too Many RequestsRate limit exceeded
500Internal Server ErrorUpstream model error — retry with backoff
503Service UnavailableModel temporarily offline