AI models & pricing
Explore model IDs, supported API endpoints and credit rates across VoidAI's public plans. Compare the models for your application before you start building.
How to read pricing. Token rates are credits per 1,000 tokens. Input and output are priced separately; cache read and write factors multiply the applicable input rate. A listed base cost is shown separately. Fixed-price models show their base fixed cost in credits.
Long-context rates and token limits appear only when supplied by the catalog. “Not supplied” means the catalog does not provide that value. See the API documentation for request and billing details.
Catalog listings are not live service health. Plan listings do not guarantee current model availability or individual access. This catalog refreshes periodically and may be cached for five minutes. Check VoidAI service status for live health information.
Claude access requires approval and is limited to approved use cases. Subject to provider policy.
Explore the model catalog
Showing 90 of 90 models
Claude
claude-fable-5
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 7,500
- Output · credits / 1k tokens
- 30,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-haiku-4-5-20251001
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 400
- Output · credits / 1k tokens
- 2,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-4-1-20250805
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 6,000
- Output · credits / 1k tokens
- 30,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-4-5-20251101
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-4-6
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-4-7
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-4-8
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-5
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
- Context limit:
- 1,000,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-opus-5-5
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 10,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
- Context limit:
- 1,000,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-sonnet-4-5-20250929
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,200
- Output · credits / 1k tokens
- 6,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-sonnet-4-6
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,200
- Output · credits / 1k tokens
- 6,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Claude
claude-sonnet-5
Anthropic · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses - Messages
/v1/messages - Token counting
/v1/messages/count_tokens
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 5,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- basic
- premium
- pro
- ultra
- enterprise
- Chat
DeepSeek
deepseek-v3.2
DeepSeek · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 200
- Output · credits / 1k tokens
- 300
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
DeepSeek
deepseek-v4-flash
DeepSeek · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 350
- Output · credits / 1k tokens
- 650
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
DeepSeek
deepseek-v4-flash-0731
DeepSeek · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 350
- Output · credits / 1k tokens
- 650
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
DeepSeek
deepseek-v4-pro
DeepSeek · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 500
- Output · credits / 1k tokens
- 1,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
DeepSeek
deepseek-v4-pro-0813
DeepSeek · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 500
- Output · credits / 1k tokens
- 1,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
FLUX
flux-kontext-max
Black Forest Labs · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 3,000
Listed public plans
- premium
- pro
- ultra
- enterprise
- Image generation
FLUX
flux-kontext-pro
Black Forest Labs · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 3,000
Listed public plans
- premium
- pro
- ultra
- enterprise
- Image generation
Gemini
gemini-2.5-flash
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 220
- Output · credits / 1k tokens
- 1,900
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-2.5-flash-image
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 220
- Output · credits / 1k tokens
- 1,900
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-2.5-pro
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 950
- Output · credits / 1k tokens
- 7,500
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3-flash-preview
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 380
- Output · credits / 1k tokens
- 2,300
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3-pro-image
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,500
- Output · credits / 1k tokens
- 9,000
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.1-flash-image
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 380
- Output · credits / 1k tokens
- 2,300
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.1-flash-lite
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 200
- Output · credits / 1k tokens
- 1,100
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.1-pro-preview
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,500
- Output · credits / 1k tokens
- 9,000
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.5-flash
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,100
- Output · credits / 1k tokens
- 6,800
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.5-flash-lite
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 220
- Output · credits / 1k tokens
- 1,900
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.6-flash
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 5,000
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.7-flash
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 5,000
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Gemini
gemini-3.8-flash
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 5,000
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GLM
glm-5.1
Z.ai · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 850
- Cache read · factor of input rate
- 0.2×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GLM
glm-5.2
Z.ai · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 850
- Cache read · factor of input rate
- 0.2×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GLM
glm-5.3
Z.ai · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 850
- Cache read · factor of input rate
- 0.2×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GLM
glm-5.3-flash
Z.ai · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 300
- Output · credits / 1k tokens
- 600
- Cache read · factor of input rate
- 0.2×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Google
gemma-4-31b-it
Google · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 40
- Output · credits / 1k tokens
- 160
- Cache read · factor of input rate
- 0.25×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-4.1
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 800
- Output · credits / 1k tokens
- 3,200
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-4.1-mini
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 60
- Output · credits / 1k tokens
- 240
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-4o
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 4,000
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-4o-mini
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 60
- Output · credits / 1k tokens
- 240
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-4o-mini-transcribe
OpenAI · Fixed pricing
Endpoints and capabilities
- Transcription
/v1/audio/transcriptions
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 20
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Transcription
GPT
gpt-4o-mini-tts
OpenAI · Fixed pricing
Endpoints and capabilities
- Speech
/v1/audio/speech
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 250
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Speech
GPT
gpt-4o-transcribe
OpenAI · Fixed pricing
Endpoints and capabilities
- Transcription
/v1/audio/transcriptions
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 50
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Transcription
GPT
gpt-5
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 800
- Output · credits / 1k tokens
- 3,200
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5-mini
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 60
- Output · credits / 1k tokens
- 240
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5-nano
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 20
- Output · credits / 1k tokens
- 80
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.1
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 800
- Output · credits / 1k tokens
- 3,200
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.2
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 800
- Output · credits / 1k tokens
- 3,200
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.3-codex
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,200
- Output · credits / 1k tokens
- 4,000
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
- Context limit:
- 400,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.4
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,600
- Output · credits / 1k tokens
- 4,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 1,600
- Output · credits / 1k
- 3,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.4-mini
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 90
- Output · credits / 1k tokens
- 300
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
- Context limit:
- 400,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.4-nano
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 80
- Output · credits / 1k tokens
- 250
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
- Context limit:
- 400,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.4-pro
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 20,000
- Output · credits / 1k tokens
- 60,000
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 25,000
Long context · above 272,000 input tokens
- Input · credits / 1k
- 40,000
- Output · credits / 1k
- 90,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.5
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 3,200
- Output · credits / 1k tokens
- 8,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 3,200
- Output · credits / 1k
- 6,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.5-pro
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 20,000
- Output · credits / 1k tokens
- 60,000
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 25,000
Long context · above 272,000 input tokens
- Input · credits / 1k
- 40,000
- Output · credits / 1k
- 90,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.6-luna
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 300
- Output · credits / 1k tokens
- 1,200
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 600
- Output · credits / 1k
- 1,800
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.6-sol
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 6,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 4,000
- Output · credits / 1k
- 9,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-5.6-terra
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 1,000
- Output · credits / 1k tokens
- 3,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 2,000
- Output · credits / 1k
- 4,500
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-6-astra
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 7,500
- Output · credits / 1k tokens
- 30,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-6-luna
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 300
- Output · credits / 1k tokens
- 1,200
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 600
- Output · credits / 1k
- 1,800
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-6-sol
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 6,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 4,000
- Output · credits / 1k
- 9,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-6.1-sol
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 2,000
- Output · credits / 1k tokens
- 6,000
- Cache read · factor of input rate
- 0.1×
- Cache write · factor of input rate
- 1.25×
- Base cost · credits
- 0
Long context · above 272,000 input tokens
- Input · credits / 1k
- 4,000
- Output · credits / 1k
- 9,000
- Context limit:
- 1,050,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-image-1
OpenAI · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations - Image editing
/v1/images/edits
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 12,000
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Image generation
GPT
gpt-image-1.5
OpenAI · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations - Image editing
/v1/images/edits
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 8,000
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Image generation
GPT
gpt-image-2
OpenAI · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations - Image editing
/v1/images/edits
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 10,000
Listed public plans
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Image generation
GPT
gpt-oss-120b
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 40
- Output · credits / 1k tokens
- 160
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
GPT
gpt-oss-20b
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 20
- Output · credits / 1k tokens
- 80
- Cache read · factor of input rate
- 0.5×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Kimi
kimi-k2.5
Moonshot AI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 120
- Output · credits / 1k tokens
- 600
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Kimi
kimi-k2.6
Moonshot AI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 120
- Output · credits / 1k tokens
- 600
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Kimi
kimi-k2.7-code
Moonshot AI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 120
- Output · credits / 1k tokens
- 600
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Kimi
kimi-k3
Moonshot AI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 3,000
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
- Context limit:
- 1,000,000 tokens
- Output limit:
- 128,000 tokens
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Midjourney
midjourney
Midjourney · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 75,000
Listed public plans
- ultra
- enterprise
- Image generation
OpenAI
omni-moderation-latest
OpenAI · Fixed pricing
Endpoints and capabilities
- Moderation
/v1/moderations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Moderation
OpenAI embeddings
text-embedding-3-large
OpenAI · Fixed pricing
Endpoints and capabilities
- Embeddings
/v1/embeddings
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 50
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Embeddings
OpenAI embeddings
text-embedding-3-small
OpenAI · Fixed pricing
Endpoints and capabilities
- Embeddings
/v1/embeddings
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 50
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Embeddings
OpenAI o-series
o3
OpenAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: Not supportedTool calling: Not supported- Input · credits / 1k tokens
- 800
- Output · credits / 1k tokens
- 3,200
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- basic
- premium
- pro
- ultra
- enterprise
- Chat
OpenAI TTS
tts-1
OpenAI · Fixed pricing
Endpoints and capabilities
- Speech
/v1/audio/speech
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 75
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Speech
OpenAI TTS
tts-1-hd
OpenAI · Fixed pricing
Endpoints and capabilities
- Speech
/v1/audio/speech
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 150
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Speech
Qwen
qwen3-235b-a22b-instruct
Qwen · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 50
- Output · credits / 1k tokens
- 500
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Qwen
qwen3-coder-480b-a35b-instruct
Qwen · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 50
- Output · credits / 1k tokens
- 500
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Qwen
qwen3.8-2.4t-a95b
Qwen · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 50
- Output · credits / 1k tokens
- 500
- Cache read · factor of input rate
- 0.1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Qwen
qwen3.8-27b:free
Qwen · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 0
- Output · credits / 1k tokens
- 0
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Recraft
recraft-v3
Recraft · Fixed pricing
Endpoints and capabilities
- Image generation
/v1/images/generations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 1,000
Listed public plans
- premium
- pro
- ultra
- enterprise
- Image generation
Sonar
sonar
Perplexity · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 200
- Output · credits / 1k tokens
- 200
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Sonar
sonar-deep-research
Perplexity · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 400
- Output · credits / 1k tokens
- 1,600
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Sonar
sonar-pro
Perplexity · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 3,000
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Sonar
sonar-reasoning-pro
Perplexity · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 600
- Output · credits / 1k tokens
- 3,000
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
VoidAI
umbra
VoidAI · Token pricing
Endpoints and capabilities
- Chat
/v1/chat/completions - Responses
/v1/responses
Streaming: SupportedTool calling: Supported- Input · credits / 1k tokens
- 20
- Output · credits / 1k tokens
- 100
- Cache read · factor of input rate
- 1×
- Base cost · credits
- 0
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Chat
Whisper
whisper-1
OpenAI · Fixed pricing
Endpoints and capabilities
- Transcription
/v1/audio/transcriptions - Audio translation
/v1/audio/translations
Streaming: Not supportedTool calling: Not supported- Base fixed cost · credits
- 10
Listed public plans
- free
- economy
- basic
- premium
- pro
- ultra
- enterprise
- Transcription