Public model catalog

AI models & pricing

Explore model IDs, supported API endpoints and credit rates across VoidAI's public plans. Compare the models for your application before you start building.

How to read pricing. Token rates are credits per 1,000 tokens. Input and output are priced separately; cache read and write factors multiply the applicable input rate. A listed base cost is shown separately. Fixed-price models show their base fixed cost in credits.

Long-context rates and token limits appear only when supplied by the catalog. “Not supplied” means the catalog does not provide that value. See the API documentation for request and billing details.

Catalog listings are not live service health. Plan listings do not guarantee current model availability or individual access. This catalog refreshes periodically and may be cached for five minutes. Check VoidAI service status for live health information.

Claude access requires approval and is limited to approved use cases. Subject to provider policy.

Explore the model catalog

Showing 90 of 90 models

  • Claude

    claude-fable-5

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    7,500
    Output · credits / 1k tokens
    30,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-haiku-4-5-20251001

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    400
    Output · credits / 1k tokens
    2,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-4-1-20250805

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    6,000
    Output · credits / 1k tokens
    30,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-4-5-20251101

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-4-6

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-4-7

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-4-8

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-5

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0
    Context limit:
    1,000,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-opus-5-5

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    10,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0
    Context limit:
    1,000,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-sonnet-4-5-20250929

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,200
    Output · credits / 1k tokens
    6,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-sonnet-4-6

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,200
    Output · credits / 1k tokens
    6,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Claude

    claude-sonnet-5

    Anthropic · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    • Messages/v1/messages
    • Token counting/v1/messages/count_tokens
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    5,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • DeepSeek

    deepseek-v3.2

    DeepSeek · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    200
    Output · credits / 1k tokens
    300
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • DeepSeek

    deepseek-v4-flash

    DeepSeek · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    350
    Output · credits / 1k tokens
    650
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • DeepSeek

    deepseek-v4-flash-0731

    DeepSeek · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    350
    Output · credits / 1k tokens
    650
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • DeepSeek

    deepseek-v4-pro

    DeepSeek · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    500
    Output · credits / 1k tokens
    1,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • DeepSeek

    deepseek-v4-pro-0813

    DeepSeek · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    500
    Output · credits / 1k tokens
    1,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • FLUX

    flux-kontext-max

    Black Forest Labs · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    3,000

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • FLUX

    flux-kontext-pro

    Black Forest Labs · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    3,000

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-2.5-flash

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    220
    Output · credits / 1k tokens
    1,900
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-2.5-flash-image

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    220
    Output · credits / 1k tokens
    1,900
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-2.5-pro

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    950
    Output · credits / 1k tokens
    7,500
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3-flash-preview

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    380
    Output · credits / 1k tokens
    2,300
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3-pro-image

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,500
    Output · credits / 1k tokens
    9,000
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.1-flash-image

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    380
    Output · credits / 1k tokens
    2,300
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.1-flash-lite

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    200
    Output · credits / 1k tokens
    1,100
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.1-pro-preview

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,500
    Output · credits / 1k tokens
    9,000
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.5-flash

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,100
    Output · credits / 1k tokens
    6,800
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.5-flash-lite

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    220
    Output · credits / 1k tokens
    1,900
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.6-flash

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    5,000
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.7-flash

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    5,000
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Gemini

    gemini-3.8-flash

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    5,000
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GLM

    glm-5.1

    Z.ai · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    850
    Cache read · factor of input rate
    0.2×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GLM

    glm-5.2

    Z.ai · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    850
    Cache read · factor of input rate
    0.2×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GLM

    glm-5.3

    Z.ai · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    850
    Cache read · factor of input rate
    0.2×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GLM

    glm-5.3-flash

    Z.ai · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    300
    Output · credits / 1k tokens
    600
    Cache read · factor of input rate
    0.2×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Google

    gemma-4-31b-it

    Google · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    40
    Output · credits / 1k tokens
    160
    Cache read · factor of input rate
    0.25×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4.1

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    800
    Output · credits / 1k tokens
    3,200
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4.1-mini

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    60
    Output · credits / 1k tokens
    240
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4o

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    4,000
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4o-mini

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    60
    Output · credits / 1k tokens
    240
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4o-mini-transcribe

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Transcription/v1/audio/transcriptions
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    20

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4o-mini-tts

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Speech/v1/audio/speech
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    250

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-4o-transcribe

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Transcription/v1/audio/transcriptions
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    50

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    800
    Output · credits / 1k tokens
    3,200
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5-mini

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    60
    Output · credits / 1k tokens
    240
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5-nano

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    20
    Output · credits / 1k tokens
    80
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.1

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    800
    Output · credits / 1k tokens
    3,200
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.2

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    800
    Output · credits / 1k tokens
    3,200
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.3-codex

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,200
    Output · credits / 1k tokens
    4,000
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0
    Context limit:
    400,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.4

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,600
    Output · credits / 1k tokens
    4,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    1,600
    Output · credits / 1k
    3,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.4-mini

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    90
    Output · credits / 1k tokens
    300
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0
    Context limit:
    400,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.4-nano

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    80
    Output · credits / 1k tokens
    250
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0
    Context limit:
    400,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.4-pro

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    20,000
    Output · credits / 1k tokens
    60,000
    Cache read · factor of input rate
    1×
    Base cost · credits
    25,000

    Long context · above 272,000 input tokens

    Input · credits / 1k
    40,000
    Output · credits / 1k
    90,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.5

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    3,200
    Output · credits / 1k tokens
    8,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    3,200
    Output · credits / 1k
    6,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.5-pro

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    20,000
    Output · credits / 1k tokens
    60,000
    Cache read · factor of input rate
    1×
    Base cost · credits
    25,000

    Long context · above 272,000 input tokens

    Input · credits / 1k
    40,000
    Output · credits / 1k
    90,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.6-luna

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    300
    Output · credits / 1k tokens
    1,200
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    600
    Output · credits / 1k
    1,800
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.6-sol

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    6,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    4,000
    Output · credits / 1k
    9,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-5.6-terra

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    1,000
    Output · credits / 1k tokens
    3,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    2,000
    Output · credits / 1k
    4,500
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-6-astra

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    7,500
    Output · credits / 1k tokens
    30,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-6-luna

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    300
    Output · credits / 1k tokens
    1,200
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    600
    Output · credits / 1k
    1,800
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-6-sol

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    6,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    4,000
    Output · credits / 1k
    9,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-6.1-sol

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    2,000
    Output · credits / 1k tokens
    6,000
    Cache read · factor of input rate
    0.1×
    Cache write · factor of input rate
    1.25×
    Base cost · credits
    0

    Long context · above 272,000 input tokens

    Input · credits / 1k
    4,000
    Output · credits / 1k
    9,000
    Context limit:
    1,050,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-image-1

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    • Image editing/v1/images/edits
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    12,000

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-image-1.5

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    • Image editing/v1/images/edits
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    8,000

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-image-2

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    • Image editing/v1/images/edits
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    10,000

    Listed public plans

    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-oss-120b

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    40
    Output · credits / 1k tokens
    160
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • GPT

    gpt-oss-20b

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    20
    Output · credits / 1k tokens
    80
    Cache read · factor of input rate
    0.5×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Kimi

    kimi-k2.5

    Moonshot AI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    120
    Output · credits / 1k tokens
    600
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Kimi

    kimi-k2.6

    Moonshot AI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    120
    Output · credits / 1k tokens
    600
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Kimi

    kimi-k2.7-code

    Moonshot AI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    120
    Output · credits / 1k tokens
    600
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Kimi

    kimi-k3

    Moonshot AI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    3,000
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0
    Context limit:
    1,000,000 tokens
    Output limit:
    128,000 tokens

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Midjourney

    midjourney

    Midjourney · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    75,000

    Listed public plans

    • ultra
    • enterprise
  • OpenAI

    omni-moderation-latest

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Moderation/v1/moderations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • OpenAI embeddings

    text-embedding-3-large

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Embeddings/v1/embeddings
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    50

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • OpenAI embeddings

    text-embedding-3-small

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Embeddings/v1/embeddings
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    50

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • OpenAI o-series

    o3

    OpenAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: Not supportedTool calling: Not supported
    Input · credits / 1k tokens
    800
    Output · credits / 1k tokens
    3,200
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • OpenAI TTS

    tts-1

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Speech/v1/audio/speech
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    75

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • OpenAI TTS

    tts-1-hd

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Speech/v1/audio/speech
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    150

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Qwen

    qwen3-235b-a22b-instruct

    Qwen · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    50
    Output · credits / 1k tokens
    500
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Qwen

    qwen3-coder-480b-a35b-instruct

    Qwen · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    50
    Output · credits / 1k tokens
    500
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Qwen

    qwen3.8-2.4t-a95b

    Qwen · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    50
    Output · credits / 1k tokens
    500
    Cache read · factor of input rate
    0.1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Qwen

    qwen3.8-27b:free

    Qwen · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    0
    Output · credits / 1k tokens
    0
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Recraft

    recraft-v3

    Recraft · Fixed pricing

    Endpoints and capabilities

    • Image generation/v1/images/generations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    1,000

    Listed public plans

    • premium
    • pro
    • ultra
    • enterprise
  • Sonar

    sonar

    Perplexity · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    200
    Output · credits / 1k tokens
    200
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Sonar

    sonar-deep-research

    Perplexity · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    400
    Output · credits / 1k tokens
    1,600
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Sonar

    sonar-pro

    Perplexity · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    3,000
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Sonar

    sonar-reasoning-pro

    Perplexity · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    600
    Output · credits / 1k tokens
    3,000
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • VoidAI

    umbra

    VoidAI · Token pricing

    Endpoints and capabilities

    • Chat/v1/chat/completions
    • Responses/v1/responses
    Streaming: SupportedTool calling: Supported
    Input · credits / 1k tokens
    20
    Output · credits / 1k tokens
    100
    Cache read · factor of input rate
    1×
    Base cost · credits
    0

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise
  • Whisper

    whisper-1

    OpenAI · Fixed pricing

    Endpoints and capabilities

    • Transcription/v1/audio/transcriptions
    • Audio translation/v1/audio/translations
    Streaming: Not supportedTool calling: Not supported
    Base fixed cost · credits
    10

    Listed public plans

    • free
    • economy
    • basic
    • premium
    • pro
    • ultra
    • enterprise