Overview
Estimate how many Convai credits an interaction consumes for a given character configuration. All amounts are in credits. Every request field is optional — anything omitted uses the defaults below. No authentication is required.
Estimates only. Actual billing follows your subscription plan and the published rate card at convai.com/pricing.
Rate card: version 3, effective 2026-08-19.
Endpoints
| GET /api/v1/estimate | Estimate via query parameters. Cacheable. |
| POST /api/v1/estimate | Estimate via JSON body { selections, assumptions }. |
| GET /api/v1/options | Every valid option key with label and rate, plus defaults and plans. |
| GET /api/v1/openapi.json | Machine-readable OpenAPI 3.1 spec. |
Request parameters — selections
What the character is configured with. In a POST body these live under selections; in a GET they are top-level query parameters.
| field | type | default | meaning |
|---|---|---|---|
| llm | string | gemini_2_5_flash | AI model, including its thinking level where the model has one (e.g. gpt_5_1_high). Every valid key is listed under Option keys below and returned by /options. Realtime (Live) models auto-imply bundled speech — see the Realtime section. |
| memory | boolean | false | Long-term memory across sessions. Flat 1-credit charge per interaction when true; memory content itself counts inside inputTokens. |
| knowledge | boolean | false | Knowledge bank retrieval. Flat 1-credit charge per interaction when true; retrieved content counts inside inputTokens. |
| stt | string | soniox_stt_streaming | Voice input (speech-to-text) option. Use included_in_realtime with Live models (speech then bills as the model's audio input tokens). |
| tts | string | convai_tts_managed | Voice output (text-to-speech) option. Use included_in_realtime with Live models (speech then bills as the model's audio output tokens). |
| animation | string | no_animation | Character animation: no_animation or convai_neurosync (lip-sync + facial expression, priced per second of speech). |
Request parameters — assumptions
Per-interaction quantities. In a POST body these live under assumptions; in a GET they are top-level query parameters. Estimates scale linearly: 500 input tokens bill exactly 500 × the model's per-token rate.
| field | type | default | meaning |
|---|---|---|---|
| inputTokens | number ≥ 0 | 4,000 | ALL context tokens the model reads each interaction: character prompt, personality, session memory, knowledge and scene content. Priced at the selected model's input rate. |
| outputChars | number ≥ 0 | 150 (37.5 tokens) | Response length in characters; 4 characters = 1 token. Priced at the model's output rate. Thinking models default higher because reasoning tokens bill as output — sending a drastically low value with a high thinking level understates the real cost. |
| sttSeconds | number ≥ 0 | 8 | Seconds of user speech transcribed per interaction. With a Live model + bundled speech, these seconds bill as the model's audio input tokens instead. |
| ttsSeconds | number ≥ 0 | 10 | Seconds of spoken response per interaction. Also drives animation duration. With a Live model + bundled speech, these seconds bill as audio output tokens. |
Examples
GET with query parameters:
curl "https://credit-calculator.convai.com/api/v1/estimate?llm=claude_5_sonnet_high&memory=true&inputTokens=6000"
POST with a JSON body (Live model, longer user speech):
curl -X POST "https://credit-calculator.convai.com/api/v1/estimate" \
-H "content-type: application/json" \
-d '{
"selections": { "llm": "gpt_realtime_1_5", "stt": "included_in_realtime",
"tts": "included_in_realtime" },
"assumptions": { "sttSeconds": 15 }
}'Response fields
| field | meaning |
|---|---|
| api_version | Semantic version of this API (1.0.0). |
| rate_card | Version + effective date of the pricing data the estimate used. |
| selections / assumptions | The fully-resolved request after defaults — what was actually priced. |
| totals.credits | Total credits per interaction (sum of all lines). |
| lines[] | Per-component breakdown: key, label, credits, human-readable detail, and latency_ms (number, {min,max} range, or null). |
| latency_ms | Estimated end-to-end response latency range for this configuration. |
| disclaimer | Estimates only — see convai.com/pricing. |
Line keys
| ai_response | The model generating the response (output tokens). Zero for Live models speaking natively — that cost appears on voice_output instead. |
| prompt | Input context processed by the model (inputTokens x input rate). |
| memory | Flat 1-credit feature charge when memory is true. |
| knowledge | Flat 1-credit feature charge when knowledge is true. |
| voice_input | Speech-to-text, or the Live model's audio input tokens when bundled. |
| voice_output | Text-to-speech, or the Live model's audio output tokens when bundled. |
| animation | Neurosync animation, priced per second of speech. |
| platform | Fixed platform fee per interaction. |
Realtime (Live) models & thinking levels
Live models (e.g. gpt_realtime_1_5, gemini_2_5_flash_live) are voice-native: with the bundled speech options selected, speech seconds bill as the model's audio tokens on the voice lines and there is no separate text-response charge. Choosing an external voice option alongside a Live model is allowed but duplicates what the model already includes.
Models with a thinking level in their name (e.g. gpt_5_1_high) spend reasoning tokens, which bill as output. Their defaults for outputChars are set accordingly — if you override with a much lower value, the estimate will understate what such a model really costs per interaction.
Errors
Invalid fields return 400 with an { "error": "..." } body whose message lists the valid values. Unknown fields are ignored.
Option keys
Generated from the live rate card; also available from /api/v1/options.
llm (AI model)
| gpt_realtime_1_5 | gpt-realtime-1.5 (beta) |
| gemini_2_5_flash_live | Gemini 2.5 Flash Live |
| gpt_realtime_mini | gpt-realtime-mini (beta) |
| deepseek_v4_flash | DeepSeek V4 Flash (non-reasoning) |
| glm_5_2 | GLM-5.2 (non-reasoning) |
| grok_4_3 | Grok 4.3 |
| grok_4_20 | Grok 4.20 (non-reasoning) |
| qwen3_6_35b_a3b | Qwen3.6 35B A3B (beta) |
| qwen3_6_27b | Qwen3.6 27B (beta) |
| gpt_5_6_luna | GPT-5.6 Luna (None) |
| gpt_5_6_terra | GPT-5.6 Terra (None) |
| gpt_5_6_sol | GPT-5.6 Sol (None) |
| claude_haiku_4_5 | claude-4-5-haiku (beta) |
| claude_4_5_sonnet | claude-4-5-sonnet (beta) |
| gpt_4_1 | gpt-4.1 |
| gpt_4_1_mini | gpt-4.1-mini |
| gpt_4_1_nano | gpt-4.1-nano |
| gpt_4o | gpt-4o |
| gpt_4o_mini | gpt-4o-mini |
| gpt_oss_120b | gpt-oss-120b (beta) |
| llama_4_maverick | Llama 4 Maverick (beta) |
| llama_4_scout | Llama 4 Scout (beta) |
| llama3_70b | LLama3-70B |
| gemini_2_5_flash | Gemini 2.5 Flash |
| gemini_2_5_flash_lite | Gemini 2.5 Flash-Lite |
| gemma_4_31b_fast | Gemma 4 31B Fast (beta) |
| gemma_4_26b_a4b_fast | Gemma 4 26B A4B Fast (beta) |
| claude_5_sonnet_none | Claude Sonnet 5 (None) |
| claude_5_sonnet_low | Claude Sonnet 5 (Low) |
| claude_5_sonnet_medium | Claude Sonnet 5 (Medium) |
| claude_5_sonnet_high | Claude Sonnet 5 (High) |
| gpt_5_1_none | GPT-5.1 (None) |
| gpt_5_1_low | GPT-5.1 (Low) |
| gpt_5_1_medium | GPT-5.1 (Medium) |
| gpt_5_1_high | GPT-5.1 (High) |
| gemini_3_1_flash_lite_minimal | Gemini 3.1 Flash-Lite (Minimal) |
| gemini_3_1_flash_lite_low | Gemini 3.1 Flash-Lite (Low) |
| gemini_3_1_flash_lite_medium | Gemini 3.1 Flash-Lite (Medium) |
| gemini_3_1_flash_lite_high | Gemini 3.1 Flash-Lite (High) |
| gemini_3_5_flash_minimal | Gemini 3.5 Flash (Minimal) |
| gemini_3_5_flash_low | Gemini 3.5 Flash (Low) |
| gemini_3_5_flash_medium | Gemini 3.5 Flash (Medium) |
| gemini_3_5_flash_high | Gemini 3.5 Flash (High) |
| gemini_3_5_flash_lite_minimal | Gemini 3.5 Flash-Lite (Minimal) |
| gemini_3_5_flash_lite_low | Gemini 3.5 Flash-Lite (Low) |
| gemini_3_5_flash_lite_medium | Gemini 3.5 Flash-Lite (Medium) |
| gemini_3_5_flash_lite_high | Gemini 3.5 Flash-Lite (High) |
| gemini_3_6_flash_minimal | Gemini 3.6 Flash (Minimal) |
| gemini_3_6_flash_low | Gemini 3.6 Flash (Low) |
| gemini_3_6_flash_medium | Gemini 3.6 Flash (Medium) |
| gemini_3_6_flash_high | Gemini 3.6 Flash (High) |
| gemini_3_7_flash_low | Gemini 3.7 Flash (Low) |
| gemini_3_7_flash_medium | Gemini 3.7 Flash (Medium) |
| gemini_3_7_flash_high | Gemini 3.7 Flash (High) |
| gpt_5_4 | GPT-5.4 (None) |
| gpt_5_4_mini | GPT-5.4 Mini (None) |
| gpt_5_4_nano | GPT-5.4 Nano (None) |
| byo_openai_compatible_endpoint | BYO OpenAI-compatible endpoint |
stt (voice input)
| included_in_realtime | ASR included in Live model |
| stt_v2_chirp_real_time | Google STT v2 Chirp (real-time) |
| nova_3_streaming | Deepgram Nova-3 (streaming) |
| no_stt_input | No STT Input |
| soniox_stt_streaming | Soniox STT (streaming) |
tts (voice output)
| included_in_realtime | Voice included in Live model |
| convai_tts_managed | Convai TTS (managed) |
| convai_cloned_voice | Convai Cloned Voice |
| cloud_tts_standard | Google Cloud TTS Standard |
| cloud_tts_wavenet | Google Cloud TTS WaveNet |
| cloud_tts_neural2 | Google Cloud TTS Neural2 |
| cloud_tts_chirp_3_hd | Google Cloud TTS Chirp 3: HD |
| cloud_tts_studio | Google Cloud TTS Studio |
| azure_neural_prebuilt | Azure Neural (prebuilt) |
| azure_neural_hd | Azure Neural HD |
| azure_custom_professional | Azure Custom Voice (Professional) |
| tts_1 | OpenAI tts-1 |
| tts_1_hd | OpenAI tts-1-hd |
| gpt_4o_mini_tts | OpenAI gpt-4o-mini-tts |
| sonic_3 | Cartesia Sonic 3 |
| eleven_multilingual_v2 | ElevenLabs Multilingual v2 |
| eleven_v3 | ElevenLabs v3 |
| eleven_flash_v2_5 | ElevenLabs Flash v2.5 |
| eleven_turbo_v2_5 | ElevenLabs Turbo v2.5 |
| external_tts_via_api_key | External TTS via API Key |
| no_tts | No TTS |
animation
| no_animation | No Animation |
| convai_neurosync | Convai Neurosync Animation |
Changelog
1.0.0 — initial public release: estimate + options endpoints, credits-only responses, realtime audio-token pricing, full prod model catalog with thinking levels.