agenticoutputs

AI Model Pricing

32 models across 9 providers

Updated Jul 27, 2026 Source: OpenRouter API Prices subject to change — verify at provider before deploying
Price Intel
Claude 3.5 Haiku input price dropped 100% $0.80 → $0.00 /1M 2026-07-27
Llama 4 Maverick input price increased 33% $0.15 → $0.20 /1M 2026-07-27
Llama 3.3 70B input price increased 30% $0.10 → $0.13 /1M 2026-07-27
DeepSeek V4 Flash input price increased 43% $0.10 → $0.14 /1M 2026-07-27
DeepSeek V3 input price increased 35% $0.20 → $0.27 /1M 2026-07-27
Sort:
Model Input /1M Output /1M
Claude Fable 5
$10.00 $50.00
Claude Opus 4.8
$5.00 $25.00
Claude Sonnet 4.6
$3.00 $15.00
Claude Haiku 4.5
$1.00 $5.00
Claude 3.5 Haiku
Free Free
GPT-5.5
$5.00 $30.00
GPT-5.4
$2.50 $15.00
GPT-5.4 Mini
$0.75 $4.50
GPT-5.4 Nano
$0.20 $1.25
o3
$2.00 $8.00
o4-mini
$1.10 $4.40
Gemini 3.5 Flash
$1.50 $9.00
Gemini 3.1 Pro
$2.00 $12.00
Gemini 3.1 Flash Lite
$0.25 $1.50
Gemini 2.5 Pro
$1.25 $10.00
Gemini 2.5 Flash
$0.30 $2.50
Llama 4 Maverick
$0.20 $0.80
Llama 4 Scout
$0.10 $0.30
Llama 3.3 70B
$0.13 $0.40
Mistral Medium 3.5
$1.50 $7.50
Mistral Large
$0.50 $1.50
Mistral Small
$0.15 $0.60
DeepSeek V4 Pro
$0.43 $0.87
DeepSeek V4 Flash
$0.14 $0.28
DeepSeek R1
$0.50 $2.15
DeepSeek V3
$0.27 $1.12
Qwen 3.7 Max
$1.48 $4.42
Qwen 3.7 Plus
$0.32 $1.28
Grok 4.20
$1.25 $2.50
Grok 4.3
$1.25 $2.50
Command A
$2.50 $10.00
Command R
$0.15 $0.60

↓ = price drop since last update · Value score = context / combined cost (higher = more tokens per dollar)

32 models

Cost Reduction Strategies

Prompt Caching

Cache system prompts with Claude and cut re-use costs by up to 90%. Write once, pay ~10% per cached hit. Critical for any high-volume agent.

Read docs →
🔀

Model Routing

Route simple tasks (classification, intent, formatting) to fast/cheap models. Reserve frontier inference for generation and reasoning. 60–80% cost reduction typical.

📦

Batch API

Anthropic and OpenAI both offer 50% off for async batch requests. Offline eval runs, bulk tagging, document processing — always batch these.

Read docs →
✂️

Output Tokens

Output tokens cost 3–5× more than input. Tell the model to be concise. Structured JSON output with strict schemas eliminates filler and preamble.

📐

Context Window

Every token in context costs money on every call. Summarize conversation history. Use RAG to retrieve only relevant chunks instead of full documents.