Modelprices
52 services · 69 calls/30d · modelprices.xyz
Modelprices · LLM Prices
modelprices.xyz
Compare AI model prices side-by-side: a live price leaderboard of LLM token cost for GPT-5, Claude 5, Claude Sonnet, Gemini 3 Pro, Llama 4, DeepSeek V4, Grok 4, Mistral Large and 2,000+ more models across 70+ providers (OpenAI, Anthropic, Google, Meta, xAI). Inference cost per token — input, output, cache and batch USD per 1M tokens — ranked, normalized into one table, cross-checked across two sources, refreshed hourly. Find the cheapest model and estimate token budgets.
Modelprices · LLM Limits
modelprices.xyz
AI model capability, specs and limits table for 2,000+ LLMs — GPT-5, GPT-4o, Claude 5, Claude Sonnet, Gemini 3 Pro, Llama 4, DeepSeek V4 and more: context window size, max output tokens, vision/audio/function-calling/reasoning support. Compare model constraints side-by-side across every provider to pick the right model for long-context, multimodal, or tool-use workloads. Refreshed hourly.
Modelprices · AI Model Pricing
modelprices.xyz
AI model pricing comparison table: what every AI model costs right now — GPT-5, GPT-4o, Claude 5, Claude Sonnet, Gemini 3 Pro, Llama 4, DeepSeek V4, Grok 4 and 2,000+ more. One normalized dataset of LLM token cost and inference pricing across OpenAI, Anthropic, Google and 70+ providers, refreshed hourly. Compare models by price, rank by cost, pick the cheapest model for any workload.
Modelprices · LLM Cheapest
modelprices.xyz
Cheapest-model query: returns the lowest-cost AI models matching your constraints — min_context (tokens), vision/function_calling/reasoning (true), max_input_per_mtok (USD cap), limit (default 5). Add expected_input/expected_output token counts to rank by blended USD per request instead of input price; exclude_preview=true skips pre-release models. Answers 'what is the cheapest model that can do X?' across 2,000+ models and 70+ providers.
Modelprices · LLM Price Changes
modelprices.xyz
LLM price-change alert feed: track when OpenAI, Anthropic, Google or any AI provider reprices GPT, Claude, Gemini or any model's inference cost per token — old vs new price, percent delta, when it happened, plus newly launched and removed models. Structured JSON diffed from hourly snapshots; ideal for cost monitoring, repricing triggers, and AI market intelligence.
Modelprices · Prices Together
modelprices.xyz
Together AI pricing table: per-token cost of every AI model Together AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Together AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Watsonx
modelprices.xyz
IBM watsonx pricing table: per-token cost of every AI model IBM watsonx serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside IBM watsonx and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Cheapest 128k Context
modelprices.xyz
Cheapest 128k-context LLM leaderboard: the 50 lowest-cost AI models with a 128,000-token context window or larger, ranked by inference cost per token across 70+ providers. Input, output and cache USD per 1M tokens with exact context window and capability flags joined in. The standard long-document tier, priced side-by-side. Refreshed hourly.
Modelprices · Cheapest Function Calling
modelprices.xyz
Cheapest tool-use LLM leaderboard: the 50 lowest-cost AI models that support function calling / tool use, ranked by token price across 70+ providers. Input, output and cache USD per 1M tokens with context window and modality flags joined in. The cheapest way to pick a model for an agent loop that needs reliable structured tool calls. Refreshed hourly.
Modelprices · Cheapest Long Context
modelprices.xyz
Cheapest long-context LLM leaderboard: the 50 lowest-cost AI models with a 200,000-token context window or larger, ranked by inference cost per token — Claude 5, Gemini 3, GPT-5, Llama 4 Scout and more across 70+ providers. Input, output and cache USD per 1M tokens with exact context window and max output tokens joined in. Refreshed hourly.
Modelprices · Cheapest Million Token Context
modelprices.xyz
Cheapest million-token-context LLM leaderboard: every AI model with a 1,000,000-token context window or larger, ranked by inference cost per token — Gemini 3 Pro, Gemini 3 Flash, Llama 4 Scout, GPT-5 long-context tiers and more. Input, output and cache USD per 1M tokens with exact context window and max output joined in. Answers 'what is the cheapest model that fits my whole corpus?' Refreshed hourly.
Modelprices · Cheapest Overall
modelprices.xyz
Cheapest LLM leaderboard: the 50 lowest-cost AI models available anywhere right now, ranked by inference cost per token across 70+ providers (OpenAI, Anthropic, Google, Meta, xAI, Mistral, DeepSeek). Input, output, cache and batch USD per 1M tokens, with context window and capability flags joined in. The one call that answers 'what is the cheapest model I can use?' Refreshed hourly.
Modelprices · Cheapest Reasoning
modelprices.xyz
Cheapest reasoning-model leaderboard: the 50 lowest-cost AI models with extended reasoning / chain-of-thought support, ranked by inference cost per token — o3, Claude 5 thinking, Gemini 3 Pro, DeepSeek R1, Qwen QwQ and more across 70+ providers. Input, output and cache USD per 1M tokens with context windows joined in. Refreshed hourly.
Modelprices · Cheapest Vision
modelprices.xyz
Cheapest vision/multimodal LLM leaderboard: the 50 lowest-cost AI models that accept image input, ranked by token price across every provider — GPT-5, Claude 5, Gemini 3, Llama 4, Qwen VL and more. Input, output and cache USD per 1M tokens with context window joined in. Answers 'what is the cheapest model that can read images?' in one call. Refreshed hourly.
Modelprices · Claude Pricing
modelprices.xyz
Claude pricing table: what every Claude model costs per token right now — Claude 5, Claude Opus, Claude Sonnet, Claude Haiku — from Anthropic and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Claude inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Command Pricing
modelprices.xyz
Command pricing table: what every Command model costs per token right now — Command A, Command R+, Command R — from Cohere and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Command inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Deepseek Pricing
modelprices.xyz
DeepSeek pricing table: what every DeepSeek model costs per token right now — DeepSeek V4, DeepSeek R1, DeepSeek Coder — from DeepSeek and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare DeepSeek inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Gemini Pricing
modelprices.xyz
Gemini pricing table: what every Gemini model costs per token right now — Gemini 3 Pro, Gemini 3 Flash, Gemini 2.5 — from Google and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Gemini inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Gemma Pricing
modelprices.xyz
Gemma pricing table: what every Gemma model costs per token right now — Gemma 3, Gemma 2 — from Google and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Gemma inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · GLM Pricing
modelprices.xyz
GLM pricing table: what every GLM model costs per token right now — GLM-4.6, GLM-4.5 Air — from Zhipu AI and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare GLM inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · GPT Pricing
modelprices.xyz
GPT pricing table: what every GPT model costs per token right now — GPT-5, GPT-5 mini, GPT-4o, GPT-4.1, o3 — from OpenAI and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare GPT inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Grok Pricing
modelprices.xyz
Grok pricing table: what every Grok model costs per token right now — Grok 4, Grok 4 mini, Grok 3 — from xAI and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Grok inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Jamba Pricing
modelprices.xyz
Jamba pricing table: what every Jamba model costs per token right now — Jamba 1.5 Large, Jamba Mini — from AI21 and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Jamba inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Kimi Pricing
modelprices.xyz
Kimi pricing table: what every Kimi model costs per token right now — Kimi K2, Moonshot v1 — from Moonshot AI and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Kimi inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Llama Pricing
modelprices.xyz
Llama pricing table: what every Llama model costs per token right now — Llama 4 Scout, Llama 4 Maverick, Llama 3.3 — from Meta and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Llama inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · LLM Context Windows
modelprices.xyz
Largest-context-window leaderboard: the 100 AI models with the biggest context windows, ranked descending — how many tokens each model can actually take, its max output tokens, modality support, and what those tokens cost. Compare GPT-5, Claude 5, Gemini 3 Pro, Llama 4 Scout and 2,000+ more on context capacity and price together. Refreshed hourly.
Modelprices · LLM Limit
modelprices.xyz
Single-model capability lookup: context window size, max output tokens, and vision/audio/function-calling/reasoning support for one AI model by id. Answers 'can this model handle my task?' in one $0.003 call.
Modelprices · LLM Price
modelprices.xyz
Single-model price lookup: current per-token cost for one AI model by id (e.g. claude-sonnet-5, gpt-5, gemini-3-pro) — input, output, cache and batch USD per 1M tokens. The cheapest way to answer 'what does this model cost right now?' inside a routing or budgeting decision. Fuzzy-matches model ids and suggests alternatives on miss. Includes provenance: provider source URL, first_observed_at, confidence tier.
Modelprices · Minimax Pricing
modelprices.xyz
MiniMax pricing table: what every MiniMax model costs per token right now — MiniMax M2, MiniMax Text — from MiniMax and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare MiniMax inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Mistral Pricing
modelprices.xyz
Mistral pricing table: what every Mistral model costs per token right now — Mistral Large, Mixtral, Codestral, Magistral — from Mistral AI and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Mistral inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · PHI Pricing
modelprices.xyz
Phi pricing table: what every Phi model costs per token right now — Phi-4, Phi-3.5 — from Microsoft and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Phi inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.
Modelprices · Prices Anthropic
modelprices.xyz
Anthropic pricing table: per-token cost of every AI model Anthropic serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Anthropic and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Azure
modelprices.xyz
Microsoft Azure pricing table: per-token cost of every AI model Microsoft Azure serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Microsoft Azure and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Bedrock
modelprices.xyz
AWS Bedrock pricing table: per-token cost of every AI model AWS Bedrock serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside AWS Bedrock and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Cloudflare
modelprices.xyz
Cloudflare Workers AI pricing table: per-token cost of every AI model Cloudflare Workers AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Cloudflare Workers AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Databricks
modelprices.xyz
Databricks pricing table: per-token cost of every AI model Databricks serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Databricks and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Deepinfra
modelprices.xyz
DeepInfra pricing table: per-token cost of every AI model DeepInfra serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside DeepInfra and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Fireworks
modelprices.xyz
Fireworks AI pricing table: per-token cost of every AI model Fireworks AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Fireworks AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Google
modelprices.xyz
Google AI Studio pricing table: per-token cost of every AI model Google AI Studio serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Google AI Studio and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Mistral
modelprices.xyz
Mistral AI pricing table: per-token cost of every AI model Mistral AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Mistral AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Nebius
modelprices.xyz
Nebius pricing table: per-token cost of every AI model Nebius serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Nebius and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Novita
modelprices.xyz
Novita AI pricing table: per-token cost of every AI model Novita AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Novita AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Openai
modelprices.xyz
OpenAI pricing table: per-token cost of every AI model OpenAI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside OpenAI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Openrouter
modelprices.xyz
OpenRouter pricing table: per-token cost of every AI model OpenRouter serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside OpenRouter and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Oracle
modelprices.xyz
Oracle OCI pricing table: per-token cost of every AI model Oracle OCI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Oracle OCI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Replicate
modelprices.xyz
Replicate pricing table: per-token cost of every AI model Replicate serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Replicate and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Sambanova
modelprices.xyz
SambaNova pricing table: per-token cost of every AI model SambaNova serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside SambaNova and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Snowflake
modelprices.xyz
Snowflake Cortex pricing table: per-token cost of every AI model Snowflake Cortex serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Snowflake Cortex and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Vercel
modelprices.xyz
Vercel AI Gateway pricing table: per-token cost of every AI model Vercel AI Gateway serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Vercel AI Gateway and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices Vertex AI
modelprices.xyz
Google Vertex AI pricing table: per-token cost of every AI model Google Vertex AI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside Google Vertex AI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Prices XAI
modelprices.xyz
xAI pricing table: per-token cost of every AI model xAI serves, in one call — input, output, cache and batch USD per 1M tokens, ranked cheapest first, with context window and capability flags joined in. Compare LLM token cost inside xAI and against other providers hosting the same model. Normalized from public sources, cross-checked, refreshed hourly.
Modelprices · Qwen Pricing
modelprices.xyz
Qwen pricing table: what every Qwen model costs per token right now — Qwen 3, Qwen 2.5, Qwen Coder — from Alibaba and every provider that resells or hosts them (Bedrock, Azure, Vertex, OpenRouter, Fireworks). Input, output, cache and batch USD per 1M tokens, sorted cheapest first, so you can compare Qwen inference cost across hosts and pick the cheapest place to run one. Normalized, cross-checked, refreshed hourly.