DeepInfra

DeepInfra

US
62 models available

Data handling can vary by model and deployment. The privacy badge above shows the more conservative of the provider-wide fallback and the policies for routes on this page; each model page shows the policy for that specific route.

Models served by DeepInfra

ModelContextInput /MOutput /MCache read /MLatencyThroughput
128k$0.53/M$2.26/M$0.37/M0.7s30 tps
128k$0.26/M$1.00/M$0.14/M0.8s6 tps
163k$0.27/M$0.40/M$0.14/M1.5s15 tps
163k$0.27/M$0.40/M$0.14/M1.5s15 tps
1.0M$0.09/M$0.19/M$0.02/M1s28 tps
1.0M$0.09/M$0.19/M$0.02/M1s28 tps
1.0M$0.06/M$0.19/M$0.02/M0.9s43 tps
1.0M$0.06/M$0.19/M$0.02/M0.9s43 tps
1.0M$0.23/M$0.68/M$0.0072/M1.8s35 tps
1.0M$1.36/M$2.73/M$0.10/M1.2s38 tps
1.0M$1.36/M$2.73/M$0.10/M1.2s38 tps
1.0M$1.36/M$2.73/M$0.10/M1.7s66 tps
1.0M$1.36/M$2.73/M$0.10/M1.7s66 tps
1.0M$0.15/M$0.44/M$0.0044/M1.4s61 tps
1.0M$0.15/M$0.44/M$0.0044/M1.4s61 tps
262k$0.07/M$0.36/M$0.01/M0.7s24 tps
262k$0.16/M$0.42/M$0.02/M2.6s10 tps
200k$0.53/M$2.10/M$0.10/M0.7s31 tps
200k$0.53/M$2.10/M$0.10/M0.7s31 tps
200k$0.42/M$1.84/M$0.08/M1s24 tps
200k$0.42/M$1.84/M$0.08/M1s24 tps
1.0M$0.59/M$1.89/M$0.11/M1.1s54 tps
1.0M$0.59/M$1.89/M$0.11/M1.1s54 tps
1.0M$0.59/M$2.63/M$0.13/M1.9s53 tps
1.0M$0.59/M$2.63/M$0.13/M1.9s53 tps
1.0M$0.08/M$0.26/M$0.02/M1.4s29 tps
128k$0.16/M$0.63/MN/A0.4s110 tps
128k$0.03/M$0.15/MN/A0.4s75 tps
131k$0.06/M$0.26/M$0.02/M0.7s64 tps
66k$0.73/M$0.73/MN/A0.6s27 tps
524k$0.47/M$1.26/M$0.10/M0.8s57 tps
524k$0.47/M$1.26/M$0.10/M0.8s57 tps
256k$0.79/M$3.68/M$0.16/M2.2s9 tps
1.0M$2.99/M$14.96/M$0.30/M2.1s19 tps
131k$0.02/M$0.04/MN/A0.8s24 tps
131k$0.10/M$0.34/MN/A0.5s14 tps
1.0M$1.05/M$3.15/M$0.21/MN/AN/A
1.0M$1.05/M$3.15/M$0.21/MN/AN/A
1.0M$0.15/M$0.29/M$0.0029/M1.4s49 tps
1.0M$0.45/M$0.91/M$0.0038/M1.8s23 tps
512k$0.29/M$1.16/M$0.06/M1.2s18 tps
512k$0.29/M$1.16/M$0.06/M1.2s18 tps
262k$0.09/M$0.42/MN/A2.4s60 tps
262k$0.09/M$0.42/MN/A2.4s60 tps
256k$0.53/M$2.31/M$0.11/MN/AN/A
256k$0.53/M$2.31/M$0.11/MN/AN/A
29k$0.06/M$0.17/M$0.03/MN/AN/A
29k$0.06/M$0.17/M$0.03/MN/AN/A
131k$0.38/M$0.42/MN/A0.4s22 tps
260k$0.34/M$3.36/MN/A0.4s40 tps
260k$0.34/M$3.36/MN/A0.4s40 tps
262k$0.10/M$1.00/M$0.10/M0.5s140 tps
262k$0.10/M$1.00/M$0.10/M0.5s140 tps
262k$0.09/M$0.58/MN/A0.6s14 tps
41k$0.08/M$0.29/MN/A0.4s22 tps
256k$0.09/M$1.16/MN/A0.6s43 tps
258k$0.47/M$3.15/M$0.23/M0.8s43 tps
258k$0.47/M$3.15/M$0.23/M0.8s43 tps
262k$0.16/M$1.97/M$0.04/M1.2s34 tps
262k$0.16/M$1.97/M$0.04/M1.2s34 tps
262k$2.10/M$6.30/M$0.21/M1.3s94.5 tps
262k$2.10/M$6.30/M$0.21/M1.3s94.5 tps

Prices are per million tokens. Selectable routes include the applicable provider-selection markup; fixed routes show the standard model price. Availability and pricing refresh continuously; each model page shows the live provider comparison.

Using DeepInfra via the API

For models with provider selection, append :deepinfra to the model ID, set "provider": "deepinfra" in the request body, or send an X-Provider: deepinfra header. Explicit provider selection adds a route-specific markup over that provider's base price; the prices in the table above already include it. Models without provider selection use the displayed model ID without a provider suffix or selection surcharge.

curl https://nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-R1-0528:deepinfra",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the API documentation for provider preferences, price-aware routing, and error behavior.