Skip to content
AI Model Radar
IndexCollected 2 h ago

Cost check

What does it cost?

Three honest ways to run the same model: pay per token for the full-precision API, or run its quantized weights on a machine you own or rent by the hour. Pick your usage and hardware — every number below decomposes into a measured price, a sourced vendor figure, or an assumption you can change.

≈ 500,000 tokens and 4 h of active compute per day

Prefilled: NVIDIA launch MSRP, Oct 2022. Edit to what you would actually pay.

Cheapest at your settings: Cloud API

Money only — the paths differ in speed, privacy and, for quantized weights, output quality.

Cloud APICheapest at your settings

per day

$0.597

per week

$4.18

per month

$18.16

per year

$218

$0.420 in / $3.00 out per 1M tokens

Your machine

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

per day

$2.01

per week

$14.07

per month

$61.15

per year

$734

$1.46 hardware + $0.550 electricity per day — runs the Q4_K_M quantized weights

Rented GPU

per day

$1.60

per week

$11.22

per month

$48.76

per year

$585

RTX 4090 at $0.401/h — cheapest measured class that fits the Q4_K_M weights

Every assumption on the table

  • + Usage: 500,000 tokens per day, of which 70% input tokens — an order-of-magnitude anchor, not a measurement.
  • + Active compute: 4 hours per day, used for electricity and rental time alike.
  • + Power draw: 450 W board power (vendor spec) plus 100 W for the rest of the machine, while active.
  • + Hardware spread over 3 years, straight-line; resale value ignored.
  • + All figures are USD. API and rental prices are the same worldwide; hardware list prices are US figures excluding sales tax (EU retail prices include VAT and differ), and electricity varies a lot by country — edit both fields to your local numbers.
  • + A month is 365/12 ≈ 30.4 days, so twelve months equal one year.
  • + API and rental prices are collected automatically and carry their collection time; hardware prices are vendor list prices (launch MSRP or current official price, as noted next to the field) you can override.
  • + The three paths differ in speed, privacy and — for quantized local weights — output quality. This table compares money only.

What this is not

Not a quote and not advice: prices move, electricity varies, and your real token volume is yours to judge. Each column shows where its numbers come from — check them against the live sources before spending money.