Skip to content
AI Model Radar
IndexCollected 2 h ago

Model · Alibaba Qwen

Qwen3 4B

4B · Q4_K_M · 2.3 GB · open weights · curated for general, coding · repo created 2025-08-05

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.Hugging Face repo (External link)GGUF file (External link)Ollama (External link)

On the radar

50

Heat Score · flat ◆ · 100% confidence

Rank among scored models
24 of 45
Trending score · Hugging Face
6
Downloads, rolling 30 days · Hugging Face
3,579,530
Downloads vs. the reading of 2026-09-05
+1.1%
Likes · Hugging Face
959
API price per 1M tokens, in / out · OpenRouter

Measured signals as of 2026-09-12 06:00 UTC · How the Heat Score is made →

Downloads, daily readingsdaily last · UTC

2026-08-30 · 3.4M2026-09-12 · 3.6M

Heat is a composite of measured signals only: each component is the model's percentile among models with a full week of history — trending level (35%), 7-day download growth (30%), 7-day trending change (15%), 30-day downloads (15%) and Hub likes (5%). Missing components renormalize the weights and lower the shown confidence; nothing is guessed. Trending and downloads come from the Hugging Face Hub — the download counter is a rolling 30-day window, not unique users — and input pricing from OpenRouter (CC BY 4.0). The full formula, thresholds and flag rules are on the methodology page.

What it needs

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

Weights, context cache and runtime overhead at four reference contexts. The total is the model's own; what differs per machine is the usable memory it has to fit into.

What it needs
ContextWeightsContext cacheOverheadTotal
8K2.31.10.94.3 GB
32K2.34.51.78.5 GB
64K2.39.02.714.0 GB
128K2.318.04.725.0 GB

All memory figures are estimates: measured quantized file size + computed context memory + runtime overhead, with a 12% safety margin on your hardware. Real usage varies with runtime version and settings.

Where it runs

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

Every tracked GPU and Mac at 8K and 32K context with a full-precision (f16) cache — the same engine as the finder and the GPU pages.

Runs comfortably (EXCELLENT or GOOD) on 14 of 14 tracked devices at 8K context.

Where it runs
Device8K context32K context
RTX 407012 GB · 10.6 GB usableExcellent fit4.3 GBGood fit8.5 GB
RTX 3060 12GB12 GB · 10.6 GB usableExcellent fit4.3 GBGood fit8.5 GB
RTX 408016 GB · 14.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 508016 GB · 14.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 4060 Ti 16GB16 GB · 14.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 5060 Ti 16GB16 GB · 14.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 5070 Ti16 GB · 14.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 309024 GB · 21.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 409024 GB · 21.1 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
RTX 509032 GB · 28.2 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
Mac mini M4 Pro 64GB64 GB unified · 44.8 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
Mac Studio M4 Max 64GB64 GB unified · 44.8 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
Mac Studio M3 Ultra 96GB96 GB unified · 67.2 GB usableExcellent fit4.3 GBExcellent fit8.5 GB
NVIDIA DGX Spark128 GB unified · 112.6 GB usableExcellent fit4.3 GBExcellent fit8.5 GB

GGUF is the quantized single-file format local runtimes load (Ollama, LM Studio, llama.cpp); our memory figures are measured from the linked GGUF file. The original repo holds full-precision safetensors — a much larger download meant for GPUs with far more memory.

Check it against your own machine →

Or rent it

Cheapest measured rental class that runs it comfortably (EXCELLENT or GOOD at 8K): A100 at $0.94/h, the median across verified on-demand offers.

2026-09-12 08:00 UTC

Live rental prices →Buy or rent? The guide →

More from Alibaba Qwen

Every tracked model →