Skip to content
AI Model Radar
IndexCollected 2 h ago

01 / Stack builder

What are you building?

Five questions. A complete stack with a reason behind every slot — assembled from rules, not by an LLM, so the same answers always give the same result.

Assembled deterministically from the same measured data as the rest of the site: local fit from your hardware and published model sizes, prices from OpenRouter, rental rates from the Vast.ai marketplace. No LLM decides this ranking.

Hybrid stack

Cloud budget: $50/mo buys about 119M input tokens/mo at Qwen3.8 27B

Hybrid stack. Main brain: Qwen3.8 27B. Fast + cheap: LFM2.5 2.6B. Local model: GLM-4.7-Flash. Runtime: llama.cpp. Cloud router: OpenRouter.

Main brain

Medium cost

Qwen3.8 27B

$0.42/M input tokens

The most-adopted model for this use case among those with a published price. We publish no quality benchmarks, so this reflects usage, not a capability ranking.

OpenRouter ↗ (External link)

Fast + cheap

No per-token cost

LFM2.5 2.6B

$0.00/M input tokens

Cheaper than your main brain, for classification, routing and other high-volume calls that do not need the strongest model.

OpenRouter ↗ (External link)

Local model

No per-token cost

GLM-4.7-Flash

Q4_K_M · ~18.9 GBMemory fit 73/100

Runs on hardware you already own, so everyday work costs nothing per token and never leaves your machine.

Hugging Face ↗ (External link)

Runtime

No per-token cost

llama.cpp

GGUF weights run under llama.cpp without a packaged runtime.

Cloud router

No cost

OpenRouter

One API across providers, so swapping a model does not become a migration. The prices on this site come from its public catalogue.

OpenRouter ↗ (External link)