Selected comparison
Local vs cloud LLM comparison
These 6 local models need an estimated 13.2–62.4 GB at Q4 and 4K context. Compare them with 4 hosted options from Claude, Astra and Gemini. Choose your hardware, then estimate monthly API token costs alongside Arena's preference ratings from 2026-10-02.
Showing 10 selected models. Local fit uses Apple M4 (24GB) at 4K context.
Usage estimate
Estimate listed API spend
Enter millions of tokens for one month. Only hosted rows with a listed standard rate receive an estimate.
Estimates use standard-tier uncached input and billed output rates. They exclude cache reads and writes, tools, batch or priority discounts, long-request surcharges and taxes. Known expired rates show Unavailable. Follow each price link for vendor notes.
Preference and fit
Selected models, highest rating first
Scroll the table sideways for local memory fit and API costs.
| Model | Arena preference rating | Local fit and Q4 memory | Dated API price | Monthly API estimate |
|---|---|---|---|---|
| Gemini 3.8 Flash (High)gemini-3.8-flash-high · Hosted | 1494.81489.7–1499.8 · 26,298 votes | Hosted only | In $0.75 / MOut $3.75 / MVerified 2026-09-27Through 2026-12-31 | $1.50 |
| Claude Opus 5 (Legacy, High)claude-opus-5-high · Hosted | 1489.71485.9–1493.5 · 59,482 votes | Hosted only | In $5.00 / MOut $25.00 / MVerified 2026-09-27 | $10.00 |
| GPT-6 Astra (Max)gpt-6-astra-max · Hosted | 1477.11470.2–1484.0 · 9,156 votes | Hosted only | In $10.00 / MOut $50.00 / MVerified 2026-09-27 | $20.00 |
| Claude Sonnet 5 (High)claude-sonnet-5-high · Hosted | 1461.81457.6–1466.0 · 45,739 votes | Hosted only | In $2.00 / MOut $10.00 / MVerified 2026-09-27 | $4.00 |
| Gemma 4 31Bgemma-4-31b · Local | 1452.81445.4–1460.3 · 6,133 votes | No fit~20.5 GB4K setup details | Not listed | Not listed |
| Gemma 4 26B-A4Bgemma-4-26b-a4b · Local | 1437.51430.0–1444.9 · 6,081 votes | No fit~19 GB4K setup details | Not listed | Not listed |
| gpt-oss 120Bgpt-oss-120b · Local | 1351.61347.2–1356.0 · 30,018 votes | No fit~62.4 GB4K setup details | Not listed | Not listed |
| Qwen3 32Bqwen3-32b · Local | 1346.91337.5–1356.4 · 3,926 votes | No fit~22 GB4K setup details | Not listed | Not listed |
| Qwen3 30B-A3Bqwen3-30b-a3b · Local | 1326.51321.7–1331.2 · 26,037 votes | No fit~20.7 GB4K setup details | Not listed | Not listed |
| gpt-oss 20Bgpt-oss-20b · Local | 1317.61311.2–1324.0 · 10,368 votes | Fits~13.2 GB4K setup details | Not listed | Not listed |
The bounds show the source confidence interval. Rating intervals can overlap, so do not assume a reliable ordering from nearby rows. Names retain the source's evaluated reasoning modes. Arena's settings can differ from a local Q4 installation.
How to read this
One source snapshot, two deployment paths
The preference ratings come from the linked Arena snapshot on 2026-10-02. Local memory uses this catalog's Q4 measurements and selected context. Monthly API usage is separate from that per-request context setting. Billed output includes reasoning tokens, which can exceed the visible answer. Memory fit does not guarantee runtime support or usable speed. Vendor API prices have their own verification dates and links per row. These are separate facts, shown together for planning, not evidence of parity.
For more hosted models and rates, use the API pricing calculator. For a broader local benchmark view, visit the local leaderboard. To size any local model for a machine, use the memory calculator. Can you run Claude locally? explains the open-weight boundary.
Methodology and dates
Arena: Bradley-Terry rating, adapted from the dataset snapshot dated 2026-10-02; rounded to one decimal here.
Memory: Q4 local-catalog estimate recomputed for the selected device and context. Memory data updated 2026-10-05.
Prices: USD per million tokens, separate from subscription plans. Vendor rates below cover the base API model; the Arena label retains the evaluated reasoning mode.
Gemini 3.8 Flash (High) (gemini-3.8-flash), verified 2026-09-27: Standard paid Gemini Developer API rates through December 31, 2026. Output includes thinking tokens; cache storage, tools, and other service tiers are excluded.
Claude Opus 5 (Legacy, High) (claude-opus-5), verified 2026-09-27: Standard Claude API rates. Cache writes, batch discounts, fast mode, and US inference pricing are excluded.
GPT-6 Astra (Max) (gpt-6-astra), verified 2026-09-27: Standard text rates. Requests over 272K input tokens have higher rates; cache writes and tools are excluded.
Claude Sonnet 5 (High) (claude-sonnet-5), verified 2026-09-27: Standard Claude API rates. Cache writes, batch discounts, fast mode, and US inference pricing are excluded.
FAQ
Frequently asked questions
Can I run a cloud model locally?
Usually not. Hosted rows represent a provider API run, while local rows link to an open-weight model with a Q4 memory estimate. The hardware selector only evaluates the local catalog models in this selected comparison.
Does an Arena preference rating measure local performance?
No. The rating comes from the listed Arena leaderboard snapshot and retains its evaluated mode. A hosted Arena run is not a local quantization benchmark, so use the rating for preference context and the memory verdict for local hardware planning.
Are the API costs a bill quote?
No. The estimate multiplies your entered input and output millions of tokens by the linked vendor rates in this snapshot. It excludes subscriptions, tools, regional pricing, taxes and any rate that is not listed.
Source retrieved 2026-10-05. Download the comparison data. Selection is editorial and does not claim a global rank. Ratings are rounded for display; source and license links above carry the full dataset context.