FIND YOUR FIT

Find the right local AI model for your task

Choose what you want to do. Explore quality, hardware estimates and sources before deciding what to run.

1,317 catalog groups22 reviewed open models with scoresCatalog checked: 2026-09-26 ↓
QUALITY × LOCAL MEMORY

Which open models lead?

Top 5 within this selection. LiveBench quality score; Q4 memory is a planning estimate, not a local test.
#ModelLiveBench / 100Memory · Q4 estimateWeight license
1Smaug MiniAbacus AI76.9~29 GBApache 2.0 ↗
2Qwen3.8 27BAlibaba75.3~29 GBApache 2.0 ↗
3Qwen3.6 27BAlibaba64.0~29 GBApache 2.0 ↗

LiveBench 2026-06-25 ↗ · Selected task: Overall · Scores collected: 2026-09-26. Positions cover only evaluated models within the selected memory budget. This LiveBench score is not the Artificial Analysis Intelligence Index, and the two scales are not directly comparable.

Speed depends on the computer, runtime and quantization. Check tokens/s and time to first token with the measured configuration.

Checking the latest saved data…

LOCAL AI WIZARD

What do you want to do?

Start with the task. Find the right model for you.

FROM THE CATALOG

Recently updated repositories

These entries update with the catalog, even without a benchmark. Repository changes are not necessarily new model releases.

SOURCE-BACKED GUIDES

Take a closer look

12 checkpoint profiles, 22 hardware configurations and 3 task comparisons. Each page keeps its source, date and limits visible.

Sources, methodology & data status
Benchmark release: 2026-06-25Data checked: 2026-09-26Verified snapshot · checking source

01 / LiveBench

The overall score is the equally weighted mean of seven category means (0–100). All scores use one release. This is not the Artificial Analysis Intelligence Index; the two scales must not be mixed.

CSV ↗ · Categories ↗ · LiveBench ↗ · Data reuse policy ↗

02 / Memory estimate

Planning formula: total parameters × 0.625 bytes, plus 20% working allowance and 8 GB reserve, rounded up in decimal GB. This is a heuristic, not a measured minimum or guarantee. Quantization, context, KV cache, vision components, runtime and offloading change requirements. MoE uses total parameters, not only active ones.

Qwen counts refer to language parameters; GLM-5.3 and DeepSeek V4 use repository tensor totals; GLM-5.3 Flash uses the manufacturer’s 320B specification. Architectures without a reviewed estimate remain outside the memory axis.

03 / Reference, not a local test

LiveBench evaluates the named model and reasoning effort. Its score is not a measurement of the local Q4 variant. Hardware links are official destinations without affiliate tracking. Stock, exact configuration and regional availability can change.

04 / Automatic updates without paid APIs

The deployed product checks scores, approved official model organizations, US prices and stock daily, and hardware specifications weekly. Due visits can also refresh safely. D1 preserves the last valid snapshot. Incompatible methodology is quarantined; unavailable sources never erase valid data. New models remain pending benchmark or insufficient data until official facts exist.