Start here

79Models Indexed
5Formats Tracked
43GPUs in Database
98.4%Avg Accuracy Retained

Data updated 2026-08-20

This week’s updates

New models, recency tags, and data cadence2026-08-20

  • 2026-08-20CLI generator now emits commands that actually run: the real GGUF repo and filename instead of a placeholder, `ollama run hf.co/…` instead of a tag that 404s, and a repo id for vLLM instead of the model's display name. llama.cpp build flags updated to the GGML_* names (the old LLAMA_* ones are ignored, giving a silent CPU-only build)
  • 2026-08-20VRAM calculator: forward mode now uses each model's own measured bits-per-weight instead of the generic per-level table, matching what reverse mode always did. GPT-OSS 20B at Q8_0 was overstated by 67% (21.1GB → 12.7GB). EXL2 3.5bpw is selectable again
  • 2026-08-18Chinese edition audit: every /zh page now declares zh-Hans in the HTML itself, the 23 guide links on /zh/cookbook stopped bouncing readers to English, and structured data on 102 Chinese pages describes the Chinese page rather than the English one
Full changelog

Explore the Site

Data Changelog

Last updated 2026-08-20

  • 2026-08-20CLI generator now emits commands that actually run: the real GGUF repo and filename instead of a placeholder, `ollama run hf.co/…` instead of a tag that 404s, and a repo id for vLLM instead of the model's display name. llama.cpp build flags updated to the GGML_* names (the old LLAMA_* ones are ignored, giving a silent CPU-only build)
  • 2026-08-20VRAM calculator: forward mode now uses each model's own measured bits-per-weight instead of the generic per-level table, matching what reverse mode always did. GPT-OSS 20B at Q8_0 was overstated by 67% (21.1GB → 12.7GB). EXL2 3.5bpw is selectable again
  • 2026-08-18Chinese edition audit: every /zh page now declares zh-Hans in the HTML itself, the 23 guide links on /zh/cookbook stopped bouncing readers to English, and structured data on 102 Chinese pages describes the Chinese page rather than the English one
  • 2026-08-18Chinese edition is now indexable: /zh/** mirrors all 113 pages with hreflang pairing, and the Chinese text is baked into the static HTML rather than swapped in after load
  • 2026-08-08Model index +4 → 79: Qwen3-VL 8B / 30B-A3B, Magistral Small 1.2, Seed-OSS 36B — picked for constrained hardware (multimodal on a 12GB card, small-active MoE for unified memory, 512K context at dual-GPU size). Qwen2-VL 7B marked superseded
  • 2026-08-08New cookbook: running GPT-OSS 20B/120B locally without re-quantizing (23 guides). VRAM calculator now offers MXFP4 — picking Q4_K_M for GPT-OSS overstated weights by ~14%
  • 2026-08-07Model index +4 → 75: GPT-OSS 20B/120B (native MXFP4), GLM-4.5-Air 106B-A12B, Devstral Small 1.1 — MoE-heavy batch for 16GB cards and unified-memory Macs
  • 2026-07-22Polish: re-rendered og.png (71+ models), all 22 cookbook guides have verified stack banners
  • 2026-07-22Cadence pack: +4 models (Gemma 3 27B, R1-Llama-8B, Phi-4, Qwen3 1.7B), superseded tags, measured/estimated labels, Hub “recent”, weekly updates, RSS, cookbook verified stack (71 models)
  • 2026-06-26UX for real traffic: job paths, mobile GPU profile, OG/favicon, honest format heat, feedback email
  • 2026-06-26Model index +4: Qwen3 4B, Qwen3-Coder 30B-A3B, Mistral Large 3, GLM-4-9B (67 total)
  • 2026-06-26QA fixes: HF stats merge on failure, ≤3B filter, CLI/VRAM tool bugs, i18n polish
  • 2026-06-26Model index +5: Qwen3 32B, 30B-A3B MoE, 235B-A22B, DeepSeek-V3, DeepSeek-R1 (63 total)
  • 2026-06-26Model index +7: Qwen3 8B/14B, Gemma 3 4B/12B, Llama 4 Scout/Maverick, Llama 3.1 405B (58 total)
  • 2026-06-25Cookbook TOC scroll highlight, code block copy, Quant Hub Markdown export
  • 2026-06-25Cookbook reading progress bar, model HF link copy, Quant Hub shareable filter URLs
  • 2026-06-25Breadcrumb nav + JSON-LD, cookbook article TOC, Quant Hub GPU quick-filter chips
  • 2026-06-25Related cookbook guides, similar-model cards on detail pages, hero latest-update badge
  • 2026-06-25Homepage explore strip, 404 page, multi-model benchmarks (Qwen 7B/32B, DeepSeek-R1 14B), llms.txt for AI crawlers
  • 2026-06-24Added About page (/about) — maintainer story, update cadence, contribution guide
  • 2026-06-24Quant Hub: default to all 51 models visible; scale stats bar; clearer GPU filter UX
  • 2026-06-24SEO: canonical URLs, JSON-LD, per-page metadata, Google/Bing verification env vars
  • 2026-06-24Quant Hub: show-all toggle when GPU profile active; Cookbook +7 guides (8GB GPU, WSL2, Docker GPU, Nginx, AMD ROCm)
  • 2026-06-24Privacy Policy + Plausible analytics, cookbook standalone pages (/cookbook/[slug]), model index expanded to 51
  • 2026-06-24Added Terms & Disclaimer page (/legal) with trademark notice and liability disclaimer
  • 2026-06-24Phase 3: HF live stats pipeline, model A vs B compare tool, cookbook expanded to 15 guides
  • 2026-06-24Expanded model index to 30+ entries; added format wizard, hardware profile, ExLlamaV2 CLI, SEO (sitemap/OG), data transparency
  • 2026-06-24Model detail pages, GPU reverse lookup, shareable VRAM calculator URLs, real homepage stats
  • 2025-06-10Initial launch: Quant Hub, VRAM calculator, CLI generator, benchmarks, cookbook

All data is manually curated and verified against community sources. Always cross-check with official Hugging Face model cards before deployment.