91 個開源模型

Local LLMs you can run

These are open chat, coding and reasoning models you can download and run on your own machine. We grade each one against the GPU or Apple Silicon detected in your browser.

GPUVRAMBWRAMCores

FLUX.2 Klein 4B

8 個月前

Black Forest Labs · 4B · Apache 2.0

Fastest open FLUX.2 — sub-second text-to-image and multi-reference editing on consumer GPUs

2.5GB·32K ctx·

Wan 2.2 TI2V 5B

1 年前

Alibaba · 5B · Apache 2.0

Unified text/image-to-video — the local sweet spot under Apache 2.0

3.1GB·4K ctx·

Z-Image Turbo

10 個月前

Alibaba · 6B · Apache 2.0

8-step distilled image model — photorealism and bilingual text on 16GB cards

3.6GB·4K ctx·

HunyuanVideo 1.5

10 個月前

Tencent · 8.3B · Tencent Hunyuan Community

Compact cinematic video model — strong faces and motion on a single 4090

4.8GB·4K ctx·

Qwen3-VL 8B

11 個月前

Alibaba · 8.8B · Apache 2.0

The community-favourite local VLM — superb OCR, receipts & captioning

5GB·256K ctx·

Qwen 3.5 9B

7 個月前

Alibaba · 9B · Apache 2.0

Multimodal Qwen 3.5 mid-size

22 AA·5.1GB·32K ctx·

LTX 2.3

6 個月前

Lightricks · 19B · LTX-2 Community

Open 4K video with native stereo audio — text, image and video-to-video

10.2GB·4K ctx·

Qwen Image 2512

9 個月前

Alibaba · 20B · Apache 2.0

Open text-to-image with strong English and Chinese typography

10.7GB·4K ctx·

GPT-OSS 20B

1 年前

OpenAI · 21B · Apache 2.0

OpenAI's open-weight MoE with configurable reasoning

15 AA·11.3GB·128K ctx·

Gemma 4 26B-A4B IT

5 個月前

Google · 27B · Gemma

Gemma 4 MoE instruct model (official)

26 AA·14.3GB·256K ctx·

Qwen 3.8 27B

1 個月前

Alibaba · 27B · Apache 2.0

Flagship dense Qwen 3.8 — native multimodal all-rounder with video understanding

52 AA·14.3GB·256K ctx·

Wan 2.2 T2V A14B

1 年前

Alibaba · 27B · Apache 2.0

Flagship open Wan 2.2 — 14B-active MoE for photoreal text-to-video

14.3GB·4K ctx·

Muse Glimmer 30B

1 個月前

Meta · 30B · Apache 2.0

Open agentic 30B distilled from Muse Spark — tool use, vision and local recovery on a single GPU

35 AA·15.9GB·128K ctx·

Qwen3-VL 30B-A3B

1 年前

Alibaba · 31B · Apache 2.0

Efficient vision MoE — 3B active, strong temporal & document understanding

16.4GB·256K ctx·

FLUX.2 Dev

10 個月前

Black Forest Labs · 32B · FLUX Non-Commercial

Flagship open-weight FLUX.2 — text-to-image and multi-reference editing up to 4MP

16.9GB·32K ctx·

MiniMax H3

1 個月前

MiniMax · 33B · MiniMax Community

Open video generation — text/image to 2K video with native stereo audio

17.4GB·32K ctx·

Agents-A1 35B-A3B

3 個月前

InternScience · 35B · Apache 2.0

Efficient multimodal agentic MoE for long-horizon search, engineering and scientific research

18.4GB·256K ctx·

Ornith 1.0 35B-A3B

3 個月前

DeepReinforce · 35B · MIT

Agentic coding MoE with a 3B active working set and self-improving training

18.4GB·256K ctx·

Qwen 3.6 35B-A3B

5 個月前

Alibaba · 36B · Apache 2.0

Big-model quality at 3B-active speed — the mid-hardware sweet spot

32 AA·18.9GB·256K ctx·

HunyuanImage 3.0 Instruct

8 個月前

Tencent · 80B · Tencent Hunyuan Community

Reasoning image model — prompt rewrite, chain-of-thought and image-to-image editing

41.5GB·4K ctx·

Llama 4 Scout 17B

1 年前

Meta · 109B · Llama 4 Community

MoE with 16 experts, 17B active params

10 AA·56.3GB·128K ctx·

GPT-OSS 120B

1 年前

OpenAI · 117B · Apache 2.0

OpenAI's flagship open-weight MoE — 52.6% SWE-bench

24 AA·60.4GB·128K ctx·

Mistral Small 4 119B

6 個月前

Mistral AI · 119B · Apache 2.0

Sparse Mistral Small 4 — 6.5B active, strong local all-rounder

20 AA·61.5GB·256K ctx·

Qwen 3 VL 235B-A22B

10 個月前

Alibaba · 235B · Apache 2.0

Flagship vision-language MoE — frontier multimodal reasoning and agentic GUI control

120.9GB·256K ctx·

Hy3

2 個月前

Tencent · 295B · Apache 2.0

Production-focused agentic MoE with strong coding, tool use and long-context reasoning

42 AA·151.6GB·256K ctx·

MiniMax M3

3 個月前

MiniMax · 428B · MiniMax Community

Native multimodal MoE — understands text, image and long video with 1M context

45 AA·219.7GB·1024K ctx·

GLM-5.3

1 個月前

Z.ai · 753B · MIT

Same 753B / 40B-active base as GLM-5.2 — post-training lifts coding and long-horizon agents, 1M context

386.2GB·1024K ctx·

LongCat 2.0

2 個月前

Meituan · 1.6T · MIT

Frontier-scale agentic and coding MoE with sparse attention and native 1M context

34 AA·820.1GB·1024K ctx·

DeepSeek V4 Pro

5 個月前

DeepSeek · 1.6T · MIT

Flagship V4 MoE — 49B active, 1M context

820.1GB·1024K ctx·

Qwen 3.8 2.4T-A95B

1 個月前

Alibaba · 2.4T · Qwen

Frontier Qwen 3.8 MoE — 95B active, 1M context

58 AA·1229.8GB·1024K ctx·

Kimi K3

2 個月前

Moonshot AI · 2.8T · Kimi

Frontier 2.8T multimodal MoE — 104B active, native video understanding, 1M context

60 AA·1424.5GB·1024K ctx·
全部模型

Qwen 3 0.6B

1 年前

Alibaba · 0.6B · Apache 2.0

Ultra-light Qwen 3 model for constrained devices

0.8GB·32K ctx·

Qwen 3.5 0.8B

7 個月前

Alibaba · 0.8B · Apache 2.0

Ultra-tiny model for embedded and edge

5* AA·0.9GB·32K ctx·

Llama 3.2 1B

2 年前

Meta · 1B · Llama 3.2 Community

Meta's smallest Llama for edge devices

1GB·128K ctx·

Gemma 3 1B

1 年前

Google · 1B · Gemma

Google's tiny Gemma for on-device

1GB·32K ctx·

Wan 2.1 T2V 1.3B

1 年前

Alibaba · 1.3B · Apache 2.0

Tiny open text-to-video — 480p clips on 8GB consumer GPUs

1.2GB·4K ctx·

Qwen 2.5 Coder 1.5B

1 年前

Alibaba · 1.5B · Apache 2.0

Ultra-lightweight coding model

1.3GB·32K ctx·

DeepSeek R1 1.5B

1 年前

DeepSeek · 1.5B · MIT

Tiny reasoning model distilled from R1

1.3GB·64K ctx·

Qwen 3 1.7B

1 年前

Alibaba · 1.7B · Apache 2.0

Compact multilingual Qwen 3

1.4GB·32K ctx·

Qwen 3.5 2B

7 個月前

Alibaba · 2B · Apache 2.0

Small multimodal Qwen 3.5

7* AA·1.5GB·32K ctx·

Llama 3.2 3B

2 年前

Meta · 3B · Llama 3.2 Community

Lightweight Llama for mobile and edge

2GB·128K ctx·

SmolLM3 3B

1 年前

HuggingFace · 3B · Apache 2.0

Lightweight multilingual reasoning

2GB·128K ctx·

Granite 4.1 3B

5 個月前

IBM · 3B · Apache 2.0

Compact enterprise model for edge and constrained environments

2GB·128K ctx·

Ministral 3 3B

9 個月前

Mistral AI · 3B · Apache 2.0

Current-gen tiny Ministral — edge chat with 256K context

7 AA·2GB·256K ctx·

Phi-4 Mini Reasoning

1 年前

Microsoft · 3.8B · MIT

Lightweight reasoning model

2.4GB·16K ctx·

Gemma 3 4B

1 年前

Google · 4B · Gemma

Multimodal Gemma with 128K context

2.5GB·128K ctx·

Qwen 3.5 4B

7 個月前

Alibaba · 4B · Apache 2.0

Small multimodal Qwen 3.5

20* AA·2.5GB·32K ctx·

Qwen3-VL 4B

11 個月前

Alibaba · 4.4B · Apache 2.0

Compact dedicated vision-language model — OCR & image chat on edge

2.8GB·256K ctx·

Gemma 4 E2B IT

5 個月前

Google · 5B · Gemma

Gemma 4 efficient instruct model (official)

10* AA·3.1GB·256K ctx·

Qwen 2.5 Coder 7B

1 年前

Alibaba · 7B · Apache 2.0

Dedicated coding model

4.1GB·128K ctx·

DeepSeek R1 Distill 7B

1 年前

DeepSeek · 7B · MIT

R1 reasoning distilled into Qwen 7B

4.1GB·64K ctx·

Gemma 4 E4B IT

5 個月前

Google · 8B · Gemma

Gemma 4 balanced instruct model (official)

12* AA·4.6GB·256K ctx·

Llama 3.1 8B

2 年前

Meta · 8B · Llama 3.1 Community

Meta's versatile 8B — great quality/speed ratio

4.6GB·128K ctx·

Qwen 3 8B

1 年前

Alibaba · 8B · Apache 2.0

Qwen 3 with thinking mode support

4.6GB·128K ctx·

Granite 4.1 8B

5 個月前

IBM · 8B · Apache 2.0

Balanced general-purpose enterprise model

4.6GB·128K ctx·

Ministral 8B

2 年前

Mistral AI · 8B · MRL

Mistral's efficient 8B model

4.6GB·32K ctx·

GLM-4 9B

2 年前

Zhipu AI · 9B · GLM-4

Multilingual model supporting 26 languages with 128K context

5.1GB·128K ctx·

Nemotron Nano 9B v2

1 年前

NVIDIA · 9B · NVIDIA Open

Hybrid Mamba2 architecture for reasoning

9* AA·5.1GB·128K ctx·

Ornith 1.0 9B

3 個月前

DeepReinforce · 9B · MIT

Self-improving agentic coding model optimized for terminal and software engineering tasks

5.1GB·256K ctx·

FLUX.2 Klein 9B

8 個月前

Black Forest Labs · 9B · FLUX Non-Commercial

Higher-quality distilled FLUX.2 — sub-second generation and multi-reference editing

5.1GB·32K ctx·

Gemma 3 12B

1 年前

Google · 12B · Gemma

Multimodal Gemma with 128K context

6.6GB·128K ctx·

Mistral Nemo 12B

2 年前

Mistral AI · 12B · Apache 2.0

Multilingual 12B with 128K context

6.6GB·128K ctx·

Gemma 4 12B IT

5 個月前

Google · 12B · Apache 2.0

Gemma 4 mid-size instruct — multimodal any-to-any

22* AA·6.6GB·256K ctx·

Phi-4 14B

1 年前

Microsoft · 14B · MIT

Microsoft's reasoning-focused model

5* AA·7.7GB·16K ctx·

Qwen 3 14B

1 年前

Alibaba · 14B · Apache 2.0

Strong all-rounder with thinking mode

7.7GB·128K ctx·

DeepSeek R1 Distill 14B

1 年前

DeepSeek · 14B · MIT

R1 reasoning distilled into Qwen 14B

7.7GB·64K ctx·

Ministral 3 14B

9 個月前

Mistral AI · 14B · Apache 2.0

Current-gen Ministral mid-size — local assistant with 256K context

11 AA·7.7GB·256K ctx·

LFM2 24B

10 個月前

Liquid AI · 24B · Liquid AI

Hybrid MoE with convolution+attention layers — 2.3B active

5* AA·12.8GB·32K ctx·

Devstral Small 2 24B

9 個月前

Mistral AI · 24B · Apache 2.0

Coding-focused model with 256K context — 68% SWE-bench

12.8GB·256K ctx·

Mistral Small 3.1 24B

1 年前

Mistral AI · 24B · Apache 2.0

Multimodal Mistral with vision support

12.8GB·128K ctx·

DiffusionGemma 26B-A4B IT

3 個月前

Google · 26B · Apache 2.0

Discrete diffusion MoE — 1100+ tok/s on H100, multimodal (text/image/video)

13* AA·13.8GB·256K ctx·

Qwen 3.5 27B

7 個月前

Alibaba · 27.8B · Apache 2.0

Flagship native multimodal Qwen 3.5

14.7GB·256K ctx·

Qwen 3.6 27B

5 個月前

Alibaba · 27.8B · Apache 2.0

Flagship dense Qwen 3.6 — native multimodal all-rounder

38 AA·14.7GB·256K ctx·

Qwen 3 30B-A3B

1 年前

Alibaba · 30B · Apache 2.0

MoE with only 3.3B active — extremely efficient

15.9GB·128K ctx·

Nemotron 3 Nano 30B

1 年前

NVIDIA · 30B · NVIDIA Open

MoE with 1M context and 3B active

15 AA·15.9GB·1024K ctx·

Granite 4.1 30B

5 個月前

IBM · 30B · Apache 2.0

High-capacity enterprise model for complex reasoning and tool use

15.9GB·128K ctx·

North Mini Code

3 個月前

Cohere · 30B · Apache 2.0

Open agentic coding MoE with 3B active — built for software engineering and terminal tasks

15.9GB·256K ctx·

Qwen 3 Coder 30B-A3B

1 年前

Alibaba · 30B · Apache 2.0

Efficient agentic coding MoE — 3B active, 256K context

15.9GB·256K ctx·

Qwen 3 32B

1 年前

Alibaba · 32B · Apache 2.0

Qwen 3 flagship dense model

16.9GB·128K ctx·

DeepSeek R1 Distill 32B

1 年前

DeepSeek · 32B · MIT

R1 reasoning distilled into Qwen 32B — sweet spot

16.9GB·64K ctx·

OLMo 2 32B

1 年前

Allen AI · 32B · Apache 2.0

Fully open research model by Allen AI

16.9GB·4K ctx·

Gemma 4 31B IT

5 個月前

Google · 33B · Gemma

Gemma 4 flagship instruct model (official)

30 AA·17.4GB·256K ctx·

Gemma 4 31B

5 個月前

Google · 33B · Gemma

Gemma 4 flagship base model (official)

17.4GB·256K ctx·

Command R 35B

2 年前

Cohere · 35B · CC BY-NC 4.0

Optimized for retrieval-augmented generation

18.4GB·128K ctx·

Qwen 3.5 35B-A3B

7 個月前

Alibaba · 35B · Apache 2.0

Efficient multimodal MoE with 3B active

18.4GB·256K ctx·

Mixtral 8x7B

2 年前

Mistral AI · 47B · Apache 2.0

MoE with 12.9B active params

24.6GB·32K ctx·

Llama 3.3 70B

1 年前

Meta · 70B · Llama 3.3 Community

Best open model at 70B class

9* AA·36.4GB·128K ctx·

Qwen 3 Next 80B-A3B

9 個月前

Alibaba · 80B · Apache 2.0

High-sparsity MoE — extreme low activation ratio for fast inference at 80B scale

41.5GB·256K ctx·

Qwen 3 Coder Next 80B-A3B

7 個月前

Alibaba · 80B · Apache 2.0

Ultra-efficient agentic coding MoE optimized for tool-calling coding agents

41.5GB·256K ctx·

HunyuanImage 3.0

1 年前

Tencent · 80B · Tencent Hunyuan Community

Largest open image MoE — 13B active, strong long-prompt generation

41.5GB·4K ctx·

GLM-4.5 Air

1 年前

Z.ai · 106B · MIT

Consumer-friendly GLM MoE — 12B active, strong agentic & tool use

54.8GB·128K ctx·

Qwen 3.5 122B-A10B

7 個月前

Alibaba · 122B · Apache 2.0

Large multimodal MoE

33 AA·63GB·256K ctx·

DeepSeek V4 Flash

5 個月前

DeepSeek · 158B · MIT

Efficient long-context V4 — 13B active, 1M context

81.4GB·1024K ctx·

Qwen 3 235B-A22B

1 年前

Alibaba · 235B · Apache 2.0

Massive MoE with 22B active — frontier quality

120.9GB·128K ctx·

GLM-4.6

1 年前

Z.ai · 357B · MIT

Large GLM MoE with strong coding and 200K context

183.4GB·195K ctx·

Qwen 3.5 397B-A17B

7 個月前

Alibaba · 397B · Apache 2.0

Largest multimodal Qwen 3.5 MoE

34 AA·203.9GB·256K ctx·

Llama 4 Maverick 17B-128E

1 年前

Meta · 400B · Llama 4 Community

Multimodal MoE with 128 experts — 17B active, 1M context

14 AA·205.4GB·1024K ctx·

Qwen 3 Coder 480B

1 年前

Alibaba · 480B · Apache 2.0

Largest open coding MoE — 35B active

246.4GB·256K ctx·

DeepSeek R1

1 年前

DeepSeek · 671B · MIT

Massive MoE reasoning model — 37B active

344.2GB·64K ctx·

DeepSeek V3.2

9 個月前

DeepSeek · 685B · MIT

State-of-the-art MoE — 37B active params

351.4GB·128K ctx·

GLM-5

7 個月前

Zhipu AI · 744B · MIT

MoE with 256 experts, 40B active — frontier-class agentic coding

381.6GB·128K ctx·

GLM-5.2

3 個月前

Z.ai · 753B · MIT

Frontier open-weight coder — top SWE-bench, 1M context

53 AA·386.2GB·1024K ctx·

GLM-5.1

5 個月前

Zhipu AI · 754B · MIT

Improved agentic coding — SOTA SWE-bench Pro, long-horizon tasks

386.7GB·128K ctx·

Kimi K2.6

5 個月前

Moonshot AI · 1.06T · Kimi

Natively multimodal 1T MoE — 32B active, frontier agentic

542.4GB·256K ctx·

Common questions

What is a local LLM?
A local LLM is an open-weight language model that runs on your computer instead of a cloud API. Prompts stay on the device, there is no usage meter, and you can keep working offline.
How much VRAM do I need for a local LLM?
A 7B–9B chat model usually fits in 8 GB of VRAM at Q4. 12–16 GB covers most 12B–27B models. 24 GB and up opens 30B dense models and mid-size mixture-of-experts.
Can I run a local LLM on a Mac?
Yes. Apple Silicon shares unified memory between CPU and GPU, so an M-series Mac can run the same open models that would need a discrete GPU on Windows.

midudev 為本機 AI 社群製作GitHub

数值由瀏覽器 API 估算,實際規格可能不同。WebGPU資料来源: llama.cpp, OllamaLM Studio.

所有產品名称、標誌与品牌均為其各自所有者的財產。Apple、NVIDIA、AMD、Intel、Qualcomm 以及本站提及的所有 AI 模型名称均為其各自持有者的商標或註冊商標。本站与上述任何公司均無關聯,亦未獲其背書。