GPT-OSS 120B

Apache 2.0

OpenAI · 117B (5.1B active) · 専家混合(MoE)

OpenAI's flagship open-weight MoE — 52.6% SWE-bench 看看你的 GPU 或 Mac 跑不跑得動 GPT-OSS 120B——最低 65.4 GB,建議 109 GB。

2025-08128K context

専家混合(MoE)

専家総数: 16
启用専家: 2
启用参数: 5.1B

量化選項

量化位元VRAM品質状態
Q2_K238 GBlow
Q3_K_M352.9 GBmoderate
Q4_K_M460.4 GBgood
Q5_K_M575.4 GBgood
Q6_K690.4 GBexcellent
Q8_08120.4 GBexcellent
F1616240.2 GBlossless

関于这個模型

OpenAI gpt-oss banner

Welcome OpenAI’s gpt-oss!

Ollama partners with OpenAI to bring its latest state-of-the-art open weight models to Ollama. The two models, 20B and 120B, bring a whole new local chat experience, and are designed for powerful reasoning, agentic tasks, and versatile developer use cases.

Get started

You can get started by downloading the latest Ollama version.

The model can be downloaded directly in Ollama’s new app or via the terminal:

ollama run gpt-oss:20b

ollama run gpt-oss:120b

Feature highlights

  • Agentic capabilities: Use the models’ native capabilities for function calling, web browsing (Ollama is introducing built-in web search that can be optionally enabled), python tool calls, and structured outputs.
  • Full chain-of-thought: Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in outputs.
  • Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs.
  • Fine-tunable: Fully customize models to your specific use case through parameter fine-tuning.
  • Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment.

benchmark

Quantization - MXFP4 format

OpenAI utilizes quantization to reduce the memory footprint of the gpt-oss models. The models are post-trained with quantization of the mixture-of-experts (MoE) weights to MXFP4 format, where the weights are quantized to 4.25 bits per parameter. The MoE weights are responsible for 90+% of the total parameter count, and quantizing these to MXFP4 enables the smaller model to run on systems with as little as 16GB memory, and the larger model to fit on a single 80GB GPU.

Ollama is supporting the MXFP4 format natively without additional quantizations or conversions. New kernels are developed for Ollama’s new engine to support the MXFP4 format.

Ollama collaborated with OpenAI to benchmark against their reference implementations to ensure Ollama’s implementations have the same quality.

20B parameter model

gpt-oss 20B

gpt-oss-20b model is designed for lower latency, local, or specialized use-cases.

120B parameter model

gpt-oss 120B

Reference

Can I run GPT-OSS 120B locally?

Can I run GPT-OSS 120B locally?
GPT-OSS 120B needs about 65.4 GB of memory at a minimum and 109 GB recommended. Open this page to grade it against your GPU or Mac, then run it with runai, Ollama or LM Studio.
How much VRAM does GPT-OSS 120B need?
At Q4_K_M, GPT-OSS 120B uses about 60.4 GB of VRAM. Higher quants need more memory; lower quants fit tighter cards with a quality tradeoff.