openai/gpt-oss
52.2
Adequate · 19 September 2026
17.3k
lines of production code
Python
with C, C++
1
measurement over time
What this system is
This system is an open-source large language model platform centered on the gpt-oss family, providing infrastructure for local inference, fine-tuning, and agent-based interaction. It features a comprehensive Responses API that supports tool calling, code interpretation, and web browsing, alongside specialized backends for Apple Metal and Triton to optimize performance on local hardware. The project also includes a Python Agents SDK, Model Context Protocol (MCP) server implementations, and evaluation frameworks to facilitate development and testing of AI applications.
Features
Add Gradio chat interface example
A new Gradio-based chat interface example has been added to the examples directory. This script provides a user-facing web UI for interacting with a local model API, featuring a chat history view, message input, and configuration options for model selection (large/small), function calling, browser search, and reasoning effort. It handles streaming responses, including reasoning text and tool calls (function and web search), displaying them in the chat stream.
examples/gradio · high confidence
Add MCP server reference implementation for gpt-oss tools
Introduces a new \gpt-oss-mcp-server\ directory containing reference implementations that expose gpt-oss tools (Python execution and web browsing) via the Model Context Protocol (MCP). This includes \python\_server.py\ and \browser\_server.py\ to run as MCP services, a \build-system-prompt.py\ script to demonstrate automatic tool discovery and system prompt generation compatible with OpenAI Harmony, and a \reference-system-prompt.py\ for direct comparison. The browser tool supports configurable backends (You.com or Exa) via environment variables.
gpt-oss-mcp-server · high confidence
Add Python Agents SDK example with MCP and local LLM support
A new example script (example.py) demonstrates how to use the Python Agents SDK to run an agent connected to a local OpenAI-compatible server (such as Ollama) and an MCP filesystem server. The example shows configuring the SDK with a local client, defining a custom weather tool, and processing streaming events including agent updates, tool calls, and message outputs.
examples/agents-sdk-python · high confidence
Add You.com as a search backend for the simple browser tool
Users can now use You.com as a search provider in addition to the existing Exa backend. This change introduces the YouComBackend class, which integrates with the You.com Search API (https://api.ydc-index.io) to fetch web and news search results, and supports fetching page content via the /v1/contents endpoint. The backend requires a You.com API key, which can be provided via the constructor or the YDC\_API\_KEY environment variable. Additionally, HTTP requests to the backend now include a User-Agent header identifying the gpt-oss version, aiding in request tracking.
_gpt\_oss/tools/simple\browser · high confidence
Added API compatibility test suite for validating tool calling and response shapes
A new \compatibility-test\ directory has been introduced to verify that API providers correctly implement tool calling and return valid response shapes. The suite uses the Agents SDK to run a set of predefined cases (covering tools like weather, database queries, and chart generation) against configurable providers (defaulting to a local vLLM instance). It validates that tool calls match expected schemas and arguments, checks for valid response structures for both Chat Completions and Responses APIs, and optionally verifies streaming events. Results are output as JSONL files with a summary analysis including pass@k metrics and counts of invalid responses or wrong-input tool calls.
compatibility-test · high confidence
Introduce custom PEP 517 build backend for pure-wheel and Metal support
A new in-tree build backend (\gpt\_oss\_build\_backend\) has been added to resolve PyPI upload rejections for Linux binary wheels. By default, the project now builds a pure Python wheel (\py3-none-any\) compatible with PyPI. Developers can opt into building native Metal/C extensions locally by setting the \GPTOSS\_BUILD\_METAL\ environment variable, which dynamically switches the backend to \scikit-build\_core\ and injects necessary build tools like CMake and Ninja. This change also adds support for editable installs via the \build\_editable\ hook.
_\build · high confidence
Metal backend introduces MoE kernels, per-context batch sizing, and async pipeline compilation
The Metal backend now supports Mixture-of-Experts (MoE) inference by adding specialized kernels for expert routing, scatter, and gather/accumulate operations, which are registered and launched during model initialization. To improve memory efficiency and concurrency, the \max\_batch\_tokens\ parameter has moved from the global model scope to the per-context scope, allowing each context to allocate activation buffers based on its own batch size rather than a fixed global default. Additionally, Metal compute pipeline state creation is now asynchronous using completion handlers and semaphores, and model weights are locked in memory via \mlock\ to prevent paging, while a control buffer enables safe kernel abortion.
_gpt\oss/metal/source · high confidence
New reinforcement fine-tuning example for gpt-oss
Added a new Jupyter notebook example in the \examples\ directory that demonstrates how to fine-tune the \gpt-oss-20b\ model using reinforcement learning (specifically the GRPO algorithm) to play the 2048 game. The notebook provides a complete, runnable guide using Unsloth for efficient training on Google Colab, including installation steps, model loading configuration, and training logic.
examples · high confidence
Behavioural changes
Add code interpreter support and fix UI bugs in Streamlit chat demo
The Streamlit chat example now supports a new 'Code Interpreter' tool, allowing users to toggle it in the sidebar and view execution status, code, and outputs (logs and images) directly in the chat interface. The maximum output token limit has been increased from 20,000 to 131,072. Additionally, the 'Browser Search' toggle is now always visible regardless of query parameters, and several f-string syntax errors and variable naming issues in the chat rendering logic have been fixed.
examples/streamlit · high confidence
Add code interpreter support and increase max output tokens
The Responses API now supports a code interpreter tool, allowing models to execute code and return logs or images as outputs; this is reflected in new event types (e.g., \ResponseCodeInterpreterCallInProgress\) and request/response types (\CodeInterpreterCallItem\, \CodeInterpreterToolConfig\). The maximum output token limit has been increased from 10,000 to 131,072. Additionally, the browser tool now defaults to the You.com backend (configurable via \BROWSER\_BACKEND\), and the browser MCP server exposes search, open, and find capabilities on port 8001.
gpt-oss · high confidence
Configurable browser backend and generation parameters
The chat interface now defaults to the You.com browser backend instead of Exa, with the ability to switch back to Exa via the BROWSER\_BACKEND environment variable. In generation, the context length for the Triton backend and tensor parallel size for the vLLM backend are now configurable via command-line arguments, and the max\_tokens limit is handled more robustly to allow unlimited generation when set to zero. Additionally, developer messages are no longer included in the chat history by default to reduce mismatch, and the demo output only prints developer message instructions when they are actually provided.
_gpt\oss · high confidence
Context now supports batch token generation and configurable batch size
The Python Context API has been updated to allow generating multiple tokens in a single call. The \Context\ constructor now accepts an optional \max\_batch\_tokens\ argument to define the batch size, which is passed to the underlying Metal engine. Additionally, the \Context.sample()\ method has changed its return behavior: instead of returning a single token ID, it now accepts a \max\_output\_tokens\ argument and returns a list of token IDs, enabling batched inference directly from the Python wrapper.
_gpt\oss/metal/python · high confidence
Evals framework adds Chat Completions sampler, basic eval, and configurable reasoning effort
The evaluation framework now supports sampling via a new Chat Completions API sampler (replacing the previous hardcoded reliance on the Responses API for all model calls), allowing users to select the backend via the \--sampler\ argument. A new \BasicEval\ is available for simple prompt-response evaluations. The \--model\ and \--reasoning-effort\ arguments now accept comma-separated lists, enabling batch evaluation across multiple models and reasoning effort levels (low, medium, high) in a single run. Additionally, the default \max\_tokens\ for reasoning models has been increased to 131,072, and model names containing slashes are now sanitized with double underscores in output filenames.
_gpt\oss/evals · high confidence
Introduce Transformers backend and refactor Metal/Ollama inference adapters
Adds a new Transformers-based inference backend that generates tokens one at a time using \model.generate\ to mimic existing behavior. Refactors the Metal adapter to simplify token generation by leveraging the underlying \Context\ for internal KV-cache reuse and batch sampling, removing manual longest-common-prefix logic. Updates the Ollama adapter to remove deprecated context-passing logic, adds a \raw: True\ flag to API requests, and fixes minor typos in comments.
_gpt\_oss/responses\api/inference · high confidence
Metal backend refactors kernel interfaces and model structures for optimized MoE and sampling
The Metal backend's internal headers have been significantly restructured to support fused kernels, Mixture-of-Experts (MoE) optimizations, and GPU-based sampling. New kernel argument structures (e.g., \gptoss\_dense\_matmul\_args\, \gptoss\_moe\_dense\_matmul\_swiglu\_args\) and command-buffer encoding functions (e.g., \gptoss\_metal\_command\_buffer\_encode\_launch\_f32\_bf16w\_dense\_matmul\_qkv\) enable fused QKV projections and optimized MoE MLP operations. The \gptoss\_model\ struct now includes threadgroup sizes for various kernels and a \lock\_memory\ flag, while the \gptoss\_context\ struct has moved \max\_batch\_tokens\ from the model to the context and added a \control\_buffer\ for kernel synchronization. Additionally, the \gptoss\_sampler\ struct has been removed, indicating sampling logic has been moved to the GPU via new \gptoss\_sample\_args\ and \f32\_sample\_fn\ kernel support.
_gpt\oss/metal/source/include · high confidence
Metal backend supports batched token generation and configurable prefill batch size
The Metal backend's C API has been updated to allow generating multiple tokens in a single call via \gptoss\_context\_sample\, which now accepts a \max\_tokens\ limit and outputs an array of tokens (\tokens\_out\) along with the count (\num\_tokens\_out\). Additionally, \gptoss\_context\_create\ now accepts a \max\_batch\_tokens\ parameter to control the maximum number of tokens processed in a single prefill batch, allowing users to tune performance versus memory usage. Several documentation typos in the header file were also corrected.
_gpt\oss/metal/include · high confidence
Metal benchmark suite: end-to-end tests and threadgroup tuning
The Metal benchmark directory now includes new end-to-end benchmarks (end-to-end.cc and end-to-end-threadgroup.cc) that measure generation performance over 100 tokens and allow tuning of attention and MLP threadgroup sizes for GPT- OSS models. Existing kernel benchmarks, such as the f32-bf16w-rmsnorm test, have been updated to pass a control buffer to the underlying Metal kernel, ensuring the benchmark accurately reflects the current API requirements.
_gpt\oss/metal/benchmark · high confidence
Python code execution support and validation error logging in responses API
The responses API now supports executing Python code via a new PythonTool, introducing dedicated state tracking for code interpreter calls (including logs and images) and reasoning outputs. The API server also adds a custom exception handler to log the raw body of invalid requests for better debugging, and the serve script now allows selecting the 'transformers' inference backend.
_gpt\_oss/responses\api · high confidence
Python tool now supports multiple execution backends including uv and local Jupyter
Users can now choose how Python code is executed within the tool by setting the PYTHON\_EXECUTION\_BACKEND environment variable. The default remains the existing Docker-based execution, but two new options are available: 'dangerously\_use\_uv' for running scripts via the uv package manager on the host, and 'dangerously\_use\_local\_jupyter' for stateful execution through a local Jupyter kernel. This change also improves reliability by capturing stderr in the uv backend and providing a clearer warning when no stdout output is generated in the Docker backend.
_gpt\_oss/tools/python\docker · high confidence
Refactor attention kernel to use TensorDescriptor and modern indexing
The Triton attention kernel has been updated to replace manual stride-based offset calculations and \tl.make\_block\_ptr\ with \TensorDescriptor\ and direct \.load\/\.store\ indexing. This simplifies the kernel signature by removing explicit stride arguments and changes how data is accessed, which may affect performance characteristics or numerical behavior for users relying on the previous implementation details.
_gpt\oss/triton · high confidence
Repository documentation and licensing updates
The repository now includes a USAGE\_POLICY file outlining responsible use and compliance requirements, and the LICENSE has been updated (with formatting adjustments to the Apache 2.0 text). The README has been significantly restructured with a table of contents, updated model descriptions (specifying 80GB GPU requirements for gpt-oss-120b), and new vLLM offline serve code examples. Additionally, the awesome-gpt-oss.md list has been expanded to include resources for AMD, llama.cpp, TVM, AWS, and Google Colab, while MANIFEST.in was added to include build artifacts.
(repo-wide) · high confidence
Support for sharded checkpoints in the local model converter
The \create-local-model.py\ script now reads weights from multiple \.safetensors\ files within a checkpoint directory instead of relying on a single \model.safetensors\ file. This allows the converter to process sharded model checkpoints, ensuring that all tensor shards are correctly located and merged into the final local model file.
_gpt\oss/metal/scripts · high confidence
Fixes
Fix import path for Metal example script
The generate.py example script now correctly imports Context and Model from the gpt\_oss.metal module, resolving a previous import error that likely prevented the example from running.
_gpt\oss/metal/examples · high confidence
Test coverage
Added test suite for Responses API and browser tool backends; Added tests for optimized Metal prefill kernels and updated kernel interfaces.
Dependencies
Updated build system and added Python agent examples
The main project's build system has been migrated from scikit-build-core to setuptools with a custom build backend (gpt\_oss\_build\_backend), and the required Python version constraint has been relaxed from \<3.13 to \>=3.12. Additionally, new example directories have been added: a Python Agents SDK example depending on openai-agents, a GPT-oss MCP server example depending on mcp, and a Node.js compatibility test directory depending on @openai/agents, ajv, and listr2.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 52.
Lenses
- Code Health 67
- Architecture 100
- Maturity 54
- Readiness 36
- Security 89
Changes since last survey
- 130 commits — 105 feature/other, 25 fixes
By area
- (root) — 61 commits
- gpt_oss/metal — 25 commits
- gpt_oss/tools — 8 commits
- gpt_oss/responses_api — 7 commits
- gpt_oss/evals — 6 commits
- (repo) — 5 commits
- examples/agents-sdk-python — 3 commits
- .github/workflows — 2 commits
- gpt_oss/chat.py — 2 commits
- gpt_oss/generate.py — 2 commits
- .github/CODEOWNERS — 1 commit
- _build/gpt_oss_build_backend — 1 commit
- examples/gradio — 1 commit
- examples/streamlit — 1 commit
- gpt-oss-mcp-server/README.md — 1 commit
- gpt-oss-mcp-server/build-system-prompt.py — 1 commit
- gpt-oss-mcp-server/python_server.py — 1 commit
- gpt_oss/triton — 1 commit
- tests/conftest.py — 1 commit
Notable commits
- fix: Fix TOML parsing errors in pyproject.toml for scikit-build configuration (#27)
- fix: Fix chat demo (#26)
- fix: Fix import for metal example (#24)
- fix: Fix start_q use in upper bound calculation (#136)
- fix: Merge pull request #16 from openai/zhuohan/fix-pypi-ci
- fix: Metal: fix KV-cache invalidation after reset+append (#163)
- fix: Metal: fix resource leak in prefill (#194)
- fix: Try fix pypi ci
- fix: Try fix pypi ci (#13)
- fix: a few typo fixes. (#102)
- fix: first fix — readme (#2)
- fix: fix build
- fix: fix ci/pypi (#30)
- fix: fix editable build (#113)
- fix: fix f string errors in streamlit chat (#73)
- fix: fix invalid import in build-system-prompt.py (#32)
- fix: fix packaging (#90)
- fix: fix streamlit & ollama demo. Add python tool (#131)
- fix: fix: Add channel parameter to PythonTool response handling (#33)
- fix: fix: Correct broken links in awesome-gpt-oss.md (#12)
- …and 110 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
openai/gpt-oss was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 7b583341fe16729127f6d5b94a7b09ccae97e1a1 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.