Private AI
gpt-oss-safeguard is a first open weight reasoning model specifically trained for safety classification tasks to help classify text content based on customizable policies. As a fine-tuned version of gpt-oss, gpt-oss-safeguard is designed to follow explicit written policies that you provide. This enables bring-your-own-policy Trust & Safety AI, where your own taxonomy, definitions, and thresholds guide classification decisions. Well crafted policies unlock gpt-oss-safeguard's reasoning capabilities, enabling it to handle nuanced content, explain borderline decisions, and adapt to contextual factors.
Added Feb 23, 2026
Model weightsContext Window
128.0K
Max Output
16.4K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.30/1M
Capabilities
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
14.9
Coding Index
20.7
GPQA Diamond
Graduate-level scientific reasoning
68.8%
Better than 50% of models compared
HLE
Humanity's Last Exam
9.8%
Last updated Aug 4, 2026
Artificial AnalysisAuto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Better than 60% of models compared
IFBench
Instruction-following benchmark
65.1%
Better than 77% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
60.2%
Better than 57% of models compared
AA-LCR
Long context reasoning evaluation
30.7%
Better than 41% of models compared
SciCode
Python programming for scientific computing
34.4%
Better than 52% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
10.6%
Better than 45% of models compared
LiveCodeBench
Contamination-free coding benchmark
77.7%
Better than 89% of models compared
AIME 2025
American Invitational Mathematics Examination 2025
89.3%
Better than 88% of models compared
MMLU-Pro
Professional and academic subject knowledge
74.8%
Better than 47% of models compared