Dedicated model · Available as managed deployment

Request a gpt-oss deployment on your own DGX Spark

OpenAI's open-weight reasoning models — 20B and 120B mixture-of-experts — built for agentic work, tool use and long chains of reasoning, under Apache-2.0. Validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.

eu-es-1 · Málaga Available as managed deployment Text Quoted per deployment gpt-oss.axforge.ai
Request deploymentTalk to an engineerSign in €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT Hardware rental plus a managed service quoted per deployment — both confirmed in writing before anything is billed.

Why AxForge

Why gpt-oss as a managed deployment

OpenAI, open weightsThe first open-weight models from OpenAI: two sizes, both released under Apache-2.0, with configurable reasoning effort and native tool calling.
20B on one machinegpt-oss-20b is sized for a single accelerator; on a dedicated DGX Spark it runs as your own OpenAI-compatible endpoint. The 120B build is a multi-GPU deployment.
Agents, not just chatStructured outputs, function calling and a reasoning mode you can dial down for latency — the pieces an agent workflow needs, on hardware only you use.

Specifications

What you get

Modelgpt-oss — openai
ModalitiesText
Sizes20.9B, 116.8B
LicenceOpen weights — apache-2.0; commercial use permitted
HardwareNVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge
Rental termHour, week, month or year
Hardware pricing€0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT
Managed serviceQuoted per deployment
RegionMálaga, Spain (eu-es-1)

Full details, benchmarks and FAQ on the gpt-oss page. Prices exclude VAT.

How it works

From sign-in to running

1Request deployment — describe your traffic, context needs and rental term.
2You receive the configuration, hardware rental and managed-service price in writing before anything is billed.
3AxForge deploys gpt-oss on a dedicated DGX Spark reserved for you.
4Point your OpenAI SDK at your own endpoint with the model name you receive.
5Adjust the term — hour, week, month or year — as your workload settles.

Request deployment or sign in to start.

FAQ

gpt-oss — common questions

Is gpt-oss on the AxForge serverless API?

Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.

Which gpt-oss size should I deploy?

gpt-oss-20b for a single dedicated machine and low latency; gpt-oss-120b when quality on hard reasoning matters more than cost — it needs a multi-GPU system, which AxForge scopes with you.

Does gpt-oss support tool calling?

Yes — function calling and structured outputs are part of the model, and they work through the OpenAI-compatible endpoint on your machine.

How fast is it on your hardware?

AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.

What does EU gpt-oss hosting cost?

Hardware by the hour, week, month or year; the managed service is quoted per deployment — both confirmed in writing before anything is billed.

Ready for gpt-oss on your own machine?

Request deployment Sign in Talk to an engineer

Explore

More from AxForge

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms