Pricing aligned to workload behavior. Token-based for inference, flexible for dedicated setups. Designed for predictable cost under real usage.
Explore the GPT-OSS-120B API for model-specific pricing and integration details.
Peak traffic and request duration are used only for the rate-limit check. Monthly cost is based on total requests and billed units.
Usage-based estimate before applicable taxes or negotiated discounts.
Deploy in our US and EU regions. Deployment in APAC is underway.
Enterprise-grade security and data isolation for all workloads.
One SDK for both serverless inference and dedicated compute.
For large-scale deployments, custom SLAs, or multi-region clusters, our enterprise team can provide volume discounts and tailored infrastructure solutions.