Home/tools/OpenAI Reduces GPT‑5.6 Sol Input Cost by 20 % and Output Cost by One‑Third
A pencil sketch of a server rack labeled “GPT‑5.6 Sol” with price tags showing $4 per million input and $20 per million output, no text, no logos.
ToolsPublished 22 August 20263 min read

OpenAI Reduces GPT‑5.6 Sol Input Cost by 20 % and Output Cost by One‑Third

OpenAI announced a new pricing structure for its GPT‑5.6 Sol model that lowers the cost of both input and output tokens.

GPT‑5.6 Sol is positioned as the frontier model in the GPT‑5.6 family, intended for complex professional tasks.

The model can process up to 1,050,000 tokens of context and generate up to 128,000 tokens in a single response.

Its knowledge base is current through 16 February 2026, providing relatively recent information for users.

Developers can select reasoning effort from none to max, with medium set as the default.

Compared with GPT‑5.5, GPT‑5.6 Sol’s input price of $4 per million tokens is 20 % lower than the $5 rate.

Its output price of $20 per million tokens is 33 % lower than the $30 rate charged for GPT‑5.5.

GPT‑5.4 remains cheaper at $2.50 per million input tokens, placing GPT‑5.6 Sol between the two tiers.

The promotional pricing is guaranteed at least until 21 November 2026, giving users a multi‑year window of lower costs.

For requests that exceed 272,000 input tokens, OpenAI applies a multiplier of 2× for input and 1.5× for output pricing.

Cached input tokens are billed at $0.40 per million, while uncached input remains at $4 per million.

Cache writes incur a charge of 1.25× the uncached input token rate.

Pricing Changes and Structure

GPT‑5.6 Sol’s pricing is calculated per token, with separate rates for input, cached input, and output.

The model’s alias “gpt‑5.6” routes API requests directly to GPT‑5.6 Sol, simplifying integration.

These rates apply across all supported endpoints, including chat completions and embeddings.

Capabilities and Access

GPT‑5.6 Sol supports text input and output, as well as image input, expanding multimodal possibilities.

Audio, video, and pure audio input are not supported by this model.

The model is reachable through a wide range of endpoints, including chat completions, real‑time translation, and embeddings.

Streaming responses, function calling, and structured outputs are all enabled for GPT‑5.6 Sol.

Fine‑tuning is not currently offered, keeping the model in a fixed‑behavior state.

Users may lock a specific version of the model using snapshots, ensuring consistent performance across deployments.

Tool Support and Rate Limits

When accessed via the Responses API, GPT‑5.6 Sol can invoke tools such as web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search.

Rate limits are tiered, with the free tier unavailable, and paid tiers scaling from 5,000 requests per minute at Tier 1 to 15,000 at Tier 5.

Token limits per minute also increase with tier, reaching up to 40 million tokens for Tier 5.

Batch queue limits rise from 1,500,000 tokens at Tier 1 to 15 000 000 000 tokens at Tier 5, accommodating large‑scale workloads.

These limits ensure fair access while allowing higher‑spending customers to submit larger volumes of requests.

The combination of lower pricing, extensive token windows, and broad tool integration makes GPT‑5.6 Sol a cost‑effective option for enterprises building sophisticated AI applications.

Why This Matters: OpenAI’s reduced pricing for GPT‑5.6 Sol lowers cost barriers for developers seeking high‑capacity models.

#tools#ai#digest#auto

This digest was compiled from:

Share this digest

Share on XWhatsAppLinkedInTelegram

People Also Ask

Share your thoughts

Reactions, corrections, or insights — all welcome.

0/2000