OpenAI Reduces GPT‑5.6 Sol Input Cost by 20 % and Output Cost by One‑Third
OpenAI announced a new pricing structure for its GPT‑5.6 Sol model that lowers the cost of both input and output tokens.
GPT‑5.6 Sol is positioned as the frontier model in the GPT‑5.6 family, intended for complex professional tasks.
The model can process up to 1,050,000 tokens of context and generate up to 128,000 tokens in a single response.
Its knowledge base is current through 16 February 2026, providing relatively recent information for users.
Developers can select reasoning effort from none to max, with medium set as the default.
Compared with GPT‑5.5, GPT‑5.6 Sol’s input price of $4 per million tokens is 20 % lower than the $5 rate.
Its output price of $20 per million tokens is 33 % lower than the $30 rate charged for GPT‑5.5.
GPT‑5.4 remains cheaper at $2.50 per million input tokens, placing GPT‑5.6 Sol between the two tiers.
The promotional pricing is guaranteed at least until 21 November 2026, giving users a multi‑year window of lower costs.
For requests that exceed 272,000 input tokens, OpenAI applies a multiplier of 2× for input and 1.5× for output pricing.
Cached input tokens are billed at $0.40 per million, while uncached input remains at $4 per million.
Cache writes incur a charge of 1.25× the uncached input token rate.
Pricing Changes and Structure
GPT‑5.6 Sol’s pricing is calculated per token, with separate rates for input, cached input, and output.
The model’s alias “gpt‑5.6” routes API requests directly to GPT‑5.6 Sol, simplifying integration.
These rates apply across all supported endpoints, including chat completions and embeddings.
Capabilities and Access
GPT‑5.6 Sol supports text input and output, as well as image input, expanding multimodal possibilities.
Audio, video, and pure audio input are not supported by this model.
The model is reachable through a wide range of endpoints, including chat completions, real‑time translation, and embeddings.
Streaming responses, function calling, and structured outputs are all enabled for GPT‑5.6 Sol.
Fine‑tuning is not currently offered, keeping the model in a fixed‑behavior state.
Users may lock a specific version of the model using snapshots, ensuring consistent performance across deployments.
Tool Support and Rate Limits
When accessed via the Responses API, GPT‑5.6 Sol can invoke tools such as web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search.
Rate limits are tiered, with the free tier unavailable, and paid tiers scaling from 5,000 requests per minute at Tier 1 to 15,000 at Tier 5.
Token limits per minute also increase with tier, reaching up to 40 million tokens for Tier 5.
Batch queue limits rise from 1,500,000 tokens at Tier 1 to 15 000 000 000 tokens at Tier 5, accommodating large‑scale workloads.
These limits ensure fair access while allowing higher‑spending customers to submit larger volumes of requests.
The combination of lower pricing, extensive token windows, and broad tool integration makes GPT‑5.6 Sol a cost‑effective option for enterprises building sophisticated AI applications.
Why This Matters: OpenAI’s reduced pricing for GPT‑5.6 Sol lowers cost barriers for developers seeking high‑capacity models.
This digest was compiled from:
Share this digest
People Also Read
- Introducing SCM: AI‑powered search across all photos and video frames on macOS
SCM brings offline, AI‑powered search for photos and video frames to macOS, offering privacy‑first media retrieval without cloud reliance.
- Introducing Kolibri: Germany’s sovereign AI model
Germany’s Aleph Alpha releases Kolibri, a sovereign German‑trained LLM designed for EU‑compliant, on‑premises use.
- GPT-6 Astra tackles World of Warcraft via agent-wow for the first time
GPT‑6 Astra completes WoW’s starting zone quests in 40 minutes using agent‑wow, showing LLMs can play complex games without game‑specific training.
Share your thoughts
Reactions, corrections, or insights — all welcome.
