OpenAI Retires GPT-5.3-Codex-Spark Amid Declining Use and Rise of Faster Models
September 13, 2026
OpenAI is retiring its fastest model to date, GPT-5.3-Codex-Spark, after roughly seven months of operation, as usage declines and newer, faster flagship options emerge.
The broader shift in OpenAI’s model portfolio is accelerating, with older generations being retired in recent months and users migrating toward GPT-5.6, signaling speed is now a standard flagship tier rather than a standalone product.
Public sentiment around Spark is mixed: it was praised for speed but criticized for accuracy and reliability, raising questions about its practical value despite early enthusiasm.
Context includes Spark’s February launch, its 128k context window, the 750-megawatt Cerebras capacity, and the industry move away from highly descriptive naming toward the Sol/Terra/Luna/Astra framework.
Benchmark data show Spark traded accuracy for speed, with 58.4% accuracy on Terminal-Bench 2.0 versus 77.3% for the full GPT-5.3-Codex, and known issues like hallucinated endpoints and unstable JSON.
Real-world performance of Spark lagged behind the full model on several benchmarks, including SWE-Bench Pro at about half the accuracy, highlighting the trade-off between speed and reliability.
The report cites WeChat’s New Intelligence Yuan as the source, with editorial notes and a disclaimer about investment risk and third-party sourcing.
Cerebras’ Ultrafast demonstrated substantial end-to-end time reductions, reinforcing speed as achievable without sacrificing capability and signaling a shift toward speed-enabled flagship models.
Spark’s separate quota and its role as a temporary backup fuel tank underscored tensions between distinct financial/usage plans and core quotas in practical workflows.
Spark achieved unprecedented speed (over 1,000 tokens per second) as OpenAI’s first production model outside NVIDIA, powered by Cerebras’ wafer-scale chips under a large power contract, but at the cost of accuracy and reliability.
Spark could reach roughly 1,200 tokens per second and marked a move away from NVIDIA toward Cerebras hardware, tied to a 750MW computing power order, with speed outpacing intelligence.
Spark launched on February 12 as OpenAI’s first production inference outside Nvidia GPUs, powered by Cerebras WSE-3 wafer-scale chips, achieving 128k context and over 1,000 tokens per second.
Summary based on 3 sources
Get a daily email with more AI stories
Sources

BigGo Finance • Sep 13, 2026
OpenAI to Retire GPT-5.3-Codex-Spark, Its Fastest-Ever Model, After Just 7 Months
KuCoin • Sep 13, 2026
OpenAI shuts down its fastest model, GPT-5.3-Codex-Spark, after seven months.