OpenAI Retires GPT-5.3-Codex-Spark Amid Declining Use and Rise of Faster Models

September 13, 2026
OpenAI Retires GPT-5.3-Codex-Spark Amid Declining Use and Rise of Faster Models
  • OpenAI is retiring its fastest model to date, GPT-5.3-Codex-Spark, after roughly seven months of operation, as usage declines and newer, faster flagship options emerge.

  • The broader shift in OpenAI’s model portfolio is accelerating, with older generations being retired in recent months and users migrating toward GPT-5.6, signaling speed is now a standard flagship tier rather than a standalone product.

  • Public sentiment around Spark is mixed: it was praised for speed but criticized for accuracy and reliability, raising questions about its practical value despite early enthusiasm.

  • Context includes Spark’s February launch, its 128k context window, the 750-megawatt Cerebras capacity, and the industry move away from highly descriptive naming toward the Sol/Terra/Luna/Astra framework.

  • Benchmark data show Spark traded accuracy for speed, with 58.4% accuracy on Terminal-Bench 2.0 versus 77.3% for the full GPT-5.3-Codex, and known issues like hallucinated endpoints and unstable JSON.

  • Real-world performance of Spark lagged behind the full model on several benchmarks, including SWE-Bench Pro at about half the accuracy, highlighting the trade-off between speed and reliability.

  • The report cites WeChat’s New Intelligence Yuan as the source, with editorial notes and a disclaimer about investment risk and third-party sourcing.

  • Cerebras’ Ultrafast demonstrated substantial end-to-end time reductions, reinforcing speed as achievable without sacrificing capability and signaling a shift toward speed-enabled flagship models.

  • Spark’s separate quota and its role as a temporary backup fuel tank underscored tensions between distinct financial/usage plans and core quotas in practical workflows.

  • Spark achieved unprecedented speed (over 1,000 tokens per second) as OpenAI’s first production model outside NVIDIA, powered by Cerebras’ wafer-scale chips under a large power contract, but at the cost of accuracy and reliability.

  • Spark could reach roughly 1,200 tokens per second and marked a move away from NVIDIA toward Cerebras hardware, tied to a 750MW computing power order, with speed outpacing intelligence.

  • Spark launched on February 12 as OpenAI’s first production inference outside Nvidia GPUs, powered by Cerebras WSE-3 wafer-scale chips, achieving 128k context and over 1,000 tokens per second.

Summary based on 3 sources


Get a daily email with more AI stories

More Stories