Bilibili Launches AI Infinite Arena, Crowdsourcing Model Tests with GPT-6 Astra Leading Initial Rankings

September 20, 2026
Bilibili Launches AI Infinite Arena, Crowdsourcing Model Tests with GPT-6 Astra Leading Initial Rankings
  • Bilibili debuts the AI Infinite Arena, a large-model evaluation platform that crowdsources real-workflow style tests from UP masters across diverse fields.

  • The arena conducts real-device, real-world tests with hundreds of large language models evaluated through dozens of tests created by platform organizers (UPs).

  • Rankings feature mainstream models like DeepSeek, Kimi, ChatGPT, Claude, Gemini, Doubao, Qwen, Hy, and MiniMax, with real-time updates and ongoing open registration for new UP creators.

  • GPT-6 Astra wins first place in the inaugural ranking and earns the “most climbs to the top” title, ahead of GLM-5.3 in round one.

  • The initial round places GPT-6 Astra at the top, with three domestic large models in the top five.

  • The arena centers on practical, real-workflow tasks—coding, reasoning, collaboration, and knowledge—without a fixed set of evaluation dimensions; topics are proposed by creators to mirror real usage.

  • There are no topic or evaluation-dimension restrictions; real workflows and self-proposed topics showcase model performance across multiple competencies.

  • Top contenders include both domestic Chinese models and overseas giants, such as DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax, signaling broad participation.

  • Rankings update in real time and registrations remain open, inviting ongoing participation from the Bilibili community.

  • Public data show AI knowledge-content consumption on Bilibili rising, with over 190 million monthly AI-content viewers and a thriving ecosystem of releases, discussions, evaluations, uses, and collaborative development.

  • The platform aims to reveal practical performance differences among models within same-topic competitions, offering an intuitive take beyond traditional benchmarks.

Summary based on 2 sources


Get a daily email with more AI stories

More Stories