Bilibili Launches AI Infinite Arena, Crowdsourcing Model Tests with GPT-6 Astra Leading Initial Rankings
September 20, 2026
Bilibili debuts the AI Infinite Arena, a large-model evaluation platform that crowdsources real-workflow style tests from UP masters across diverse fields.
The arena conducts real-device, real-world tests with hundreds of large language models evaluated through dozens of tests created by platform organizers (UPs).
Rankings feature mainstream models like DeepSeek, Kimi, ChatGPT, Claude, Gemini, Doubao, Qwen, Hy, and MiniMax, with real-time updates and ongoing open registration for new UP creators.
GPT-6 Astra wins first place in the inaugural ranking and earns the “most climbs to the top” title, ahead of GLM-5.3 in round one.
The initial round places GPT-6 Astra at the top, with three domestic large models in the top five.
The arena centers on practical, real-workflow tasks—coding, reasoning, collaboration, and knowledge—without a fixed set of evaluation dimensions; topics are proposed by creators to mirror real usage.
There are no topic or evaluation-dimension restrictions; real workflows and self-proposed topics showcase model performance across multiple competencies.
Top contenders include both domestic Chinese models and overseas giants, such as DeepSeek, Kimi, ChatGPT, Claude, Gemini, DouBao, Qianwen, Hy, and MiniMax, signaling broad participation.
Rankings update in real time and registrations remain open, inviting ongoing participation from the Bilibili community.
Public data show AI knowledge-content consumption on Bilibili rising, with over 190 million monthly AI-content viewers and a thriving ecosystem of releases, discussions, evaluations, uses, and collaborative development.
The platform aims to reveal practical performance differences among models within same-topic competitions, offering an intuitive take beyond traditional benchmarks.
Summary based on 2 sources

