GPT 6 Astra Might Dethrone Opus 5.1
GPT 6 Astra could challenge Opus 5.1 with stronger reasoning, coding and agentic capabilities, potentially reshaping the race among the leading AI models.
The AI industry does not slow down, and this week delivered three stories in three days, each deserving its own article. A possible preview of OpenAI’s next major model. A controversy at Anthropic that has developers furious. And a massive open-source release from Tencent that most people outside China have not yet heard about.
Here are all three, clearly separated, with the rumour labelled as rumour and the facts labelled as facts.OpenAI’s Next Model Is Apparently Almost Here
On August 1, OpenAI did something unusual. In a research post about ten unsolved mathematical problems, the company mentioned almost in passing that the results were achieved by “an internal version of Astra, our next major model.” No release date. No specs. No pricing. Just a name and a classification: next major model.
That has sent the AI community into a sustained speculative frenzy, and leaks have been accumulating since.
According to leakers on X, including one who goes by the handle Leo and has a track record of OpenAI information, Astra has graduated from OpenAI’s internal testing phase, known internally as dogfooding. It is reportedly now being distributed to a small group of OpenAI partners under a codename: Ultima Alpha. The reported plan is to expand access if partner feedback is positive, with a wider public launch window potentially opening around early September.
OpenAI has officially confirmed that none of this is true: no release date, no model card, no API, no pricing. The company has not even confirmed whether Astra will be called GPT-6 or something else entirely. Earlier this year, a model codenamed Spud was widely expected to be GPT-6. It shipped in April as GPT-5.5. OpenAI has a consistent track record of releasing under version numbers that deflate pre-release speculation.
What makes the Astra leaks worth paying attention to despite the uncertainty is the qualitative output circulating online. Developers and communities who claim to have accessed early checkpoints have shared generated interfaces, 3D interactive environments, and SVG illustrations that look significantly more capable than anything currently publicly available. The code generation outputs in particular look like a meaningful step forward: complete interactive interfaces built from single prompts, with design decisions, spacing, and visual hierarchy that current models typically get wrong. If these outputs are real and representative, frontend and web development is the area where the jump would be most visible.
The model is also apparently being tested alongside an updated image generation model, with both potentially launching in a similar window.
Take all of this with appropriate scepticism. OpenAI has not said any of it. But the combination of the official August 1 naming, the leak volume, and the quality of the outputs circulating has made this the most credible pre-release Astra signal to date.
Anthropic Announced a Raise That Is Actually a Cut
This one is fully confirmed, and it is a case study in how not to communicate with your user base.
On August 29, Anthropic’s developer account posted the following: “Starting September 14, we’re permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans.”
That sounds like good news. It is not, and developers caught it immediately.
Here is the math. On May 13, Anthropic introduced a temporary 50 per cent increase in Claude Code’s weekly usage limits. The company extended that temporary boost four times over the summer, meaning users experienced it for months. It stopped feeling temporary. It started feeling like the new normal.
So when Anthropic announced a permanent 25 per cent raise, the relevant comparison for most users was not the baseline from before May. It was the 150per centt of baseline they have been running on all summer. And 125 per cent of the old baseline is roughly 17 per cent less than 150 per cent of the old baseline.
Anthropic eventually acknowledged this in a follow-up post: “Compared to today, this works out to a 17% reduction in weekly limits on Claude Code.” The follow-up clarification came after the original post attracted significant backlash. Anthropic employees admitted internally that the messaging should have led with the reduction rather than burying it. Several developers announced they were cancelling their subscriptions.
The practical situation: if you are a heavy Claude Code user who has been running at full capacity through the summer, your available usage drops by roughly 17 percent on September 14. If you rarely hit the weekly cap, the change may not affect you at all. Anthropic says interface updates are coming that will give users better visibility into their remaining weekly quota, though no timeline has been provided.
The timing is particularly awkward. Anthropic is rolling out this reduction at the same moment OpenAI is apparently preparing to launch what could be a significant new model. Developers who have built their workflows around Claude Code are now being asked to do more with less, precisely when the competitive pressure is highest.
Tencent Just Released a Monster, and It Is Free
On August 28, Tencent’s Hunyuan team released Hy4 preview, a 770-billion-parameter open-source model licensed under Apache 2.0. Anyone can download it, run it, modify it, and build commercial products on top of it.
The 770 billion number deserves context, because it is somewhat misleading on its own. Hy4 is a Mixture-of-Experts model, which means its total parameter count is spread across a large pool of specialised sub-networks. Still, only about 49 billion of those parameters are actually activated for any given request. The rest sit dormant. This is the same design principle that makes DeepSeek V4 Pro efficient despite its enormous total size: you get the depth of a massive model with the computational cost of a much smaller one.
What Tencent is positioning as the model’s defining purpose is productivity in real workloads. Not benchmark scores, though those are strong. The internal testing used a blind evaluation with 163 engineers across 203 engineering tasks, where Hy4 outperformed GLM 5.3 and Kimi K3, both considered among the strongest open models in their respective categories. The evaluations are Tencent’s own and have not yet been independently replicated, which is worth noting.
The context window exceeds one million tokens. The model supports coding, long-form document analysis, and scientific research. It has been integrated into Tencent’s existing products including WorkBuddy, CodeBuddy, and Yuanbao, and is available globally through Tencent Cloud and OpenRouter.
There is also something in the Hy4 technical documentation that is not mentioned enough: this is the first Hunyuan model to participate in optimising its own training pipeline. Hy4 was used during its own development to help automate improvements to training methods, and it separately optimized its own inference infrastructure, resulting in a measured 31.8 percent throughput increase. A model that helped design the process that produced it is not a marketing claim. It is a specific, documented mechanism, and it is a meaningful step in a direction the entire field is moving.
Two weeks of free access to Hy4 are available through WorkBuddy and CodeBuddy from launch. API access is available through Tencent Cloud at $0.834 per million input tokens.
The honest caveats: the model can be slow for complex queries and sometimes oververifies its own answers before responding, as Tencent acknowledges in the release notes. The benchmark results are entirely company-reported, and independent evaluation is pending. And as with any model trained on data from a Chinese company, enterprise teams should evaluate the data sourcing and governance documentation before deploying it in sensitive contexts.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0