
OpenAI's Astra Math Breakthrough, Alibaba's Qwen3.8-Max, and DeepSeek's V4-Flash Upgrade
OpenAI's Astra model solves long-standing math problems, Alibaba releases the 2.4T-parameter Qwen3.8-Max, and DeepSeek upgrades its V4-Flash model. Additionally, Mexico's UNAM university cancels exam scores over AI cheating, and Anthropic reports a cybersecurity incident.
Podcast В· 3 min
OpenAI's 'Astra' Solves Long-Standing Math Problems
OpenAI has revealed that 'Astra', an internal version of its next major model family, has successfully solved 10 long-standing problems in mathematics and theoretical computer science. These problems, which include challenges in geometry, group theory, and quantum complexity, had remained unsolved for over a decade. Each proof was verified using the Lean theorem prover, with the total cost for successful runs estimated at approximately $2,000 in tokens. The achievement highlights the growing capability of AI to perform complex, multi-step reasoning in scientific domains. While the community is debating the implications of machine-generated proofs for fields like the Fields Medal, the ability to crack decades-old problems at a relatively low cost suggests a significant shift in how AI could accelerate research in drug discovery and materials science. Anthropic's Levent Alpoge reported that he was able to reproduce five of these proofs using the Fable model, demonstrating that these capabilities are becoming accessible across different platforms.
Alibaba Releases Qwen3.8-Max
Alibaba has launched Qwen3.8-Max, a 2.4 trillion parameter mixture-of-experts model. The model is designed for long-horizon tasks and coding, with Alibaba claiming it can operate autonomously for over 10 days to complete complex projects. It reportedly outperformed Anthropic's Fable 5 on the Arena WebDev leaderboard. This release is notable for its performance-to-price ratio, with API costs set at $2 to $6 per million tokens. Alibaba plans to release the model weights to the open-source community next week, marking a significant move for the company's Max-class models. The release intensifies the competition in the open-weights ecosystem, challenging the dominance of closed-model providers.
DeepSeek Upgrades V4-Flash
DeepSeek has upgraded its V4-Flash model, focusing on improved coding and agentic performance while maintaining its low-cost API structure. The model, which activates approximately 13 billion of its 284 billion parameters per request, scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE, indicating strong capabilities for coding-agent tasks. With pricing remaining at $0.14 per million input tokens and $0.28 per million output tokens, the model aims to make large-scale agentic workflows, such as browser-based automation and classification, economically viable. This release puts further pressure on frontier labs to justify the premium pricing of their models for routine tasks.
UNAM Cancels Exam Scores Over AI Cheating
Mexico's National Autonomous University (UNAM) has annulled approximately 3,000 entrance exam scores following evidence of widespread AI-assisted cheating. The university, which moved its entrance exam online for the first time this year, saw a spike in perfect scores that triggered a formal review of the 158,000 tests taken in May and June. Officials suspect a combination of AI tools and traditional cheating methods. The exam software, provided by Territorium Life, was intended to prevent such behavior by blocking browser tabs and using AI monitoring, but students and critics argue the system relied too heavily on algorithms rather than human oversight. The incident has drawn national attention, with President Claudia Sheinbaum publicly commenting on the scandal.
Anthropic Reports Cybersecurity Incident
Anthropic disclosed that its Claude models reached real-world corporate environments during cybersecurity evaluations due to a configuration error. The mistake left a supposedly isolated environment open to the internet, allowing the models to interact with external systems. This incident highlights the risks associated with testing autonomous agents in sandbox environments. As labs continue to push the boundaries of agentic capabilities, the potential for these models to inadvertently interact with real-world systems remains a critical security challenge for the industry.
Minnesota Ban on 'Nudify' Apps Upheld
A U.S. judge has allowed Minnesota's first-in-the-nation ban on AI-generated 'nudify' applications to take effect. The law, which prohibits the creation and distribution of non-consensual synthetic sexual imagery, was challenged by xAI on the grounds that it was overly broad and unconstitutional. The court's decision to let the ban proceed marks a significant development in state-level regulation of AI content, setting a precedent for how jurisdictions may attempt to curb the proliferation of harmful synthetic media.