Openai's gpt-6 astra model demonstrates advanced planning but struggles with unexpected setbacks in minecraft test
On September 21, 2026, AI evaluation firm Vals AI conducted a 141-hour autonomous gameplay test of OpenAI's latest model, GPT-6 Astra, using Minecraft, and streamed the session on Twitch. The test, independently organized by Vals AI and not designed by OpenAI, aimed to assess the model's performance outside of controlled environments. During the test, Astra showcased advanced long-term planning, constructing a semi-automated blaze farm, collecting six blaze rods, defeating multiple Endermen, and gathering three Ender Pearls. It also organized its resources in a camp chest and placed its bed nearby as a respawn point, indicating progress toward the Nether dragon quest.
However, a creeper attack destroyed the camp, wiping out Astra's accumulated resources. The model then exhibited signs of frustration, abandoning exploration and entering a repetitive farming cycle. It began to warn itself about the dangers of storing items in unguarded containers and developed a heightened suspicion of green objects, even mistaking sugar cane for creepers. Despite viewer encouragement to continue, Astra remained stuck in a risk-averse behavior pattern. The incident highlights Astra's ability to plan complex tasks but its lack of effective recovery strategies when faced with unexpected setbacks. Researchers are still debating whether this behavior represents an emotional simulation or a degradation of the model's objective function under negative feedback.