Announcing GPT-6.7 Galaxy Ultra xhigh Reasoning With 1M Context

Wait, you do more than write emails? And you outperform previous models by 3.9%? Well, here are the three ways to make meaningful progress in AI development.

Published July 10th, 2026

Announcing GPT-6.7 Galaxy Ultra xhigh Reasoning With 1M Context

Our new model outperforms our previous models on almost every benchmark. It has extended (i.e., safer and better) safeguards that prevent malicious use while preserving industry-leading performance across benchmarks and agentic tool use. Please consult our chart:

Test
GPT 6.7 Galaxy Ultra xhigh Reasoning (1M Context) GPT 6.5 Sol Pro xhigh (1M Context)
SWE-Bench Pro Biggest (for now) Big
Agentic Engineering Biggest (for now) Big
Vibe Evaluation Biggest (barely) Big
Apocalypse Incoming v3 (ai-v3) Biggest Big

I’m obviously being satirical, but unfortunately, most recent model releases can be boiled down to this. I mean: Fable 5, Opus 4.8, GPT 5.6, Sonnet 5. They’re all the same. Here’s Sonnet 5’s, from the Anthropic announcement:
Announcing GPT-6.7 Galaxy Ultra xhigh Reasoning With 1M Context Sonnet 5 Announcement.png

And GPT 5.6, from OpenAI:
Announcing GPT-6.7 Galaxy Ultra xhigh Reasoning With 1M Context GPT 5.6 Sol.png

They’re all the same. It’s getting old!

Having worked extensively with AI for about six months now, and with a background in computer science, I feel like every model release is… More of the same. Every release comes with a wave of developers espousing the cool things that this new model can do—“I left such-and-such new model on overnight, and it did such-and-such cool thing”—which creates clout and the impression that this new model is noticeably and significantly better than its predecessor.

Here’s a thought.

If I hid the reasoning traces and all human-readable output, and only showed you a finished product, would you be able to tell whether something was made by Fable 5, Opus 4.8, Opus 4.7, or Sonnet 5? Could you tell the difference between GPT 5.3, 5.4, and 5.5?

My guess is no. There’s a reason benchmarks have become so prevalent, and it’s because there’s no longer any way for an average developer to see a noticeable difference in ability between LLMs. In the days before agents, agent teams, whatever, every new model or feature was a noticeable and significant deviation from the norm. When Github Copilot first started offering AI-powered autocomplete, or was first able to generate a couple of lines of code, it was huge. A few years later, ChatGPT could generate a whole script! And then Copilot and Claude Code were able to write code and control your computer while doing it, and now Claude can interact with and work from your Slack workspace. Each of these steps is significant in that it changes the form factor and expands the capabilities of AI models—but when developments in AI models aren’t paired with developments in how AI is used (or capable of working), we get these model announcements:

OpenAI releases new model that weightsmaxxes and sizemogs Anthropic’s no-longer-state-of-the-art model.

Okay.

But if Anthropic could probably get away with serving Opus 4.8 to a software engineer while claiming that it’s Fable 5, what are we doing? What does an extra 3.9%[1] on some benchmark actually translate to in user experience changes? Thinking about this and talking to people around the world (in places where self-driving cars are still science fiction) has given me a couple of thoughts on the subject of AI development. I’ll summarize them like this:

  1. Our models are good enough now for the vast majority of applications.
  2. The vast majority of the world doesn’t understand that “AI will take your job” is no longer science fiction. (For very avoidable reasons.)
  3. There are three main directions for truly novel and meaningful AI development.

Our models are good enough

In the days of GPT 3.5-Instant, when we still had models with less than (fewer than?) a billion parameters—current SoTA models are in the trillions of parameters, now—we had problems that prevented industrial use. Early models couldn’t do arithmetic very well, they hallucinated information, they had frequent breakdowns, etc. It was a whole thing. Now, though, our frontier models are solving unsolved problems in mathematics, splicing human DNA to create custom cancer treatments, and writing a huge part of the world’s software. They’ve even been used to find vulnerabilities in Cloudflare, the primary provider of internet infrastructure!

This is to say that, if you want to have voicemail that answers in your voice and can schedule meetings for you, there’s really nothing stopping you from doing it. If you want to have a website chatbot that doesn’t suck, you can have it. You could even have an AI girlfriend if you wanted one.[2]

We are no longer in the days of GPT 3.5-Instant. Our modern models are good enough; whether they perform 3.9% better than their predecessor or not makes little, if any, difference. (In practical applications.)

Yes, AI will take your job

As more and more companies, people, and—importantly—investors realize that our models meet this threshold of being usable, money flows into adopting or developing novel applications of AI. If you don’t believe me, consider, for example, Allbirds: the shoe company whose stock price increased sevenfold after it rebranded to “Smartbird” and became an AI company. Money flows to AI.

And this money flow will cause innovation, as it always does, and consequently a loss in jobs. Software engineers said their jobs were safe two years ago; now, tech companies like Meta and Fortune 500 company Block lay off thousands of their technical staff. Block’s CEO, after firing 40% of Block’s workforce, explained that the primary reason was AI—and that he believes other companies will follow suit. Shortly after, Coinbase followed suit. Software engineering might not be so safe now, huh?

Increasingly, corporations are feeling the pressure to incorporate automation. This pressure spans accounting, law, software engineering, computer chip design (!), and it’s a matter of time before other industries feel the same pressures—and pressure to automate inevitably leads to automation and, consequently, layoffs.[3]

There is a commonly-held belief in some circles that there will always be a need for humans in… whatever. People have filled the blank with all sorts of things (teaching, research, and customer service, as examples). As far as I can tell, this belief is simply false.

So long as our machines have limitations, there is going to be a guy who thinks, “What if we could automate this?”—and to be honest, thinking like this is what brought us here in the first place. It’s why we have self-driving cars. It’s what we have to thank for the Industrial Revolution. But it also means that it’s a matter of time until, one year from now or twenty, your job will be replaced by AI (or some future derivative of it).

Though I imagine AI will cause a record-breaking loss in jobs, I’m not too worried about the future. I think the world will have time to adapt,[4] and overall the transition will be similar to the automation revolutions (such as the arrival of the Internet) we’ve experienced in the past. People will lose jobs. It’ll be a shame. But that’s the inescapable toll of progress.

Perhaps the more pressing problem, though, is that at the moment the vast majority of the world doesn’t believe that AI—which is only good for writing emails, right?—is all that. I’ve spoken to people who haven’t even used ChatGPT—and in the circles that read my writing, this is quite unfathomable. I’ve found that there are three camps of people when it comes to AI:

  • The people who have not used AI, or have used it at only a surface level (e.g. via ChatGPT Free, AI Overview in Google, or Apple Intelligence); I will call these the Outsiders.
  • The people who pay for cheap (the $1–$30 dollar plans) AI, and use it considerably heavily. I’ve found that somehow this group disproportionately contains students. The Marginals.
  • The people who pay for the most expensive (more than $99 a month) plans and use those to the maximum. The Insiders.

Here’s a fun fact of the day: of OpenAI’s weekly active users, 94.4% use the free tier. The vast, vast majority of the internet-capable human race is an Outsider.

And this is a problem. At the Outsiders’ level, “AI” is Google AI Overview, news about Grok generating some new thing involving minors, and pieces from technically illiterate mainstream media. For example, this piece from WIRED[5] about how humans are safe in the realm of fact checking uses worthless models and actively misleads people about what’s possible with AI. As the article points out, Google’s AI Overview and ChatGPT Free genuinely suck. In fact, while I was writing this article, Overview failed me, too:

IMG_7262.jpeg

And if you were to hypothetically Google, “How many stairs on the Inca Trail?” it would tell you, confidently, how many steps (70,000) there are, not how many stairs (9,862, according to my calculations) there are. But the difference between ChatGPT 4o (from early May of last year), which the article uses to back up its claim, and GPT 5.5 Pro (or the difference between Gemini and Claude Fable 5), is large.[6] So large, in fact, that when I first thought about how to describe it, the first words that came to mind were “larger than your mom.”

I’m sure she’s a nice, polite, attractive woman; I mean no insult. But at the frontier of artificial intelligence, models are autonomously solving Erdos problem after Erdos problem. They are literally splicing human DNA. They can generate (nearly) any website you can think of in the span of less than an hour, and deploy it to the world in, maybe, 30 minutes. Modern AI can drive cars across the country. And so the Insiders, who talk about “agent swarms” and think “autoresearch” is last year’s news, are scared about the future. But these people make up less than 0.1% of the global population, and their articles don’t show up in WIRED or the New York Times.

There is a reason these companies are valued in the trillions, and it’s only partly due to the fact that with 80 minutes and a hundred dollars, you can produce a publishable paper in graduate-level mathematics.

But at the Outsiders’ level, AI is the thing that the big bad data centers are built for. “Each ChatGPT query guzzles a bottle of water,” says the mainstream media[7] to Outsiders:
Announcing GPT-6.7 Galaxy Ultra xhigh Reasoning With 1M Context Washington Post Bottle of Water.png
(This graphic is from The Washington Post. See Footnote 5.)

And now we have social media, mainstream media, politicians, influencers, and normal people all campaigning to ban data centers in a way not dissimilar to the pushes against nuclear energy and supersonic flight. A LaGuardia-to-SFO flight could take two hours; instead, it takes five and a half. Because of poor media and an inadequate, outdated understanding of modern AI’s capabilities, the vast majority of people lean anti-AI. This means they don’t see what’s coming—and being surprised by the future is worse than having time to adapt.

Three directions for meaningful AI development

I say that in a very grandiose way, but the reality is that even in AI circles most people don’t know what’s coming. Those who claim they know have either read Asimov, Gibson, or some threads on Twitter. AI models can solve increasingly high percentages of benchmarks that most humans would struggle on already;[8] we’re just making them smarter and smarter with no clear understanding of what we’re going to do once we’ve maxed out the benchmarks. Do we… make new, harder benchmarks?

It’s this question—what does a slightly higher number on a benchmark really mean, if our models are already good enough, and the main issue is that the public doesn’t want to build data centers or believe in AI—that I gave a little bit of thought to.

I only gave it a little bit of thought because the answer is pretty obviously “A higher number basically means nothing.” Anthropic can be serving us Opus 4.7 instead of Fable 5, and nobody would be the wiser. This tells me that chasing benchmarks will not translate to noticeable, significant, or meaningful changes in the field of AI. Instead, as we’ve seen in the past, real newsworthy changes come from one of three directions in AI development:

  1. New architectures,
  2. New capabilities, or
  3. Speedups and cost reductions.

It should be obvious why new architectures constitute newsworthy AI research: imagine we invented a successor to reinforcement learning, or world models finally showed signs of being comparable in capability to language models. Every AI enthusiast would know! There is a reason you can say “Attention is…” and most people familiar with AI research will finish the phrase with “All you need”: everyone knows the Attention Is All You Need paper that finally got modern large language models off the ground. This is all a verbose way of saying that if we tried something other than large language models, and it performed better (or had a different set of limitations) than current frontier models, it would be big news.

New capabilities also make for interesting and meaningful progress. The fact that many Slack users can now ping @Claude and have Claude start working on something completely autonomously, or can be pleasantly surprised by Claude reacting to their Slack messages with fun emojis, or interact with Claude as though it were a work-from-home coworker… This is a very, very strange time we live in. But every major step in AI’s form factor and capabilities has been newsworthy and created quality of life improvements across industries. As a short list of noteworthy AI developments that have fallen into this “new capabilities” category (and as a result been widely reported on in tech circles):

  • AI that can generate 3D models for game developers,
  • Image and video generation by Midjourney, Dall-E, ChatGPT, and Gemini,
  • Web design by AI (such as v0, Base44, Claude Code/Design, and Codex),
  • Slack integrations,
  • Code autocomplete,
  • Financial audits (like OpenAI finance tools),
    And, most recently, AI-in-messages capabilities like Poke, Meta AI, and OpenClaw. Future additions to this already long list are surely right around the corner! I imagine it won’t be long until Claude/ChatGPT can control your house or do something equally crazy.

Finally, perhaps the least interesting of the directions: speedups and cost reductions. While boring, these make a real difference in the usability and adoption of AI. The Fable 5 session running on my machine right now, which has been running for about an hour, has cost $267.43 in API usage. That is a staggeringly high amount of money. Open-source models, though, simply aren’t feasible to run on consumer hardware like my three-year-old laptop yet. This leaves a huge amount of potential improvement on the table: wouldn’t it be nice if your phone could run the world’s most intelligent technology to give you near-instant answers to any question you have? Faster models would make things like Claude for Fusion 360 actually usable! Honestly, I’m out of examples. I think there’s just so much potential in continued research into speed and cost optimizations.

And perhaps… Perhaps there will always be a need for humans who write article-length thought pieces about AI. With that, I’d like to announce GPT 6.8 Ultra Mega Extra High with 2M Context. It’s even bigger, better, and safer than GPT 6.7.


  1. Real number! It’s the difference between GPT 5.6 Sol Ultra and Claude Mythos 5.↩︎

  2. Read my short story, All You Need↩︎

  3. Here, I should specify that AI is necessarily “automation.” We replace humans’ jobs (or parts of their jobs) with machines.↩︎

  4. I think this will be the case even in a Recursive Self-Improvement (RSI) scenario. A footnote is not the right place to explain why, so. An article for another time.↩︎

  5. Article can be found permanently on archive.is.↩︎

  6. To be precise: GPT 4o mini (which is the free ChatGPT model) scores 4.0% on Humanity’s Last Exam. Fable 5 scores 53.3%.↩︎

  7. The Washington Post’s 2024 reporting: https://www.washingtonpost.com/technology/2024/09/18/energy-ai-use-electricity-water-data-centers/. (Found here on archive.is). Maybe democracy dies in poorly-researched infographics too?↩︎

  8. Is this a three or a five? https://mint.westdri.ca/ai/pt/img/mnist_digit.png. AI models can score up to 99.8% on the MNIST dataset. And this is one of the easiest benchmarks around!↩︎

Liked this article?

More articles