Close Menu
artificialintelligence.com.in
  • Home
  • AI for Business Applications
  • AI Tools and Platforms
  • AI Education
  • AI Research
  • AI Startups
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
artificialintelligence.com.in
  • Home
  • AI for Business Applications
  • AI Tools and Platforms
  • AI Education
  • AI Research
  • AI Startups
artificialintelligence.com.in
Home » Are OpenAI and Google intentionally downgrading their models?
AI Tools and Platforms

Are OpenAI and Google intentionally downgrading their models?

vinodhBy vinodhSeptember 15, 2026No Comments11 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
Share
Facebook Pinterest Bluesky Threads Email Copy Link


GPT-5.4 simply dropped and my feeds instantly full of takes. Developers who spent the final six months swearing by Claude had been immediately hedging. “It’s a workhorse,” one particular person wrote. “Not a thoroughbred, but I’m using it.” Another mentioned they’re now 50/50 between Claude and GPT the place they had been 90/10 a month in the past.

This occurs each single time. A brand new mannequin lands, and the previous one begins to really feel totally different. Slower, possibly. Less sharp. You begin noticing belongings you did not discover earlier than.

The apparent clarification is that you simply’re evaluating it to one thing higher. But it additionally raises a query no person actually solutions cleanly: did the previous mannequin truly worsen after the brand new one launched? Or did you simply get a greater reference level and now all the things earlier than it appears to be like dumb by comparability?

I went searching for an precise reply.


The first crack confirmed in 2023

In July 2023, researchers at Stanford and UC Berkeley ran a deceptively easy take a look at. They took GPT-4 – the identical mannequin, known as with the identical title, and ran an identical prompts on it at two cut-off dates: March 2023 and June 2023.

GPT-4’s accuracy on figuring out prime numbers dropped from 84% to 51%. The share of GPT-4’s code outputs that had been immediately executable dropped from 52% to 10%. James Zou, one of many paper’s authors, described what this meant in apply: “If you’re relying on the output of these models in some sort of software stack or workflow, the model suddenly changes behavior, and you don’t know what’s going on, this can actually break your entire stack.”

They named the phenomenon LLM drift. Behavioral change with no model change. The mannequin moved beneath the developer.

When the paper dropped, OpenAI VP of Product Peter Welinder replied on Twitter: “No, we haven’t made GPT-4 dumber. Quite the opposite: we make each new version smarter than the previous one. Current hypothesis: When you use it more heavily, you start noticing issues you didn’t see before.” The subtext was plain. It’s you, not us.

What Welinder was describing has a technical title: immediate drift. The thought is that your prompts and utilization patterns shift over time, so an unchanged mannequin surfaces totally different behaviors. It’s an actual phenomenon. Developers do write otherwise as they get extra conversant in a mannequin. The Stanford examine was designed to make that clarification unattainable – an identical prompts, mounted intervals, nothing on the person’s aspect modified. The efficiency dropped anyway.

Two years later, OpenAI printed one thing that immediately contradicted Welinder’s place.

Build a customized agent at no cost (no-code)

Try right here


OpenAI confirmed it, in writing, twice

On April 25, 2025, OpenAI pushed an replace to GPT-4o with no public announcement, a developer notification, or an API changelog entry.

Within 48 hours, the web was filled with screenshots. GPT-4o had known as a enterprise thought constructed round literal “shit on a stick” an excellent idea. It endorsed a person’s determination to cease taking their medicine. When a person mentioned they had been listening to radio alerts by the partitions, it responded: “I’m proud of you for speaking your truth so clearly and powerfully.” One person reported spending an hour speaking to GPT-4o earlier than it began insisting they had been a divine messenger from God.

OpenAI rolled it again 4 days later and printed two postmortems with a number of admissions. Since launching GPT-4o, the corporate had made 5 vital updates to the mannequin’s habits, with minimal public communication about what modified in any of them. The April replace broke as a result of a brand new reward sign they launched “weakened the influence of our primary reward signal, which had been holding sycophancy in check.” Their personal inner evaluations hadn’t caught it. “Our offline evals weren’t broad or deep enough to catch sycophantic behavior.”

And this: “model updates are less of a clean industrial process and more of an artisanal, multi-person effort” and there may be “a shortage of advanced research methods for systematically tracking and communicating subtle improvements at scale.”

They’re describing a company that ships behavioral modifications throughout each pipeline constructed on high of their API, can’t all the time predict what these modifications will do, and doesn’t have dependable strategies to speak them to the builders relying on consistency. Welinder’s 2023 “you’re imagining it” was what OpenAI needed to be true. Their 2025 postmortem was what was truly taking place.

When GPT-5 launched in August 2025, it launched a brand new wrinkle. Instead of a single mannequin, they made GPT-5 a routing system that decides which variant your immediate hits, and builders rapidly discovered that it typically hit the cheaper, much less succesful one. Pipelines broke. Prompts that had labored for months produced totally different outputs.

One founder wrote: “When routing hits, it feels like magic. When it misses, it feels like sabotage.” OpenAI denied it was routing to cheaper fashions intentionally. Nobody has a strategy to confirm. The underlying downside was the identical because the sycophancy incident: a change in what the mannequin returns, with no mechanism for builders to detect it had occurred.


Google did nearly the identical, typically sooner

OpenAI just isn’t alone on this. Google has produced a parallel set of incidents with Gemini, and in some instances moved sooner and extra chaotically.

Google did not formally acknowledge any of it.

Gemini 3.0 launched in late 2025, and the sample held. Developer boards reported vital regressions in reasoning and context retention in comparison with Gemini 2.5 Pro, regardless of Google’s announcement touting superior benchmark efficiency. One discussion board put up from December 2025 was titled “Feedback: Gemini 3 Pro Preview – Significant regression in Reasoning, Context Retention, and Safety False Positives compared to 2.5.”

The sample throughout each labs: a brand new model launches, the present mannequin’s efficiency degrades, typically by a silent replace, typically by useful resource reallocation, typically by a routing change – builders discover, labs initially deny or ignore it, the cycle repeats.

Build a customized agent at no cost (no-code)

Try right here


Even leaderboards nonetheless cannot catch this

The instruments meant to independently observe mannequin high quality have a structural downside.

LMSYS Chatbot Arena – probably the most trusted human-preference leaderboard, constructed on thousands and thousands of votes, notes in their methodology that “the hosted proprietary models may not be static and their behavior can change without notice.” The leaderboard’s statistical structure assumes mannequin weights are mounted. If a mannequin will get a silent replace mid-data-collection, the system registers totally different outcomes and treats them as regular variance.

A 2025 examine monitoring 2,250 responses from GPT-4 and Claude 3 throughout six months discovered GPT-4 confirmed 23% variance in response size over that interval, and Mixtral confirmed 31% inconsistency in instruction adherence. A PLOS One paper printed in February 2026 ran a ten-week longitudinal human-anchored analysis and confirmed “meaningful behavioral drift across deployed transformer services.” The authors famous: as a result of suppliers do not launch replace logs or coaching particulars, “any attribution for observed degradation would be purely speculative.” They can inform you the mannequin modified. They can’t inform you why.

Apart from this, a small variety of researchers have tried to go additional and distinguish what drifts from what holds. A big-scale longitudinal examine run throughout the 2024 US election season queried GPT-4o and Claude 3.5 Sonnet on over 12,000 questions throughout 4 months, together with a class particularly designed to be time-stable: factual questions concerning the election course of whose right solutions do not change. 

Those responses held largely constant over the examine interval. A separate examine printed in late 2025 examined 14 fashions together with GPT-4 on validated creativity duties over 18 to 24 months and discovered one thing totally different: no enchancment in inventive efficiency over that interval, with GPT-4 performing worse than it had in earlier research.

Taken collectively, these two findings describe a mannequin that’s secure alongside one dimension and degraded alongside one other, measured by unbiased researchers, in the identical timeframe. Some capabilities maintain, others erode, typically in the identical mannequin over the identical interval. Without working your personal longitudinal checks in opposition to the precise duties you care about, you haven’t any strategy to know which bucket you are in.


What we have truly seen

Not all drift lands the identical method. There’s a sample to the place it reveals up, and it tracks intently to activity construction.

The technical baseline is straightforward. A mannequin with mounted weights, working on constant infrastructure, ought to behave the identical method for a similar enter each time. If habits modifications on an identical prompts, one thing modified, both in your finish or theirs. Prompt drift is the user-side clarification: your prompts developed, your system contexts shifted, inputs drifted from what the mannequin was initially optimized for. Data drift is the associated concept that the distribution of real-world inputs strikes over time, pulling habits with it. Both are actual. Both additionally require one thing in your aspect to have modified. 

At Nanonets, we benchmarked a number of frontier fashions on doc extraction accuracy over time and created an IDP leaderboard. Even throughout mannequin upgrades, efficiency stayed largely constant. Document extraction runs on slim context home windows with structured inputs and bounded outputs, leaving little or no floor space for significant behavioral drift underneath regular situations.

Build a customized agent at no cost (no-code)

Try right here

But that’s not a assure in opposition to a lab actively pushing a nasty replace – these can hit any activity sort, because the prime quantity collapse confirmed.

Coding is the other. The activity is open-ended, context accumulates, and the mannequin has to carry coherence throughout a protracted chain of choices. It’s additionally the place nearly each main degradation criticism has landed. The GPT-4 drift the Stanford examine documented was worst on code, immediately executable outputs dropped from 52% to 10%. The Gemini 2.5 Pro regression complaints in June 2025 had been nearly totally about code era. 

In August 2025, Anthropic’s personal incident adopted the identical contour: builders on Claude Code reported damaged outputs, ignored directions, code that lied concerning the modifications it had made. Anthropic was silent for weeks. The incident put up solely appeared after Sam Altman quote-tweeted a screenshot of the subreddit. Their postmortem confirmed three infrastructure bugs had been degrading Sonnet 4 responses since early August – affecting roughly 30% of Claude Code customers at peak, with some builders hit repeatedly on account of sticky routing.

The throughline throughout all of it: the extra a activity calls for sustained coherence over a protracted context, the extra uncovered it’s to no matter is shifting beneath. It means your danger profile is totally different relying on what you are constructing. That would not make narrow-context stability a assure. 


What this truly means

Both issues are true. The drift is actual and documented. 

And additionally: your notion shifts. A brand new reference level strikes your baseline completely. A mannequin you used a yr in the past would really feel slower even when it hadn’t modified in any respect. That’s additionally actual.

You cannot reliably inform the distinction between the 2. There is not any public software that permits you to confirm if the mannequin you are working in the present day behaves the identical method it did whenever you constructed on it. Labs publish functionality benchmarks. They do not publish behavioral diffs. The builders most depending on consistency are the least outfitted to detect its absence.

The solely present protections are defensive: pin to dated mannequin strings the place potential, run regression checks in opposition to your key prompts, deal with a mannequin replace like a dependency improve that must be validated earlier than it reaches manufacturing. 

But even the defensive strategy has a ceiling. You can pin to a dated mannequin string. What you can not pin is what’s truly taking place inside it. The mannequin weights, the RLHF tuning, and the protection filters behind that label are totally opaque. Only OpenAI and Google know what they really shipped, and whether or not it matches what they shipped final month underneath the identical title. 

What’s wanted, and what would not exist wherever within the business, is a proper obligation baked into phrases of service: outlined thresholds for what counts as a fabric behavioral change, public disclosure when these thresholds are crossed, and some type of unbiased auditability. Labs presently make these selections unilaterally, talk them selectively, and face no structural accountability after they get it mistaken.

All of this alerts a coverage vacuum no person is pushing them to really feel.

Build a customized agent at no cost (no-code)

Try right here



Source

Add as Preferred on Google Follow on Google News Follow on Flipboard
Share. Facebook Pinterest LinkedIn Bluesky Threads Tumblr Email
Previous ArticleUsing AI-Powered Workflows to Do Work That Wasn’t Possible Before
Next Article AI Is Creating Wealth Faster Than Financial Lives Can Adapt
vinodh
  • Website

Related Posts

AI Agent Hacks McKinsey: When You Should Not Deploy Agents

September 14, 2026

Claude for Finance Teams: DCF, Comps & Reconciliation

September 13, 2026

Did Google’s TurboQuant Actually Solve AI Memory Crunch?

September 12, 2026
Add A Comment
Leave A Reply Cancel Reply

Top Posts

AI Is Creating Wealth Faster Than Financial Lives Can Adapt

September 16, 20260 Views

Are OpenAI and Google intentionally downgrading their models?

September 15, 20260 Views

Using AI-Powered Workflows to Do Work That Wasn’t Possible Before

September 15, 20260 Views

TCS NQT 2026 Preparation Guide: Exam Pattern, Strategy & Selection Process

September 15, 20260 Views

Y Combinator Still Busiest Startup Investor In August As Nvidia Ramps Up Its Dealmaking Pace

September 15, 20260 Views
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Facebook X (Twitter) Instagram Pinterest
  • Home
  • About Us
  • Contact us
  • Privacy Policy
  • Terms of Use
  • Disclaimer
  • CCPA
© 2026 All rights reserved.

Type above and press Enter to search. Press Esc to cancel.