NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Livenerf: Has Opus 5.5 been nerfed yet? (github.com)
jug 4 hours ago [-]
We also have Nerf Bench:

https://www.bridgebench.ai/nerf-bench

They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.

This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

comboy 3 hours ago [-]
But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.
user3939382 3 hours ago [-]
Anthropic A/Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly/5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
jacquesm 1 hours ago [-]
How did pissing off your customers ever become a business model?

I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.

Imagine the power company being able to decide how much you consume and at which price point.

none_to_remain 7 minutes ago [-]
I find it amazingly rich that they bill you for """thinking""" tokens and now you don't even get to see them, they're gonna train the thing to sing "99 Bottles of Beer on the Wall" to itself before it starts work.
pixelready 50 minutes ago [-]
Step 1: Oligopoly Step 2: Regulatory Capture Step 3: Profit
herval 1 hours ago [-]
> How did pissing off your customers ever become a business model?

Airlines, banks, health insurance…

tccole 36 minutes ago [-]
So very low margin businesses with hogh amounts of regulations.
apitman 2 hours ago [-]
Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
ffsm8 1 hours ago [-]
Claude code supposedly has otel you can set via env. I haven't set it up, so I'm just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost

It's meant for their test env I think, so is not documented to my knowledge

braingravy 2 hours ago [-]
Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0
LimitExperience 1 hours ago [-]
[dead]
avazhi 1 hours ago [-]
Nerfbench isn't helpful if it's 3 days old.
andriy_koval 3 hours ago [-]
Usage bench is also very useful! Thank you for doing this!
hamandcheese 37 minutes ago [-]
...is it? I'm looking and it seems like it doesn't have any data. It might be useful if they keep it up.
andriy_koval 33 minutes ago [-]
Yeah, I guess they started it today..
Grimblewald 4 hours ago [-]
I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
eulgro 3 hours ago [-]
Your comment makes no sense. How and why would a local model be nerfed anyway...?
r_lee 3 hours ago [-]
he's saying that he notices a difference between local (not nerfable) and hosted ones, so that it's not as likely to be just placebo
jacquesm 1 hours ago [-]
That and 'loss' may have been intended to be 'launch'.
Razengan 3 hours ago [-]
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..

Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced

Centigonal 3 hours ago [-]
wouldn't less compute result in slower inference, rather than worse performance?
latentsea 3 hours ago [-]
They could potentially quantize the model and run it at lower quality taking less VRAM.
poizan42 1 hours ago [-]
My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.
btown 3 hours ago [-]
The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.
JohnBooty 1 hours ago [-]
I assume there's classification going on where a really basic "Hi how are you?" style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.

I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?

I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.

nightpool 3 hours ago [-]
why is that more likely?
zxilly 2 hours ago [-]
Because they already did so. The model in Codex will get lower `juice` than API version.
jackmott42 2 hours ago [-]
There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn't even happening is happening.

fuck

Rapzid 4 hours ago [-]
The vibe bro science is this always happens on every release, every Tuesday, and twice on Sunday.

Of course it's almost entirely unsubstantiated BS.

fbrncci 3 hours ago [-]
Well now it’s being substantiated!
Rapzid 3 hours ago [-]
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
somenameforme 2 hours ago [-]
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
Rapzid 2 hours ago [-]
Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.

So this is a case of extraordinary claims requiring extraordinary evidence.

And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.

mrandish 59 minutes ago [-]
> They sell HEAPS through their API access

The claim is that they nerf subscription accounts not API.

somenameforme 1 hours ago [-]
If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.

Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.

Rapzid 57 minutes ago [-]
> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account

I'm sorry, but yeah. The official account is a harness regression and some platform bugs.

Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!

I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.

So where is the evidence?!

airstrike 24 minutes ago [-]
You're giving way too much credit to "enterprises" both noticing and publicly airing out their dissatisfaction

Not too mention these companies could easily offer one product to enteprises and another to everyone else

Model nerfing is real

Rapzid 10 minutes ago [-]
Uh huh. OMG you're so right, it's soooo real ;) ;) ;)

I was so certain it wasn't, based on the complete lack of evidence.

But then you said it's real. NVM, I don't need evidence! Somebody said it's real!

This place has fallen off.

stackghost 2 hours ago [-]
To me that sounds exactly on-brand for Big Tech in general and Sam Altman in particular.
jackmott42 2 hours ago [-]
No one here has suggested that the conspiracy theory is stupid (But I will, it is stupid), we are pointing out that the people actually measuring model performance have not found the nerfing before launch conspiracy to be true. in short, yall dumb, shut up.
stackghost 1 hours ago [-]
I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people

> in short, yall dumb, shut up.

no u

Rapzid 48 minutes ago [-]
It's speaks to how low the bar has dropped.

Everyone wants to be a software engineer, until it's time to do software engineering shit.

You know, like scientific method shit we learned in 5th/6th grade.

It's the great bro science incursion.

stackghost 28 minutes ago [-]
>Everyone wants to be a software engineer, until it's time to do software engineering shit.

The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.

sheepscreek 18 minutes ago [-]
> It could also mean nothing happened and people are pattern-matching on noise.

It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.

The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.

What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.

But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.

johnfn 4 hours ago [-]
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

prodigycorp 4 hours ago [-]
Incorrect.

Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

Your chart is wrong.

simonw 4 hours ago [-]
> Anthropic has admitted to nerfing in the past

Where?

QwenGlazer9000 3 hours ago [-]
Earlier in march/aprile, there was a regression in Claude code.

Unintentional tbf.

prodigycorp 4 hours ago [-]
Man this was last year and some Claude subreddit drama that I can’t furnish offf the top of my head but maybe one of the historians remember it.
p-e-w 3 hours ago [-]
Noone can seem to remember anything with certainty when asked to actually substantiate these claims.
computerex 3 hours ago [-]
erinnh 3 hours ago [-]
I mean there is a direct link two comments down from here from 30 minutes before your comment: https://news.ycombinator.com/item?id=49902477
p-e-w 2 hours ago [-]
That’s NOT Anthropic admitting to “nerfing” their model as claimed above (which implies intent), that’s a regression which they quickly fixed.

Christ this forum has become intellectually dishonest.

prodigycorp 3 hours ago [-]
I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.

It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.

winwang 2 hours ago [-]
You can just have your agent find the evidence, review it, copypaste it.
consumer451 4 hours ago [-]
I asked a historian:

Two postmortems, neither quite "admitted to nerfing":

Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]

April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]

So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.

[0] https://status.claude.com/incidents/72f99lh1cj2c

[1] https://anthropic.com/engineering/a-postmortem-of-three-rece...

[2] https://texxr.com/handle/claudedevs

source: https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2

Spooky23 28 minutes ago [-]
> "We never reduce model quality due to demand, time of day, or server load."

That just means they don’t reduce model quality for those reasons.

They didn’t mention other reason, for example, “Make more money”.

what 2 hours ago [-]
> bugs

There are no bugs, just happy little accidents.

consumer451 2 hours ago [-]
u/bcherny does sound a bit like Bob Ross now that you mention it.
johnfn 4 hours ago [-]
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.

I am more skeptical about the compute provider claim - do you have any evidence of that?

r_lee 3 hours ago [-]
I noticed that a few weeks back 5.6 sol would regularly glitch out and start speeding random words or loop and then the next day it'd be fine

and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence

swader999 4 hours ago [-]
Right, and it would be simple to un-nerf or shadow nerf by any kind of angle they want.
gobdovan 4 hours ago [-]
There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: https://www.anthropic.com/engineering/april-23-postmortem

Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.

johnfn 3 hours ago [-]
You are correct, I softened the wording a bit. And thanks for the heads up on the typo!
fendy3002 26 minutes ago [-]
if that's the theory people won't keep using 4.6. Personally I've felt the nerf for 4.8, when 5.0 is (near) launching. And my theory of a model being nerfed several days / weeks after launching has to do with the number of users. At launch there won't be too many users so the computing power per user is huge. As time goes, users and agent has been adjusted to newer model, the computing power per person gets reduced
HawtAds 4 hours ago [-]
It's very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it's a just matter of a single GPU runtime layer bug/update to break inference outputs. The model weights don't necessarily change/get quantized.
bitexploder 4 hours ago [-]
I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
Rapzid 3 hours ago [-]
I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.

Note: I know what quantization is so don't hold back.

r_lee 3 hours ago [-]
I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity

like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions

dannyw 3 hours ago [-]
Inference isn’t flat 24x7, peak hours have more usage, but you buy/rent servers; not servers only for peak hours.

At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.

API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.

Rapzid 2 hours ago [-]
During peak hours requests queue and inference slows. During off-peak they can move systems over to training.

Where is the evidence they are "nerfing" the models due to request volume?

Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.

winwang 2 hours ago [-]
That's a good observation, though I'd say here that two things could be true at the same time. But, I do personally believe that most of the reported nerfing is the case of your chart + latent evidence-less complaining. Honeymoon phases are real.
hbn 4 hours ago [-]
I have not been doing increasingly complex things since Opus 4.6 when models got really good.

My work at my job has stayed the same. But the model quality has varied.

They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.

frde_me 4 hours ago [-]
> I have not been doing increasingly complex things since Opus 4.6 when models got really good.

This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)

With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.

usef- 3 hours ago [-]
What sorts of things, if you can say? Is it a similar sized/complexity codebase? Most projects do become larger and/or more complex over time. And most people's standards do creep up as they learn.
johnfn 4 hours ago [-]
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.

> It’s not a crazy conspiracy that the same model can be stupider

Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.

Computer0 4 hours ago [-]
Open AI admits to such here: Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions -https://news.ycombinator.com/item?id=49804316#49809266
eek2121 4 hours ago [-]
Admittedly, I didn't click your link, however, based on what you've stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.

What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.

Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.

There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.

I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.

Grimblewald 4 hours ago [-]
Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.

My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.

How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?

486sx33 4 hours ago [-]
[dead]
physicallyIllfr 4 hours ago [-]
Its because they're addicts and addicts always grow numb and immune to their fix, needing more dopamine. They want to feel what it felt like the first time.

By the way, dont for a second think LLM hourly limits are all about revenue, they're playing into this psychology. They hire literal gambling UX designers, they want to turn you all into addicts. They want to make you reliant.

Want to run your llm like a slot machine? They'll let you do that spin the generation on a multiple, get 6x results, pick your favorite. Feel that high.

Just know you can get that same hit of dopamine by fostering your own intelligence and creating something with it. Token dealers are just selling you the shortcut, straight to the reward, short circuiting the the natural process.

Bad times ahead for many. This shit isnt good for your brain. And you all know the truth, you just wont admit it. Its doing damage, making you lazier, less intelligent.. Making you an addict.

mwigdahl 3 hours ago [-]
That’s right, and it’s been like this ever since we stopped programming in assembly language. Programmers’ brains used to grow manly and strong on a strict diet of manual memory management and custom stack frame handling. Once we transitioned to soft, weak modern languages like C it’s been all downhill.
2 hours ago [-]
physicallyIllfr 2 hours ago [-]
Ahh, the very original and compelling comparison of llms to compilers.

You're a genius, did you think of that yourself? Or did you local token dealer teach you that?

kdkdjcjejxowjdj 2 hours ago [-]
Ahh, the even more original and compelling “no true Scotsman” argument.

A classic. You go, buddy, use all the cliches you want! Whatever you need to make you feel like a big clever H4cK3r.

physicallyIllfr 1 hours ago [-]
Wild that you signed up for a whole sock acct to make this comment lmfao
cheevly 3 hours ago [-]
You are on actual drugs my dude.
physicallyIllfr 2 hours ago [-]
Keep frying your brain with llms.
cindyllm 1 hours ago [-]
[dead]
nico 4 hours ago [-]
Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often

The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more

I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output

jacquesm 1 hours ago [-]
ChatGPT tends to modulate the speed at which you can type. If it is a bug they probably should have fixed this months ago, so I'm guessing it is intentional.
mlh496 1 hours ago [-]
I noticed this, too. I suspect their classifiers which prohibit certain tasks or require user permission for others is the reason for this.
fendy3002 33 minutes ago [-]
subagent perhaps? afaik subagent by default uses sonnet and perhaps the 5.5 uses different permission definition
pkaye 4 hours ago [-]
I just use auto mode but there is also some config settings for more fine grain control. The model could even help you customize them.
Computer0 4 hours ago [-]
I think most people are on 'auto' mode nowadays.
4 hours ago [-]
xlayn 4 hours ago [-]
The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level.

Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..

I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1

I had the exact same feeling every time they have a new big release

solarkraft 4 hours ago [-]
Did you just say “his”?
Madmallard 4 hours ago [-]
What does HAWAI mean here?
smbullet 6 minutes ago [-]
Have a whack at it?
judge2020 4 hours ago [-]
I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
gr_norm 4 hours ago [-]
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.

The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.

zeroonetwothree 4 hours ago [-]
If it can’t even tell apart Opus 5 and 5.5 (according to the readme) then it’s not useful
solfox 4 hours ago [-]
It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".

After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.

I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?

onemoresoop 39 minutes ago [-]
If servers/resources are overwhelmed it should result in slower responses not degraded quality, or at least have an option for the user to chose from. I'd almost always prefer to wait than get broken or poor results. Even a warning would help, I'd at least not waste my time.
aabhay 5 hours ago [-]
Only ten day interval? I felt Astra got nerfed within a week
refulgentis 4 hours ago [-]
It's always around a week

(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)

bbg2401 4 hours ago [-]
Brilliantly put. Hedonic adaptation is exactly what it appears to be.

It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.

I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the temporary thrill of finding something marginally better than what you have.

mattgreenrocks 3 hours ago [-]
It really feels like it is more about the novelty than actually doing the work sometimes.
aloukissas 4 hours ago [-]
already getting poorer results today
winwang 2 hours ago [-]
Complete anecdote, and nothing to do with relative nerfing or not: Opus 5.5 has been surprisingly good for me (including the past couple hours), especially for following research-level questions/directions.
whs 4 hours ago [-]
I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
madeofpalk 4 hours ago [-]
I've always used Enterprise per-token billing for Claude Code and I've never understood these nerf complaints. I've never noticed any slow downs at certain times of day, or a gradual decline in quality.
ENGNR 4 hours ago [-]
There’s probably contractual guarantees in the enterprise plans. My understanding of the subscriptions is they can swap the models out if any of them is getting too heavily loaded for a period of time
Computer0 4 hours ago [-]
Open AI aims to have a stable API and admits to meddling with effort levels and such for subscriptions here: https://news.ycombinator.com/item?id=49804316#49809266
gaigalas 1 hours ago [-]
The nerfing/quantization strategy is unsustainable. The first lab to not do it wins (short term). The Anthropic pause on Fable might just have been that.

My gut tells me this involves an undisclosed, never-released grandparent model (higher-class than Fable/Astra level, roughly unsellable due to unfeasible cost). That grandparent model is distilled into lower models, of which Opus 5.5 might be an instance of.

That also guarantees protection against distilling a core business. You never make your prime weights available to the public, you only make distillings themselves available.

The downside of this strategy is that you spend a lot of compute on something that you never release, but it might be just the right play (for now) for closed weight companies.

It's a gut feeling, I have zero hard evidence to back it up (it's what I would do as them).

LeoPanthera 4 hours ago [-]
n=1 is useless. The output is not deterministic.
solenoid0937 4 hours ago [-]
Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
sumedh 15 minutes ago [-]
Anthropic has admitted in the past about bugs in the harness after users complained.

Links have been provided by others in this post.

solfox 4 hours ago [-]
No, whether or not it's intentional, maybe can be debated. But there's definitely an experience of a model losing horsepower quickly after launch.
jyoung8607 4 hours ago [-]
Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.

If so, please share. This should be measurable, and I'm glad this project is measuring it.

Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.

Craighead 4 hours ago [-]
prove it
bradfa 4 hours ago [-]
Literally the point of the linked GitHub repo.
nimchimpsky 4 hours ago [-]
[dead]
nba456_ 4 hours ago [-]
No there isn't.
voiceeh 4 hours ago [-]
You mean to write YOU haven't experienced this. Many have indeed experienced this.
doginasuit 4 hours ago [-]
I'm not sure "experienced" is carrying much weight here. This is why there are benchmarks, anecdote doesn't mean very much.
nba456_ 4 hours ago [-]
No you didn't.
omani 4 hours ago [-]
who is paying you to say that?
nba456_ 4 hours ago [-]
Mr. Dario himself
4 hours ago [-]
AnimalMuppet 4 hours ago [-]
The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
jascha_eng 4 hours ago [-]
Yeh it's absurd that people claim this all the time. It's some crazy conspiracy theory and when you ask for examples nothing ever shows up.

It would be economical suicide from anthropic and OpenAI to actually need models intentionally.

But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.

wccrawford 4 hours ago [-]
I was just wondering if, like certain processors, bugs get fixed and the speed goes down. Like, they find it's doing things it shouldn't, restrict it, and harm the throughput.
raincole 4 hours ago [-]
It's not a hot take at all. Every benchmark shows that.
dude250711 4 hours ago [-]
Suuuuure...
empath75 4 hours ago [-]
Yeah people push the models to the limits of what they are capable of almost instantly.
3 hours ago [-]
apt-apt-apt-apt 4 hours ago [-]
Fable 5 seems like it got nerfed when 5.1 came out.
avazhi 1 hours ago [-]
This explains a lot actually.

First two days of this thing was like working with Einstein, then about 24-36 hours ago I started getting frustrated at bullshit that hadn't been a problem before. It was so egregious that I checked to make sure I was still on Opus 5.5 Max.

octoberfranklin 4 hours ago [-]
OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests.

Open models are the endgame.

paradox460 4 hours ago [-]
So make it so all clients pretend to be livenerf at first. Same as the ol' pretending to be Google UA for free articles
octoberfranklin 4 hours ago [-]
I'm not talking about HTTP headers.

The classifier is a model; it examines the actual prompt.

They already do this for the safety "guardrails".

xpct 3 hours ago [-]
This is actually why I've been reluctant to setup my own degradation trackers. I'm afraid it might be too much of a time investment for something that's much easier for them to detect.
onlyrealcuzzo 3 hours ago [-]
This is bad data at its finest.

Truly, madly, deeply sloppy.

lqstuart 2 hours ago [-]
You used Claude to make some slop to see if Claude is getting worse…?
j45 3 hours ago [-]
New model releases that have positive reviews should come with a nerfalert reminder service to make hay until it's shaped and shaped and shaped.
4 hours ago [-]
colordrops 4 hours ago [-]
This repo already has too much visibility now. Anthropic will soon benchmaxx it.
gigatexal 5 hours ago [-]
This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.
bethekidyouwant 4 hours ago [-]
People just tend towards conspiracies you have to actively fight it.
folayii 36 minutes ago [-]
[dead]
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 03:43:10 GMT+0000 (UTC) with Wasmer Edge.