We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.
"GPT-6 models have a tendency to be low initiative beyond what the prompt tells them"
This is a plus in my opinion.
dpc_01234 20 minutes ago [-]
Nothing better than staffing your prompts/skills with long lists of all the things that you don't want, but agent might think you might want even if you didn't ask.
minimaxir 1 days ago [-]
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
joshstrange 1 days ago [-]
> 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
redox99 1 days ago [-]
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
Aeolun 23 hours ago [-]
It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.
threecheese 22 hours ago [-]
Can you elaborate? I've been getting great usage out of my $200/mo plan, and thought I'd try a reset (first time) which was expiring just for giggles. Am I going to get only 20% of it effectively?
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
Aeolun 11 hours ago [-]
I can’t say what will happen to you, but yes, that has been my experience. It is better to wait for your normal full limit to return, because if you use a banked reset you get only 1/5th of the tokens but you still have to wait the full week afterwards for it to reset. 20% would be fine if it didn’t also reset the date your normal reset fires.
seunosewa 5 hours ago [-]
Banked resets do expire if you don't use them, so use them anyway.
nkmnz 4 hours ago [-]
Did you “earn” that reset on a lower tier?
mattkenefick 8 hours ago [-]
How do you get 1 day of usage with Astra?
I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?
redox99 8 hours ago [-]
Currently spending a lot of tokens programming the AI for my videogame.
1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol
jorblumesea 1 days ago [-]
yeah I use sol constantly and have done maybe $15 of spend in the past week. it's solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.
shimman 1 days ago [-]
"You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
trio8453 23 hours ago [-]
> "You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
shimman 6 hours ago [-]
I don't find it appropriate at all, especially regarding technology that workers deeply hate and are skeptical of.
If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.
trio8453 43 minutes ago [-]
Sides? Hate? This is all very emotional. Try to put the facts down plainly and see how ridiculous it is --
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?
How is 2 better in any way for anyone involved? Why would you, as a paying customer, holding it wrong, want other people to keep that information from you?
Anonasty 12 hours ago [-]
Literally the prompting and task definition is main variable how LLM's performs. There are literally millions of examples of vibe coders and new AI adopters who run out of tokens since they don't know how the LLM's work.
onlyrealcuzzo 1 days ago [-]
If you're compacting every 5 minutes, you have a workflow problem - period.
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
ngruhn 1 days ago [-]
Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
SyneRyder 24 hours ago [-]
Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
But that 250k context worth way more than 1M in terms of how well it's utilized, so actually I do like codex trying to keep you at that sweet spot.
bjord 11 hours ago [-]
> unattended overnight claude opus session
yes, exactly
SyneRyder 11 hours ago [-]
Not sure I understand if this was meant as a slight against Claude? Or agreement?
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
bjord 2 hours ago [-]
neither, actually, just that unattended "oneshot" sessions are incredibly token inefficient
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
rrvsh 21 hours ago [-]
Its really not tiny; you can't compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it's perfectly serviceable
kaoD 5 hours ago [-]
Absolutely not my experience. I've been a long time Claude user and at work people slowly started using Claude too. I canceled Claude for my personal account due to the Astra hype, but I'm back after a single month and the absolutely ridiculous context window is one of the many reasons.
I barely compact at work in a very complex monorepo (neither with Fable 5.1 nor Opus 5.5), and yet in my personal greenfield project Astra keeps compacting all the time, to the point of it being unusable.
Benjamin_Dobell 20 hours ago [-]
The context window is configurable. I've been using ~600k for months. No, not API pricing, on a Codex sub.
~/.codex/config.toml
model = "gpt-6.1-sol"
model_context_window = 700000
model_auto_compact_token_limit = 630000
jeremyjh 23 hours ago [-]
I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding tasks use a Luna max agent.
I’ve also found compaction not to be a problem when it does happen.
threecheese 22 hours ago [-]
How do you trigger this? I've been messing with OMP lately for funsies.
jeremyjh 21 hours ago [-]
Its in settings under Tools->Checkpoint/Rewind. I don't know why its not enabled by default and actually forgot I had to enable it. But its a great feature that can really stretch context.
onlyrealcuzzo 22 hours ago [-]
If it's compacting every 5 mins, you're going to notice it in your cache miss ratio and your costs...
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
rrvsh 21 hours ago [-]
It doesn't - try it first
sally_glance 24 hours ago [-]
Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
jmalicki 7 hours ago [-]
Use more subagents.
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
jaktet 6 hours ago [-]
Subagents will inherit the context window at the point in which they are spawned, but it sounds like you're more referring to orchestrating/conducting/managing multiple agents?
jmalicki 5 hours ago [-]
Both... even subagents inheriting the context window doesn't cost a huge amount if the context window was never that large, but yes orchestrating/conducting/managing multiple agents is even better though higher thought cost (but the newer claude agents are really good at this in my experience, part of why I am using Claude a lot lately despite the models being more expensive that ChatGPT's for the same performance when taken alone).
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I've always seen them be prompted, but this is something high priority on my list of unknowns to understand better...
Gareth321 14 hours ago [-]
> Cache doesn't help you much when you are compacting every 5 minutes...
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
RugnirViking 13 hours ago [-]
iirc you can still turn the compaction limit up in codex, though they don't make it easy. It costs way more when you use "large context" though, more than the ~256k that codex allows by default. You can also use the large context via the api directly
AmazingTurtle 1 days ago [-]
you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.
manmal 1 days ago [-]
Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.
apitman 1 days ago [-]
You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.
exfalso 9 hours ago [-]
what. I use Astra xhigh, sometimes max, never ran out of tokens on the 100$ thing. I'm using pi though which is by definition harder better faster stronger than claude code/codex.
codewithcheese 1 days ago [-]
you can config codex to compact at a higher context limit
_davide_ 1 days ago [-]
As a reference i burn 1% percent for every 40 minutes of sol on average
antonvs 1 days ago [-]
Try Gemini. It’s so cheap I often use my personal AI Pro account for corporate work, and most of the time it doesn’t matter.
ChickeNES 1 days ago [-]
Gemini is dumb as hell though, it's not like for like
Marha01 1 days ago [-]
Gemini 3.8 Flash is actually pretty good.
Foobar8568 1 days ago [-]
cheerleader hallucinating agent. That's Gemini.
TuxSH 1 days ago [-]
Exactly half as expensive as Opus 5.5 in every API pricing metric
bigwheels 1 days ago [-]
And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
dotancohen 1 days ago [-]
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?
notatoad 23 hours ago [-]
My side by side evaluation this week was to build a tool for mounting my app’s UI components in a headless chrome and feeding mock data into them, for the purpose of taking screenshots for help docs. Not super complicated, but a real task I needed done.
I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.
Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.
peterbell_nyc 1 days ago [-]
You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
Starlevel004 1 days ago [-]
> OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural.
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
beering 1 days ago [-]
This news and thread is about 6.1 Sol, not 6 Sol. You haven’t even had time to do a fair comparison yet.
TuxSH 1 days ago [-]
> Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
sobiolite 1 days ago [-]
Are you comparing Opus 5.5 with GPT-6 Sol or GPT-6.1 Sol? Because they are different models.
mmis1000 1 days ago [-]
For my personal experience, antropic model have better user experience except for 4.7 and 4.8 though. 4.7 and 4.8 feels like expensive downgrade of 4.6 to me (I didn't know why these two should even exist)
However it's less willing to obey your instruction so it's less usable for general runtine flows.
krzyk 1 days ago [-]
For me Anthropic models from 4.7 to 5 including where bad and ate tokens like crazy. Task delivery was worse than GPT 5.6 and token usage was 2-3x higher.
Looks like 5.5 is the new 4.6
Infinity315 1 days ago [-]
I'm not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me.
phoghed 1 days ago [-]
I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it.
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
toasty228 1 days ago [-]
Try it, it's that good compared to openai current offering.
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
copperx 1 days ago [-]
The usage allowances are now insane, like they were when the Max plans were introduced. The $100 plan is usable again for real tasks.
rspeele 1 days ago [-]
While I have no experience comparing this brand-new model, OpenAI themselves call it "near-Astra" intelligence. I set Astra and Opus 5.5 independently working on the same large research/coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both.
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
agar 1 days ago [-]
This was a very interesting, informative, and well-written comment (and experiment). Thank you.
this_user 23 hours ago [-]
Astra doesn't just burn token at an insane rate, it is also strangely high maintenance when using it. Occasionally, you have to keep prodding it to keep working. Then at other times, it will disappear down some rabbit hole, trying to resolve increasingly hypothetical issues. It feels like you constantly have to keep it on track, while Opus is just churning through tasks.
rrvsh 21 hours ago [-]
Yes, I really don't like Astra - 5.6 models seemed to perform at literally the same level with less opaque prose; I guess Astra is great if you're working on insanely hard mathematical problems (or are fooled by its masked sycophancy) but for coding 5.6 seems to have better taste. I hope that they course correct or at least offer models that do better for coding, or even better that this oligopoly ends
chaostheory 23 hours ago [-]
> I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer).
Going on a slight tangent, I find that I get the best results when I force Codex models (Astra/Sol) and Claude models (Opus/Fable) to consult each other (just have them build a simple skill). There are tasks that neither one can fully solve on their own, but their differences are large enough to make a difference when they collaborate.
rspeele 22 hours ago [-]
I strongly agree!
My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now, but either way having 2 "smart" models available from 2 different companies is a boon.
Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.
They’re comparing against the previous model, not the newly released one (6.1). Why do that on a thread about the new model, I don’t know.
1 days ago [-]
ex1fm3ta 1 days ago [-]
benchmarks.
AndrewKemendo 1 days ago [-]
Only takes 5-10 minutes to test your favorite one shot comparison prompt.
edgyquant 1 days ago [-]
Can you give an example? For me I find that one shot prompts are pretty good it’s only when working with large codebases and complex, multi prompt workflows, that I find the real limitations of models
AndrewKemendo 1 days ago [-]
Yeah the whole Pelican riding the bike is the best obvious one
squidbeak 1 days ago [-]
If 5-10 minutes is enough, you need a more ambitious one-shot goal.
rrvsh 21 hours ago [-]
I found that 6 Sol is dogshit; have you tried o5.5 vs. 5.6 sol? curious to hear if your experience is still the same in that regard
jauntywundrkind 1 days ago [-]
A pity I have to use claude code to try this, that I can't use the tools I know and love and have built around (opencode).
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
17 hours ago [-]
dom96 1 days ago [-]
Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.
Opus 5 scoring higher than 5.5 makes me question of the value of this benchmark to real world usage
dom96 1 days ago [-]
Well, it is genuine.
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
pvab3 7 hours ago [-]
gpt 6 Sol was already supposedly better and 50% cheaper than 5.6 Sol right? I didn't understand why they were keeping 5.6 Sol
vcryan 8 hours ago [-]
People's volume and approach varies. I'm a happy customer and I use my entire double max subscription on planning and analysis and have other models doing all my implementation work because I would burn through my subscription in a day or less. It's difficult to calculate, but I'm something like 10-20 billion token per week consumer and I can't use a US-based model to do this volume of implementation work.
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
verdverm 1 days ago [-]
cache is typically 10%, is this OAI setting a new level at half, 5%?
crazylogger 1 days ago [-]
The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).
sscaryterry 1 days ago [-]
[flagged]
user43928 1 days ago [-]
It's obviously true.
With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.
Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.
I have seen you make these immature accusations that users here are OpenAI employees multiple times today.
sscaryterry 1 days ago [-]
It is not obviously true. Please provide real proof. OpenAI's customers are tired of their BS.
JimDabell 1 days ago [-]
> most people, get hardly a days usage out of a 20x account
This is not even remotely true.
peterbell_nyc 1 days ago [-]
This is the distribution of usage. Spin up a bunch of loops or fire a semi-autonomous factory at a project and it's pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.
If you're running 2-3 parallel agent session with a few sub agents and waiting for you to prompt them, you'll have a very different experience!
JimDabell 1 days ago [-]
> Spin up a bunch of loops or fire a semi-autonomous factory at a project and it's pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.
This is a tiny minority of people, not “most people”.
minimaxir 1 days ago [-]
if an openai employee is reading this plz hire me i am unemployed and i need a job
(Usage limits are entirely dependent on what you're doing with them. If you're not running it on 1 million LoC codebases you can get a lot of mileage out of even a 5x account particularly with the recent cheap models)
whatifitoldyou 1 days ago [-]
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
munksbeer 9 hours ago [-]
Sort of. Hopefully the trend continues and we continue to get advances in the cheaper and open source models, but by all accounts, inside the frontier labs, they get to use much better models that they haven't released.
My main worry is that we get to a point where they have something much, much smarter than anything public and access is gated by extraordinarily high costs.
That probably can't happen though right? Inference is surprisingly cheap compared to the training.
7 hours ago [-]
wyre 5 hours ago [-]
I don't think the limitation is the price of inference of extraordinary private models, but rather scaling to the demand of the model and if that's the case the labs would keep the models a secret, unless they need it for marketing purposes like we saw Anthropic marking Mythos.
jumploops 23 hours ago [-]
The Chinese labs have shown that distillation is incredibly effective, but the major US frontier labs haven’t (yet) been incentivized to shrink their models in the same way.
This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.
ismael_rr 20 hours ago [-]
Distillation is also a broad term - I think most specifically, it refers to training a smaller model on a larger/better model's full output token distribution rather than normal pretraining, which only can access the next token in the data that was actually used.
It's also used to describe the SFT bootstrapping for posttraining, which is what people generally refer to as Chinese labs "distilling".
I would almost guarantee that smaller US frontier models (ex Luna/Sonnet) are distilled from their respective large models.
nevir 19 hours ago [-]
Isn't that roughly what Haiku/Sonnet/Luna/Terra are?
nicce 21 hours ago [-]
> The Chinese labs have shown that distillation is incredibly effective, but the major US frontier labs haven’t (yet) been incentivized to shrink their models in the same way
Have they, actually? A lots of speculation and claims but what is the level of admittance?
albrewer 6 hours ago [-]
Qwen 3.8-27B being on par with Opus 4.6 from ~18 months earlier is a sort of case in point, right?
ashleyn 6 hours ago [-]
If there is high upfront training cost (filtering, classifying data, etc) then what we may find is a situation similar to drug pricing where the cost of the end product remains high and the motivation to keep it closed is there to justify the research input. Any similarity to the pharma industry doesn't inspire confidence in cheapness and openness.
Asooka 4 hours ago [-]
I don't think it will get quite that bad. Making your own pharmaceuticals is practically impossible, because you cannot buy the machines or chemicals needed without jumping over lots of regulatory hurdles. Training your own AI model is just a question of money and time. The hurdles there are mostly technical - how do you read the entire Internet without getting banned. Unless you envision a future where AI training itself is regulated.
jillesvangurp 7 hours ago [-]
There are several moats:
- Training data, harnesses, and processes, that drives the quality of the models. Several companies have those. Some of those companies are in China.
- Infrastructure and funding for running those. That's a scarce commodity currently, training the latest frontier models cost billions of dollars apparently. And if you don't have the infrastructure already and don't have the suppliers on speed dial, good luck getting anything.
- Infrastructure for running inference for running what comes out of those. Several of the key providers of this infrastructure are using their own in house chips for this now. At scale this means huge cost savings.
If you start from scratch without infrastructure, there are a bunch of open source things you can find. But beyond that, you'll have a lot of catching up to do. That's the very definition of a moat.
reticulates 5 hours ago [-]
YC is pumping out training data startups, some of the fastest growing startups are training data startups because OpenAI and Anthropic and Meta and Google are paying them billions.
Infrastructure is readily available. Anthropic and OpenAI don’t own anything, they’re just paying for compute. You might not be able to buy 100k GPUs right now but you can rent it.
You only need to look at how Jev immediately became one of the highest volume models after launch on OpenRouter to see how this industry is moatless.
OpenAI and Anthropic employees frequently leave to start their own labs and raise hundreds of millions for it which they can use to immediately pay for infrastructure and training data. At most OpenAI and Anthropic have… brand and talent.
If anything, OpenAI and Anthropic are heavily disadvantaged because they have huge long term financial commitments that have backed them into a corner, likewise the regulatory pressure… startups have none of that.
pessimizer 4 hours ago [-]
> You might not be able to buy 100k GPUs right now but you can rent it.
And I'm going to keep telling people that it is not unreasonable for individuals to spend $30-40K on an AI rig if they are as productive and useful as many people imagine, my friends and family can take advantage of my rig when I'm not using it, and I'm independent of some snooping corporation - totally private.
There is little reason for these things to be concentrated in far away data centers other than that pooled usage is more elastic (but if I'm sharing my rig with many other people, AI is agentic now, and I can lay off some of my compute to others and they can lay it off on me I'm pretty elastic on my own.) The real reason is because we moved to a rentier economy before AI with the cloud and locked-down phones, and it is a very profitable model that doesn't care if it's rational.
You could subscribe to training material like you would subscribe to a newspaper. The most important training material would be your system learning off you, so training will have to be done in place (we can't rely forever on injecting specifics into contexts.)
These things should be in people's houses, and in the winter you should be using them as radiators. Don't think the power grid can handle it? Put solar on the roof and have batteries at home; helps both you and the grid.
rubslopes 18 hours ago [-]
Yes. I have several subscriptions for cost saving and I'm all the time swapping models according to the cost/benefit necessary for the task at hand.
17 hours ago [-]
pessimizer 5 hours ago [-]
> I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
This depends on what governments do. Crony capitalism in the US can lead to bans on home AI even if they have to root all general purpose computers to do it. The same is inevitable in Europe due to both total elite disdain of democracy and their slavish service to US interests.
Even if Western governments just lock down and antitrust-exempt the frontier labs without rooting personal computers in order to keep out Chinese/open models, I'm in far from a comfortable position when I'm depending on the Chinese government to protect my freedom.
gradus_ad 1 days ago [-]
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
mixdup 1 days ago [-]
Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
luma 1 days ago [-]
Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
OliveronData 1 days ago [-]
> ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
dwaltrip 21 hours ago [-]
RL is part of the model’s training. It changes the weights.
What distinction are you drawing?
kqr 10 hours ago [-]
I think the distinction is between "improving the g factor" and "adapting a given level of g to perform certain types of work better".
OliveronData 10 hours ago [-]
tl;dr it changes the weights, it does not add new ones.
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
dwaltrip 6 hours ago [-]
Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.
OliveronData 6 hours ago [-]
Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
luma 1 days ago [-]
I didn't use the word LLM. I'm talking AI capability, you're focused on this or that current approach to AI. I think it's fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years.
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
neta1337 1 days ago [-]
It is a bit harsh to call it knocking down considering all facts
john_strinlai 1 days ago [-]
>Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
do you think it will be exponential forever?
RobCat27 1 days ago [-]
I think we'll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical/physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I'm young enough that I'm expecting with the rate that we are advancing, I will see AI / LLM analogues developed and run on a quantum computer in my lifetime.
spathi_fwiffo 1 days ago [-]
I think the bottleneck will be the current one.
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
1 days ago [-]
sebzim4500 22 hours ago [-]
Eventually the heat death of the universe will come, so clearly any prediction made needs some kind of time frame attached.
I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)
chamomeal 24 hours ago [-]
Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off
trentnix 1 days ago [-]
Yep. I've made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark's Jarvis at our fingertips.
What a time to be alive.
neta1337 1 days ago [-]
Incredible how many times I read similar comments over the years, containing 'on the cusp' and 'what a time to be alive'. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users
trentnix 1 days ago [-]
Not a single thing? In my house, we are using LLMs to:
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
FiberBundle 1 days ago [-]
> I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
senderista 20 hours ago [-]
You can use LLMs to make your software better, but it's easier to use them to make it worse.
semiquaver 18 hours ago [-]
Companies need to radically change to be able to take advantage. Most companies are afraid to do that and are letting their engineers serve as slow meat proxies, doing software development basically the same way as before.
robryan 10 hours ago [-]
Probably because it is in flux, there is a large scale reorganisation of software around agents going on.
digdugdirk 1 days ago [-]
The difference now is that they've hit the "good enough" point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
famouswaffles 1 days ago [-]
>The difference now is that they've hit the "good enough" point.
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
digdugdirk 8 hours ago [-]
Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation? Or is most of the economic value in the stuff that already exists, and can be performed nearly as well by qwen/deepseek/kimi/GLM/etc?
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
famouswaffles 6 hours ago [-]
>Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation?
Replacing white collar work would be worth dozens of trillions of dollars at minimum. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B Annual Revenue Run Rate in 200 days. And Anthropic is growing so fast they're on pace to hit $100B in Annual Revenue.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic and OpenAI are still growing enterprise usage heavily. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit about what qwen does. And specialized models often perform worse than generalized ones.
willchis 1 days ago [-]
This is how I feel about it. I've stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I'm not the only one either, from comments above like
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
holbrad 22 hours ago [-]
I think this is just a case of the bitter lesson that increasing compute just makes all these predictions meaningless. LLMs just keep going when everyone predicts them to fail constantly.
interestpiqued 1 days ago [-]
4 years is not that long in the grand scheme of things to be fair
dgellow 1 days ago [-]
Those points were true at the time and most are still true now. But they aren’t predictions.
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
agoodusername63 22 hours ago [-]
The amount of irrationality I see in the economy with AI makes me more convinced that the wall street bankrollers know very well they're throwing money into a pit, but it's a pit they're gambling will turn into some world hunger ending AI (that will somehow also keep them making money off of scarcity)
never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.
dgellow 12 hours ago [-]
I’m pretty convinced Wall Street is clueless and relies mostly on vibes. It’s the same people who thought all SaaS would become unnecessary after seeing a Claude code demo. Though eventually they will have to ask for the ROI, and that’s where the whole thing will calm down
JacobAsmuth 1 days ago [-]
The new hardware (TPU v8 and VR) are more expensive but they are significantly cheaper per flop. e.g. many multiples more performance for only 2x the price.
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
moosehater 1 days ago [-]
I was thinking the same thing in terms of running out of data a few months ago. But aren't most gains in the past year+ due to reinforcement learning in some form? Which doesn't need "fresh data" per se, as the model effectively creates the data as it goes. As long as engineers can come up with proper environments, tasks/goals, rewards, and actions, I don't really see data being a limit to model improvement in an agentic sense. Maybe as a knowledge base
dumberquestions 1 days ago [-]
[dead]
dcchambers 1 days ago [-]
[dead]
CuriouslyC 1 days ago [-]
It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.
omalled 1 days ago [-]
I agree that they're not hitting a plateau and I see it in my reserach. I had a math/code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
djdjdkdkfk 22 hours ago [-]
I am a lawyer not a coder, but for me the new models make the exact dumb mistakes they did in 2023. Everytime I come here I feel I'm in an alternate reality.
epihelix 7 hours ago [-]
IANAL, but it sounds as though law hallucinations are the final frontier. Sorry.
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
4 hours ago [-]
RussianCow 21 hours ago [-]
I'm a software engineer and I'm in a similar boat. The models have definitely gotten better, but all the latest frontier models have been within the same order of magnitude of usefulness for many months now. Opus 4.5 was a huge boon in productivity, and I certainly write less code by hand than I did when it came out, but I can't say that my workflow has changed drastically for many months. Every model still takes some amount of babysitting to ensure it's doing the right thing, and they all make silly mistakes sometimes.
robryan 10 hours ago [-]
You get no better result from Opus 5.5 than Opus 4.1 (frontier this time last year)?
LPisGood 1 days ago [-]
Nvidia can start putting weights in silicon if model development slows down.
theturtletalks 1 days ago [-]
I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
serf 1 days ago [-]
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup 1 days ago [-]
The plateau doesn't have to be perfectly flat, but it's not a straight line upward anymore either (kind of like our work on transistors, where we've kind of hit the bounds of speed in clock cycles but are improving on miniaturization and power efficiency)
password54321 1 days ago [-]
It took 6 years to solve ARC-AGI 1, 1 year to solve ARC-AGI 2 and 6 months to solve ARC-AGI 3.
delillos 1 days ago [-]
Those version numbers don't necessarily correspond to equal increases in "difficulty", though.
password54321 1 days ago [-]
Correct, the benchmark became exponentially more difficult as it progressed from pattern matching puzzles to games.
JacobAsmuth 1 days ago [-]
Very true, the sharp increase in difficulty (as measured by human passrate plummeting from 1->2 and again from 2->3) gives an even more stark view of AI capabilities over time.
semiquaver 1 days ago [-]
What universe do you live in that you can look at the past six months and see anything like a plateau in capability?
Edit: removed a comment that was uncharitable and rude, for which I apologize.
arctic-true 1 days ago [-]
Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before.
famouswaffles 1 days ago [-]
I don't have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you'll get nowhere. And i've never seen the term 'brute-force' more abused than these LLM discussions. Basically none of the results have been brute force.
semiquaver 1 days ago [-]
Agreed. You can’t “brute force” reality, which has an infinitely large state space. A million monkeys won’t write Shakespeare and all that
usef- 22 hours ago [-]
Do you use them? I feel like you're talking about the news headlines. But in daily, direct, individual usage (one model at a time) I've found they've improved tremendously over the last few months, for tech/programming work at least
arctic-true 9 hours ago [-]
Yes, I do. They’ve definitely gotten better, I wouldn’t dispute that. But we haven’t gone from useless toys to Skynet in the last 6 months, progress is slower and steadier than that.
CamperBob2 1 days ago [-]
"This machine-intelligence stuff is overrated, they are just using <insert particular machine-intelligence technique here>" isn't the resounding verdict it may have sounded like when you typed it.
mixdup 1 days ago [-]
Not that they've hit it but that they are approaching it. The time to panic and steer the narrative is before you hit the iceberg, not after
1 days ago [-]
phoghed 1 days ago [-]
People have been saying this since GPT-4.
ActionHank 1 days ago [-]
Have we honestly seen that great a leap in the last 6 months, or just better application of what we had 6 months before that.
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
CuriouslyC 1 days ago [-]
The difference between 6 months ago frontier and now frontier in 3d modelling, graphics and video editing is night and day.
ActionHank 1 days ago [-]
Just because there are new capabilities, doesn't mean they've pushed passed the plateau, they've just expanded where the previous solutions work.
We've gone from 80% in some places to 80% in some more places.
coderenegade 16 hours ago [-]
I don't think the math community thinks there's been a plateau. The models have gone from being a joke to being able to crank out proofs to research grade problems.
CuriouslyC 1 days ago [-]
> We've gone from 80% in some places to 80% in some more places.
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
ActionHank 1 days ago [-]
Just checked your website, you really drank all the koolaid huh?
usef- 22 hours ago [-]
Statements like these remind me how much each of us are in our own unique information bubbles. People seemed very excited about analysing the differences of each model on my feeds.
Have you used recent ones on any large projects or coding issues? They've improved tremendously lately in my experience, in terms of implementing non-trivial things, debugging, handling legacy/complex codebases, etc. I usually keep a text document of things that models failed to do/fix and recent ones have wiped it all away.
colechristensen 1 days ago [-]
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
xienze 1 days ago [-]
> sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead 1 days ago [-]
The game is already over
nbardy 14 hours ago [-]
Have you even tried the new models? Opus 5.5 is a clear leap. Go back a year ago and try them and tell me there is any sort of plateau.
azan_ 1 days ago [-]
> Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
People were talking about plateau for years already.
djfjkfkffkkf 1 days ago [-]
China will do to llms what they did to german cars
bogrollben 1 days ago [-]
I guess I'm out of touch. What did china do to german cars?
CamperBob2 1 days ago [-]
Outcompeted them badly. Sent Porsche packing and BMW bawling.
zigzag312 8 hours ago [-]
By being subsidised by the government.
tngranados 8 hours ago [-]
Are German manufacturers not subsidised? There are plenty of electrification grants, direct to manufactures and to consumers that end up improving sales for them
CamperBob2 5 hours ago [-]
That argument doesn't hold much water in my opinion. VW AG is partially state-owned, with a double-digit percentage held by the federal state of Lower Saxony.
Meanwhile, in the US, companies like Boeing, GM, and Intel will never be allowed to experience more than minor financial inconvenience before the government bails them out with protectionism, guaranteed loans, and outright subsidies.
I don't see a material difference between how the Chinese government treats their strategically-important industries and the way we do here in the West. Terms like "socialism," "Communism," and "capitalism" are just fodder for Fox News camp-followers.
Razengan 1 days ago [-]
What they'll do to LLMs. Keep up
coderenegade 16 hours ago [-]
There's still a sizable gap between OpenAI, Anthropic, and the Chinese labs. If anything, it's getting bigger. We haven't seen Chinese models solve the types of problems that ChatGPT and Claude are able to solve.
bendtb 8 hours ago [-]
But 95% of all problems are fairly banal, i.e. if the AI in my dish washer, fridge, radio, bicycle computer, etc. just have GPT 5.5 intelligence for making sure temperature, water, direction, etc. are 20% better controlled than yesterday it will be an enormous win for ordinary people (and the ressources we consume).
2yrrr 6 hours ago [-]
Correct the implicit gamble of frontier labs is to displace humans - and the only firm you should trust with this is an American one.
It’s not happening and the imminent bust is coming. Strap in while the music gets turned up (dots, ipo etc) and people decide to leave the partayyy!
ok123456 1 days ago [-]
We can only hope.
Razengan 1 days ago [-]
Why make a new account just to post this comment?
It's not even anything controversial..
simlevesque 1 days ago [-]
They may work for one of the big AI labs.
necovek 1 days ago [-]
Or a German car company?
Razengan 1 days ago [-]
Maybe they're a Chinese AI car?
neta1337 1 days ago [-]
Are there no new users expected?
alch- 1 days ago [-]
Called djfjkfkffkkf? Not really.
jorblumesea 1 days ago [-]
This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.
it's also why there have been so many calls for regulation and slowdowns.
LeBit 1 days ago [-]
Yup.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
nozzlegear 1 days ago [-]
Exactly what I've been doing. I don't need the all-powerful GPT-6 Math Scoopa, or Opus T-1000, just to write react, svelte and C# for me; my local Qwen3.8 is more than capable, and I can switch to Deepseek and GLM on OpenRouter when I need speed. I just pop in to read the comments on HN for the latest drama and navel gazing, then I click the Hide button and move on. Couldn't give a wooden nickel what their latest and greatest models are capable of anymore, it's just PR buzz.
0cf8612b2e1e 1 days ago [-]
There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
nojito 1 days ago [-]
Great for the consumer.
I remember when bandwidth was super expensive and now it’s dirt cheap.
It would be weird if consumers were completely price insensitive.
simianwords 1 days ago [-]
?! this model launch was around 10% of the dev day and the other time was spent on Dots and things other than models.
minimaxir 1 days ago [-]
That makes sense. There's not really much else you can say about it.
jimbob45 1 days ago [-]
Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
the_duke 1 days ago [-]
The GPT 6 release was ... not great.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
wkcheng 1 days ago [-]
I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
equinumerous 18 hours ago [-]
If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don't mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don't think benchmarks yet capture the "human" factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
stldev 1 days ago [-]
My experience as well.
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
rrvsh 21 hours ago [-]
Hard agree
I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes
dannyw 9 hours ago [-]
Adaptive thinking was a good idea though. The old method of manually specifying how many thinking tokens you wanted as budget was just silly. The rollout might not have been great, but the change is good.
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
keyle 18 hours ago [-]
I'd even go one more, 5.4 was great until 5.5, which was a rug pull.
100% in agreement. I pay for OAI sub and also use Codex exclusively at work for the past 8 months.
I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).
Then when opus 5.5 came out, again same thing + far cheaper and faster.
I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.
I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.
Anthropic models are very thoughtful and have so much taste all around.
stasomatic 9 hours ago [-]
I cancelled Claude because of its thoughtful prose. I prefer one liner responses from OAI models.
bitexploder 1 days ago [-]
I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
trentnix 1 days ago [-]
That's not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit.
r0l1 1 days ago [-]
Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
chronogram 1 days ago [-]
Same here. Astra has been the best thing I've seen. Astra on Low has been my favourite thing so far. Higher levels just mean more cruft, not useful.
twotwotwo 16 hours ago [-]
I am always uncertain about impressions, but mine agree with this. I liked Luna 5.6 on Amazon Bedrock (which got >100 tps) for doing well-specced tasks fast. 6 seems to both be served slower by Bedrock and may spend more turns/tokens to get to the same place, so...not as fun.
And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.
FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.
One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!
moshegramovsky 1 days ago [-]
100% hard agree.
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
jrflo 1 days ago [-]
I'm in the same boat, I'll give 6.1 a shot but I'll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
skeptic_ai 19 hours ago [-]
Anthropic 20x plan only refers to 5h interval. Not the weekly quota. Very sketchy
ozgung 1 days ago [-]
Maybe OpenAI was the only one pacing the frontier.
nxc18 1 days ago [-]
How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
user43928 1 days ago [-]
They never nerfed any model after release.
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.
nxc18 1 days ago [-]
I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
holbrad 22 hours ago [-]
I've heard very little positive press around Sol 6, with a ton of people preferring Sol 5.6 instead.
I haven't used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it's not just a tune.
sebzim4500 22 hours ago [-]
I never used sol-6.0 in part because everyone kept talking about how bad it was.
Astra is clearly far better than anything prior though, so I'm not sure what you mean really.
sigbottle 1 days ago [-]
How large of codebases are you working on? The models have gotten good enough to 1 shot stupid "trivial" throwaway integration projects with 0 handholding (was having RL'd garbage in late 2025), and I'm actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It's quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn't feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it's dumb.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
moshegramovsky 1 days ago [-]
I work on a very large code base (millions of LOC) and I've had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I've used can definitely get that refactor done, but not autonomously. It needs to be small slices. I've yet to see it 1 shot anything really complicated.
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.
I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.
nxc18 1 days ago [-]
It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
moshegramovsky 1 days ago [-]
This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
sigbottle 1 days ago [-]
> It’s 50/50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
beebmam 1 days ago [-]
gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.
diego_sandoval 23 hours ago [-]
Agree.
GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.
4b11b4 19 hours ago [-]
Didn't even bother trying it yet
pampas 1 days ago [-]
That's my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.
jsw97 1 days ago [-]
After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
koyote 23 hours ago [-]
I think the fact that Sol 6 appeared higher than Sonnet 4 on benchmarks shows that benchmarks are completely rubbish and useless.
I've never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.
jstummbillig 1 days ago [-]
Eh. What? Is this common sentiment?
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
phoghed 1 days ago [-]
In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
copperx 1 days ago [-]
Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?
Marha01 1 days ago [-]
They should etch it into an ASIC. The first model worthy of that honor.
Eridrus 1 days ago [-]
Sol 6 definitely feels kind of dumb and worse than 5.6
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
nicce 1 days ago [-]
When GPT 6 Sol & Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can't trust it to do anything big alone anymore without babysitting.
the_duke 1 days ago [-]
On r/codex the sentiment seems to be quite wide-spread.
Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop
NorthSouthNorth 1 days ago [-]
I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
setnone 1 days ago [-]
yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it's just roll out pains
sunaookami 1 days ago [-]
gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: https://help.openai.com/en/articles/10306912-sharing-feedbac...
soulofmischief 1 days ago [-]
I have had the same exact experience. I feel like I'm working with 5.3 again. It is alarming how degraded the experience has become over the last month.
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
btbuildem 1 days ago [-]
That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
jeffybefffy519 1 days ago [-]
Its almost like the "frontier" is a load of marketing bullshit and we should ignore it....
revolvingthrow 1 days ago [-]
There was a model called Astra-Minor, found in the files a few days ago. I assume Sol 6.1 is this, as a last minute panic rename due to Sol 6 being underwhelming while Opus 5.5 turned out really strong. I can't really explain releasing Sol 6 in any other way, especially mere days ago.
nsingh2 1 days ago [-]
I don't understand why they didn't call 6-Sol just 6-Terra. It was 5.6-Terra level pricing with a perf jump.
r0b05 1 days ago [-]
I don't understand any of this naming man. It just gets more confusing.
mattmaroon 21 hours ago [-]
Tall, Grande, Venti should be the standard for everything.
alooPotato 19 hours ago [-]
When i read the first few words of your message i thought you were going to say that the model naming is just as bad as the coffee cup size naming.
albedoa 17 hours ago [-]
That's not incompatible with "it should be the standard for everything". We are desperate for a standard, however bad it is.
ulimn 12 hours ago [-]
A few words which everyone would understand, like "small, medium, large"? I wish we had something, yeah...
makapuf 10 hours ago [-]
Or just tshirt sizes
14 hours ago [-]
FooBarWidget 16 hours ago [-]
I find those (Italian?) names super confusing as a non-Italian. I automatically read "tall" in English and think "ok it's tall as in high so this gotta big" -> "wait, what is 'grande' then and how is that different from tall?" -> "venti? Genshin Impact character??"
quietbritishjim 13 hours ago [-]
Of course the point of their joke is that those are super confusing too.
But the origin is size inflation. The original Starbucks sizes were short and tall. Shorts are still available, in at least a few outlets, despite having not been on the menu for at least 20 years. (For a while a few years ago here in the UK, I even saw short as a column on the menu at one, but that's very rare.) Then they added grande, literally meaning large but maybe meaning more like "extra large". Then when they added an even bigger size, since large was already taken, they just stated the actual size "venti" = 20 (fluid ounces).
(The first time I ever went to a Starbucks had the same confusion as you. I didn't know what an espresso was but at least its sizes of "single" and "double" were easier to understand, so I got a double and was a bit surprised at how small my drink was!)
mattmaroon 3 hours ago [-]
What’s funny about it to me is that it now isn’t confusing because we’ve spent 20 years with it. It’s nonsense but we’ve adapted to it.
jnrk 14 hours ago [-]
Venti is 20 in Italian. 20 oz? IDK
Razengan 1 days ago [-]
I like it. It's cool and better than Opus, Fable, Sonnet etc.
Like who can figure out the ordering? With Luna < Terra < Sol < Astra it's obvious at a glance.
I propose for some third company to name after monsters: Cyclops < Minotaur < Ettin < Cerberus < Hydra < Kraken < Nyarlathotep (the AGI singularity stage)
recursive 1 days ago [-]
> With Luna < Terra < Sol < Astra it's obvious at a glance.
It.. is?
n8m8 1 days ago [-]
It’s not… why do I need to translate from Latin (or wherever these come from) to understand them? Opus, Sonnet, and Haiku do the same thing, and are widely known words. Not to mention all LLMs do is generate tokens; I prefer the homage to writing over a space reference.
Razengan 24 hours ago [-]
Boy wait till you find out where "opus" comes from.. Do you realize how many words from Latin you used in that comment?
When was the last time you heard anybody say "opus" or "sonnet" before this?
Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
clhodapp 23 hours ago [-]
It's not supposed to be about superior and inferior, it's supposed to be about bigger and smaller.
Fable kind of screwed that up, though.
quietbritishjim 13 hours ago [-]
> When was the last time you heard anybody say "opus" or "sonnet" before this?
Not a lot, but occasionally. I was familiar with the words.
> Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
It really doesn't.
The one real weakness of Anthropic's naming is that Fable doesn't fit. A fable is typically a fairly short story whereas I think of an opus (if a written work) as being at least a long book and maybe a whole collection of books.
RugnirViking 11 hours ago [-]
you just know they considered "novel" and "epic" internally and decided against both for being too normal
n8m8 20 hours ago [-]
TIL everyone knows Terra means the planet earth (not a prefix for ground) and everyone knows Astra means a star that isn’t the sun
Razengan 19 hours ago [-]
Wait till you learn that earth means the planet and also the ground
n8m8 18 hours ago [-]
So you agree that Luna vs Terra is less clear than Haiku vs Sonnet?
FooBarWidget 16 hours ago [-]
I find Luna vs Terra more clear. I know what a haiku is only because it hss some sort of hype factor among certain crowds. But I have no idea what a sonnet is, I heard of the word but never seen or written one. Opus vaguely reminds me of something in Greek mythology but also no idea what that, let alone the apparent implication that it's supposed to be longer than a sonnet.
quietbritishjim 13 hours ago [-]
I'm not a big fan of poetry but I think it's bit depressing if you've never seen a sonnet or even heard of the concept.
Shall I compare thee to a summer’s day?
fragmede 17 hours ago [-]
If you want to be pedantic, Earth is the planet and earth is ground/soil.
Razengan 15 hours ago [-]
Also in some languages without letter case though
Centigonal 19 hours ago [-]
haiku (the model family) was not inherently inferior to its larger, more expensive siblings. Sometimes you prefer a cheap, fast model, just like how sometimes you prefer a shorter poem.
bethekidyouwant 22 hours ago [-]
Everyone who speaks English knows Luna Terra Sol and Astra… no one cares what you prefer
recursive 19 hours ago [-]
I always thought I knew English. I suppose it's better to find out now.
quietbritishjim 13 hours ago [-]
But those aren't actually English words. Are you sure you know English?
n8m8 20 hours ago [-]
“I’m right and you’re wrong” type reply
hxugufjfjf 16 hours ago [-]
Yeah, the redditification of HN is ripe in these AI related threads.
theanonymousone 15 hours ago [-]
> Opus, Sonnet, and Haiku ... are widely known words.
You are serious?
wmichelin 1 days ago [-]
yes, each of these things are larger than the other
Razengan 1 days ago [-]
If that holds then the ultimate model must be called Urmom
epihelix 18 hours ago [-]
Of course, by perceived size the order would be:
terra, sol == luna, astra
sydd 1 days ago [-]
and wth is an "astra"?
Sohcahtoa82 23 hours ago [-]
My brain immediately translated "Luna -> Terra -> Sol -> Astra" to "Moon -> Earth -> Sun -> Galaxy", an overall growing of size. While sure, "astra" doesn't directly translate to "galaxy", it's the base of the word "astronomical" and "astronomy".
I feel like people who don't get it immediately are just being deliberately obtuse.
urams 23 hours ago [-]
Astra means stars.
The more obvious, to me, interpretation was distance from us so I would've expected Terra -> Luna -> Sol -> Astra.
When I think of space and astronomical scales, I think of distance as much more common concern.
confidantlake 20 hours ago [-]
yup I thought the same but had no idea wtf astra was. Some meteor belt I wasn't aware of or something? A moon on Jupiter? Idk wasn't obvious to me.
recursive 23 hours ago [-]
I don't doubt that you feel that way, and there's no way I can really provide you any evidence. But I wasn't trying to be obtuse.
ricericerice 22 hours ago [-]
i think you underestimate how much you "just getting it" stems from being terminally online - this is a HN model release thread after all
23 hours ago [-]
albedoa 17 hours ago [-]
> I feel like people who don't get it immediately are just being deliberately obtuse.
Okay. I don't believe that you got it immediately.
It's a language-nerd joke he made. You have to know some Latin basically or enough Latinate roots. Those are the Latin-ish words for, in the same order: moon earth sun stars. So it's obvious in the sense of, it's increasing in size (assuming stars is the collection of all of them while the sun is just one of all the stars or one example of its parent category).
voiceeh 1 days ago [-]
Bigger => Bigger, makes sense to me.
zaphirplane 23 hours ago [-]
It isn’t ? ;) celestial bodies Are memorable
recursive 4 hours ago [-]
Yes. Also, celestial bodies are not ordered. And remembering them doesn't do anything for estimating the power level of a particular LLM named after it.
OrangeDelonge 22 hours ago [-]
Oh thats funny, I previously thought Luna > Terra because I assumed the order was “further away from earth = better”. I guess that speaks for itself.
Yizahi 14 hours ago [-]
I forgot about Terra's model existence and after Astra release I was temporarily confused about Astra and what's that do, because I had an idea that maybe it was about observed size, so Astra should be the smallest maybe?
Honestly, the old boring business lingo of adding Pro and Max suffixes would have been clearer for both companies.
Also, cyclops is assuredly bigger than minotaur in almost any reference. And what's "ettin"?
Wowfunhappy 20 hours ago [-]
Haiku, Sonnet, and Opus also made sense to me, but Fable ruined it. It isn't at all clear that a fable is bigger/grander than someone's magnum opus—actually, I'd think probably the opposite.
I don't think it gets better if you swap "fable" for "myth" either, although when you add "-os" to the end it suddenly sounds grander so there's that. Mythos isn't a model I can use, though.
derpyzza 8 hours ago [-]
an opus isn't a magnum opus fyi. magnum opus is a phrase with means one's greatest work, where opus in latin means "work". but in this context opus refers to an ordered series of musical compositions
globular-toast 15 hours ago [-]
It's the classic "this one goes to 11" problem.
ngruhn 1 days ago [-]
There is also more room upwards. Galaxy? Quasar? Filament?
anon7725 13 hours ago [-]
“Local Globular Cluster” - just rolls off the tongue.
Razengan 11 hours ago [-]
Apparently they already have "Aeon" in the pipeline.
Then they could probably move to proper nouns: Sagittarius, Andromeda..
iamdelirium 1 days ago [-]
I mean, Haiku < Sonnet < Opus < Fable/Mythos makes just as much sense. They're larger and larger works.
machomaster 1 days ago [-]
Many people would get confused on the order of Opus, Fable and Mythos. In my mind, they are even from different groups; fable is a synonym forba fairytale (perhaps with songs), mythos is communicating importance and status instead of a size, while opus is the only one which in my mind communicates a big size. "What did you think about the Tolstoy's book? It was not a book, but a real opus, a tedious, incomprehensible, sluggish monumental opus."
bargainbin 22 hours ago [-]
Oh my god I’ve got it!
Bug < Ticket < Story < Epic < Theme
Quarrel 24 hours ago [-]
Waiting on Limerick to drop.
HDThoreaun 1 days ago [-]
opus fable sonnet haiku is way more obvious than luna terra sol astra imo
rendang 22 hours ago [-]
mythos opus sonnet haiku, yes. But why in the world would a "fable" be bigger than an "opus"
globular-toast 15 hours ago [-]
Is it? The order if Luna and Terra in particular is not obvious to me. I would have put them the other way around, especially given the "to the moon" meme.
phist_mcgee 23 hours ago [-]
Maybe if English is your first language it does.
NolF 1 days ago [-]
[dead]
PepegaRoach 23 hours ago [-]
[dead]
sick_of_slop 1 days ago [-]
[dead]
outside1234 1 days ago [-]
"When in doubt, baffle them with b*llshit"
jrflo 1 days ago [-]
Terra is the "missing middle" model and had no positive brand recognition. Sol was for intelligence, Luna was for efficiency. Luna max was cheaper and smarter than terra light and sol light was better than terra max.
swalsh 1 days ago [-]
I mean, when I upgraded my pipeline from terra to sol saying "it's the same price basically!" I was excited. Probably would not have felt as excited if it was just a version bump.
Not sure that's why they did it. But that was my experience.
hawk_ 1 days ago [-]
Because they asked the model what it wanted to be called?
ttul 1 days ago [-]
Was it a last-minute panic, or just OpenAI releasing an update when the had a bit more training under their belt to make 6.1-sol a whole lot better? Either way, I'm extremely pleased and will be giving this model a shot.
antonvs 1 days ago [-]
[flagged]
jumploops 1 days ago [-]
That seems likely, in the API GPT-6.1 Sol requires reasoning, just like Astra, whereas GPT-6 Sol (and Luna) allow "none"
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
spwa4 1 days ago [-]
Sol 6 was pretty good the first 4 days or so of it's release. Then it was probably dialed back to a lower effort level and it became really bad.
Nevin1901 1 days ago [-]
I love free market competition. We're getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence
1 days ago [-]
jeffybefffy519 1 days ago [-]
Spotted the person who hasn't used these yet....
s3p 7 hours ago [-]
This is worse than GPT4 32k? which was significantly more expensive?
phist_mcgee 22 hours ago [-]
That free market is going to explode when the AI bubble bursts.
bopou 8 hours ago [-]
Safe to assume you are actually shorting the market and not just mindlessly babbling about the impending explosion?
Did Medium not get a response, or is this a display issue?
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
UnboundedContex 1 days ago [-]
I always notice this too. Getting it right seems (psychologically for me anyway) to be a big part of "a good pelican" whenever I look at these. But doesn't always seem to correlate with increasing intelligence of models (measured via benchmarks, experience with the model etc.).
It's not frontier pelican without the back leg behind the bike frame IMO.
RugnirViking 11 hours ago [-]
yeah. Even though high got the leg correct, the pedal the rear leg is on is in front , giving the impression it's somehow reaching under the chain and frame with it's leg. Not great
simonw 24 hours ago [-]
Sorry about that, markdown bug, now fixed.
dankben 1 days ago [-]
Let's be honest, they're all guessing when it comes to rendering order
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
This test became useless. On day one everything renders good, try again after 2 days, we will see garbage results
nananana9 15 hours ago [-]
That would make it pretty useful actually.
nicolamanzini 1 days ago [-]
[dead]
ChromeUltron 22 hours ago [-]
[flagged]
phpnode 1 days ago [-]
What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?
az226 1 days ago [-]
Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
toasty228 1 days ago [-]
Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one
copperx 1 days ago [-]
It feels like an updated 4.6. It's fantastic.
RGS1811 6 hours ago [-]
I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
copperx 5 minutes ago [-]
[delayed]
copperx 1 days ago [-]
> I'm not the only one
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
Aboutplants 1 days ago [-]
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
sockaddr 1 days ago [-]
This is it.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
vividfrier 1 days ago [-]
[dead]
pythonaut_16 1 days ago [-]
Maybe process maturity too.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
scrollop 1 days ago [-]
Probably one of the factors.
Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.
Luckily it's not a mistake as now we have access to
.
.
.
dots.
(and sol 6.1, it seems)
geeky4qwerty 1 days ago [-]
jokes on me, I pay for all the subscriptions.
killingtime74 22 hours ago [-]
Of course they do. The real money makers are not subscription users, but the API users and you can just switch with the model selector.
mckirk 1 days ago [-]
No, we're pacing ourselves to have the time to evaluate the impact each new model could have, obviously.
orbital-decay 1 days ago [-]
Versions is marketing, snapshots/minor variations are easy and the number must go up. Release timing is another OAI's marketing tactic.
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
sharpshadow 1 days ago [-]
Response to DeepSeek’s technical paper and competition.
LPisGood 1 days ago [-]
Which paper are you referring to?
wg0 1 days ago [-]
What's that in summary?
Wheen 1 days ago [-]
Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.
just goes to show that OpenAI in fact did not innovate on a single thing for the better part of a year (one could argue two) and instead keeps immitating what it sees doing others successfully with the tech, all in a very transparent attempt to get people lubed up for their IPO.
jonatron 1 days ago [-]
Probably just the singularity, no big deal
anotha_one 1 days ago [-]
[dead]
MisterMunchkin 1 days ago [-]
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
dannyw 8 hours ago [-]
Both labs are spying? Employees hang out at the same bars, have overlapping social circles, etc. Alcohol does what alcohol does.
denysvitali 1 days ago [-]
They're pacing the frontier
blmarket 1 days ago [-]
and seems like they're claiming Sol/Opus are not frontier (and only Astra/Fable are)
jchw 1 days ago [-]
It is the only way to reduce prices while making it look like a good thing.
lxgr 1 days ago [-]
Wanting to have the newer model than the competitor, presumably.
dandellion 1 days ago [-]
The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
lxgr 1 days ago [-]
We swear, We Really Wanted To Make An "ASI" Model But This Is Literally The Way The Weights Dragged Us This Time
SwabbyNat74 1 days ago [-]
Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!
ChromeUltron 22 hours ago [-]
no patrick,
m̶a̶y̶o̶n̶n̶a̶i̶s̶e̶ a point release of the slopbot is NOT a̶n̶ i̶n̶s̶t̶r̶u̶m̶e̶n̶t̶ RSI
mynameisjonny_ 1 days ago [-]
The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out
motoboi 1 days ago [-]
New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.
jesse_dot_id 1 days ago [-]
No.
anotha_one 1 days ago [-]
[dead]
tjwebbnorfolk 1 days ago [-]
Competition
mattnewton 1 days ago [-]
Anthropic’s IPO?
esafak 1 days ago [-]
Productivity is increasing as models get smarter; we are ascending the singularity. I'm serious.
colpabar 1 days ago [-]
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
infamouscow 1 days ago [-]
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
agluszak 1 days ago [-]
They're releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5
system2 1 days ago [-]
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 1 days ago [-]
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
andybak 1 days ago [-]
DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.
copperx 1 days ago [-]
I was working exclusively with DS 4.1 Flash until Opus 5.5 got me back to a sub. I was disillusioned with what was available.
thraway3837 1 days ago [-]
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
senderista 20 hours ago [-]
Sounds like your dream workflow could replace you with a PM?
system2 14 hours ago [-]
I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus's cost.
white_dragon88 18 hours ago [-]
[dead]
franzcoughka 1 days ago [-]
[dead]
A_D_E_P_T 1 days ago [-]
Looking at the token prices, if this is half as good as 6-Astra for 3D model creation in Blender, it's going to be an absolute game changer.
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
CuriouslyC 1 days ago [-]
From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It's slightly worse at spatial detail but better at aesthetics and little touches.
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
kroaton 1 days ago [-]
That dude is a grifter. His German friend is even worse.
2yrrr 6 hours ago [-]
The poster above carries this tone in his posts where he is the divine one. I bet on a lot of stuff he ain’t got a clue what he’s talking about but hopes people like you don’t catch him out.
ekun 1 days ago [-]
How is it with animations?
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 1 days ago [-]
You need to use the Blender MCP. There is an official plugin for this now, so the third party one can be avoided.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
A_D_E_P_T 1 days ago [-]
I've only tried animating models in Astra-6, and I was quite impressed! It's rarely able to one-shot things perfectly, but it usually gets pretty close.
lukan 1 days ago [-]
Have you tried fable? (I did small experiements and was satisfied, but maybe there are reasons to switch?)
A_D_E_P_T 1 days ago [-]
No, because I always hit my Fable quota (Max 20x) in 12 hours on simpler tasks, and I'd hate to need to buy tokens at API pricing.
therealdrag0 1 days ago [-]
After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.
Aboutplants 1 days ago [-]
“OpenAI's new Pro 500 plan offers OpenAI's highest usage allowance and comes with access to its new "Ultrafast" feature — it also costs $500 per month.
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
They're really (finally?) starting to behave like a company bleeding money.
Our VC-backed subscription days are numbered
glaslong 1 days ago [-]
Alas, I did enjoy burning investor money on my taxis, movies and tokens.
onlyrealcuzzo 1 days ago [-]
> Our VC-backed subscription days are numbered
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
m3kw9 1 days ago [-]
I'm ok with whatever price they give out given they are not a monopoly and have competition, the lock in is minimum for me. This means they have legit reasons to send us this price plan. I don't believe they would shoot themselves in the foot when there is cut throat competition (Claude/opensource) out there.
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
1 days ago [-]
LeBit 1 days ago [-]
Let’s pray Chinese models are not banned.
Madmallard 1 days ago [-]
how could u even ban them? lol
girvo 1 days ago [-]
Same way they’ve banned a lot of Chinese networking hardware: make it impossible for companies to use it.
cromka 11 hours ago [-]
Meanwhile Huawei is doing just fine everywhere else
killingtime74 24 hours ago [-]
anyone can run it on their own laptop
girvo 20 hours ago [-]
That doesn’t help you when your laptop is a corporate one.
ricericerice 22 hours ago [-]
got a link to a laptop that can run K3?
killingtime74 16 hours ago [-]
I got lots of links to models you can run on a laptop.
user43928 1 days ago [-]
Unfortunate that the Ultrafast is only available with the $500 subscription.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
cactusplant7374 1 days ago [-]
Ultrafast uses 6x the usage. They probably realize that people will complain if the plan limits are too low. In any case, the TCO of the newer chips is supposedly lower. Hopefully everyone is on ultrafast eventually.
torginus 1 days ago [-]
I think that's by design - they're going to IPO soon so if they can get a significant percentage of users to switch from the $200 to the $500, they can 2.5x projected revenue.
glub 1 days ago [-]
Yeah, that's not going to happen. They are more likely to lose a lot of customers, unless Anthropic does the same thing.
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
latentsea 1 days ago [-]
For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can't justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it's only getting better from here.
glub 1 days ago [-]
Yes, either US AI corps reduce the cost of their top tier personal subscriptions down to what people are already paying for other expensive personal apps (e.g. Adobe), so ~$50-100, or open weights are going to eat their lunch very quickly. We're not there yet, as current hardware doesn't allow you to do things like multiple parallel agents, but we'll get there soon enough.
$500 for the old $200 is definitely a fumble.
latentsea 1 days ago [-]
I have multiple GPUs now as a way to solve that.
rrvsh 21 hours ago [-]
Surely you realize how rare the ability to do this is
latentsea 18 hours ago [-]
I do not. I had 1 GPU and I had an expensive subscription. I simply cancelled it, and purchased a second figuring if I was going to spend the money anyway I'd rather have something to show for it at the end of the day. I'm not unique or special in my capability to do this. I figure I may as well purchase at least one GPU per year equivalent to what I would have spent on SOTA model subscriptions for that given year.
RussianCow 21 hours ago [-]
People keep saying this but it's just patently not true, or at least not apples-to-apples. You can't seriously compare Qwen 3.8 27B to Fable or Astra. Even if local models get better, so will the frontier, and you'll always be at a disadvantage.
Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.
There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.
latentsea 21 hours ago [-]
You don't need SOTA. You need a model that can accomplish your task. Qwen3.8-27B isn't comparable to SOTA, but can I use it and accomplish most of my tasks with? Yup.
The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.
RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.
RussianCow 3 hours ago [-]
> You don't need SOTA. You need a model that can accomplish your task.
I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.
> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.
seizethecheese 1 days ago [-]
People said the same about $200 a month. I think the ceiling is probably much higher. Companies regularly spend 10% or more of employee cost on offices, SaaS, equipment. I could see these costs going to 10% of white collar income.
glub 1 days ago [-]
> People said the same about $200 a month
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
seizethecheese 1 days ago [-]
I’m not quite sure I understand the API usage point as it relates to regular customers.
glub 1 days ago [-]
You could only use it on chatgpt.com
Now you can use it in coding harnesses that call the API.
kadushka 9 hours ago [-]
Just as $500 makes no sense now
Why? You can use in codex, right?
adonese 1 days ago [-]
Very risky to do so especially considering how well is opus 5.5.
scottLobster 1 days ago [-]
You think these guys care about risk?
cromka 11 hours ago [-]
Their investors do
Computer0 24 hours ago [-]
I am skeptical that individuals on the $200 and $500 plans make up that meaningful of a portion of revenue.
kadushka 9 hours ago [-]
They use a meaningful portion of compute.
moregrist 1 days ago [-]
This is pretty typical product positioning. You want to sell to both high-end and low-end users, so you offer products at a few price points. Then it turns out that that middle is a much better fit for most users. So you start making the middle a worse fit to push most of those users into the higher tiers.
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch 1 days ago [-]
The 200$ plan was appealing because you got 4x usage for 2x the price.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
RussianCow 21 hours ago [-]
> it sounds like a fumble by OAI.
The vast majority of their revenue comes from large businesses buying for their teams, which are almost certainly not going to juggle lower tiers of different subscriptions to save a few bucks.
5555watch 3 hours ago [-]
Maybe. But developers are also private people who will also play with their private subs and projects. In my opinion, their private experience might influence some corporate level decisions.
Being grandfathered by OAI and happy is not the same as having both, and noticing "hmm maybe Claude is much better for my case, Ill suggest that to our manager"
nananana9 15 hours ago [-]
If we're heading to a world where AI spending for companies will be close to salary spending - which is questionable, but is the only way future in which OAI/Anthropic survive - you will most certainly have people whose full-time job it is to juggle providers and figure out how to save a few percent this month.
jpadkins 1 days ago [-]
This is what I did. Hope it works out. The other benefit is you have a more natural method to avoid lock in. A lot of "improvements" to the agent harness I believe are attempts to build customer lock in.
latentsea 1 days ago [-]
At $500 per month, it's cheaper to just buy GPUs and use local models.
tripleee 1 days ago [-]
Have you looked at the prices of GPUs lately?
latentsea 1 days ago [-]
Yup. I got an R9700 recently for exactly this reason. Figured if I'm going to spend $2400 a year I may as well have something to show for it at the end of it.
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
tripleee 1 days ago [-]
What models are you running locally? Are you banking on them improving or do you think they're good enough today? 32GB of VRAM there wouldn't be close to enough to run the best local models.
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
latentsea 21 hours ago [-]
Qwen3.8-27B is a huge step up from Qwen3.6-27B. That release only happened relatively recently but that felt like the 'Opus 4.5' release turning point that SOTA models experienced back when that came out. It was the first time I felt like local models are actually good enough to use as daily drivers now. So, it was only after that point that I switched.
Qwen3.8-Flash-Next is better still if you can run fast enough. If you have a dual R9700 setup you certainly can. That model is even better.
Qwen4-27B has been announced but not released yet. I'm super pumped for it because I already use 3.8 as my daily driver at home for all my personal stuff, so I'm definitely happy to take an increase in capability.
There is clearly still room for improvement in local models on consumer hardware. With the Qwen 27B models, If you have at least a 5070 Ti I think you can get away with running a small Q4 quant if you use KV cache streaming. The 24GB cards can run Q4 comfortably. If you have a 32B card you can run Q6 comfortably. If you have 48GB ~ 64GB of VRAM you can Q8 comfortably. Using llama.cpp Vulkan let's you pool VRAM across cards (even AMD and NVIDIA etc), so my machine has a 5060 Ti and an R9700.
A dual R9700 rig is really the sweet spot right now with the vLLM-radiance fork. If you can swing a 5070 Ti in there as well to retain some CUDA access, then all the better. That's basically the equivalent to spending 2 years on a subscription, but gets you a system that can run Qwen-3.8-Flash-Next and of course the even more capable Qwen4-Flash when it releases. At the end of the two years it'll run even better models I'm sure.
I'm all in on local now.
honkycat 1 days ago [-]
Wow, canceling my sub. Lets see how Claude is doing these days.
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 1 days ago [-]
Well, here is a breaking-news for you: the 20x from Claude is not a 20x on the weekly usage, it's a 20x on the 5h usage, while the weekly usage is simply double the $100 plan...
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
MCArth 1 days ago [-]
If you've used both you know the OpenAI plans don't compare to Anthropic plans _at all_. Claude code subscriptions are probably worth 4x as much in API spend compared to the same OpenAI subscription tier.
nostrebored 1 days ago [-]
I think you have probably started using OpenAI recently -- one draw used to be that it was really, really hard to ever hit limits. If you did, you probably had usage resets available.
I think this is still true provided you're not using Astra.
machomaster 24 hours ago [-]
The shitty thing about OpenAI's resets is that, unlike Anthropic, they also reset the limit (on the next natural weekly reset). It means that of you pushed the reset button 5 days into the week, you only get 2 days' (2/7 of weekly) worth of extra tokens.
cromka 11 hours ago [-]
Most recent two resets were banked?
andriy_koval 1 days ago [-]
people say this, but I am wondering if there is benchmark/dashboard which actually measures this?
the_duke 1 days ago [-]
It used to be bad, but right now with the 200$ Claude sub I find it pretty hard to blow past the session limit.
You have to do a lot of things in parallel.
InsideOutSanta 1 days ago [-]
Yeah, Fable is essentially unusable, it just burns through quota, but Opus 5.5 is great. The $200 plan goes a long way.
diffuse_l 1 days ago [-]
OpenAI 20x wasn't 20x even before that change. I got a lot more from Claude 5x than Codex 20x...
spiderice 1 days ago [-]
You are literally completely flipping reality. Codex was, in fact, 20x. It was Claude that was not 20x until they got caught.
diffuse_l 1 days ago [-]
I'm describing what I got from 20x Codex vs Claude 5x.
Codex is just not worth the money, at least for me.
What's flipped is the value you get for each of those
enraged_camel 1 days ago [-]
You are painting half of the picture, perhaps on purpose? The other half is this: Opus 5.5 is significantly better than both Sol 6.1 and Astra, and with the newly increased limits across the board, it is quite difficult to run out (unless you're spamming agents at Max effort). So it is a much, much better deal than OpenAI's Pro 100.
WinstonSmith84 1 days ago [-]
> Opus 5.5 is significantly better than (..) Sol 6.1
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
1 days ago [-]
cmrdporcupine 1 days ago [-]
"For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. " - Dario a couple weeks ago.
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
spiderice 1 days ago [-]
> According to the company, existing subscribers will keep their current limits for a time, and will later receive a one-time credit to help them make the most of their new reduced allowances
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
honkycat 1 days ago [-]
I just got an email telling me this isn't true. They're immediately cutting my 200, which I've had for like a year.
ndbe 1 days ago [-]
[dead]
surgical_fire 1 days ago [-]
OpenAI is deeply unprofitable, particularly on those pro plans.
The only way is for prices to go up. Way up.
mrtesthah 1 days ago [-]
It really does look like OpenAI is trying to gradually get rid of their subscription plans. Every week there is noticeably less usage available to them while each new model release boasts substantially cheaper API token pricing. If this continues then the two pricing models will eventually be at parity.
surgical_fire 1 days ago [-]
The subscription plans are a huge money sink, that obviously will have to go away or be priced at a ridiculous level to make sense.
machomaster 24 hours ago [-]
This is not true at all, at the most fundamental level. There is a reason why all the businesses (IT, gyms, cars, restaurants, streaming services, music, games, stores, apps, food delivery, magazines, newspapers, shaving blades, parfume, etc) are doing everything in their to get subscribers and are willing to decrease prices in order to get customers who are paying the monthly (or even better, a yearly) fee.
surgical_fire 23 hours ago [-]
This is delusional.
The cost of providing the tokens for a heavy user (and let's be frank, the people paying $200 are likely heavy users) is many, many times more than the $200 recurring revenue they generate.
machomaster 7 hours ago [-]
Let's be frank, you don't know what you are talking about.
Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
You can actually check the approx. financials of OpenAI and Anthropic. The growth is insane.
There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
surgical_fire 4 hours ago [-]
rofl, and I am the one that doesn't know what he is talking about.
> Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
DeepSeek increased prices substantially not long ago. I find their profit margins hard to inspect considering I have very little idea what sort of environment they may get in China (from cheaper energy to government subsidies). I honestly doubt you have any insight here as well.
> You can actually check the approx. financials of OpenAI and Anthropic.
No you can't. They are not publicly traded, and they constantly and selectively leak bullshit metrics, from extremely unclear ARR, to extremely deceiving EBITDA. You willingly eat their bullshit and call me a picky eater in return.
> There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
I see no reason to believe (much less any actual evidence) that OAI or Anthropic have any path to profitability.
If inference (particularly for subscriptions) was in anyway as profitable as you claim today, they wouldn't need private investment rounds like crazy nor they would be desperate to offload this hot potato in an IPO.
82% margins lol. Are you telling me that if you created a machine that turns 1 dollar in 5 what you would do is dillute your ownership of the machine instead of using these fabulous profits to expand the business?
proxysna 1 days ago [-]
I am yet to spend $200 on deepseek this year. Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500. Deepseek is faster, IMO intelligence difference is negligible and it so much cheaper that i no longer care about how much i use it. I never hit any daily/weekly quota or anything like that while working or tinkering. At this point i am OK with being 6 months behind the "frontier", purely on bang-for-buck basis and who cares which shadowy government gets my data.
rapind 1 days ago [-]
So I took Deepseek V4.1 Flash for a spin maybe 2 weeks ago now (before Luna 6 and Sol 6 were announced), and I racked up $100+ in about 2-3 days. It was pretty great, but it uses way more tokens (TPS is fast, but it's way more tokens per turn) than Sol 5.6 which I found to be about it's equivalent at the time (on medium or high, with DS on max). My cache rate was around 98-99%.
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
jbellis 1 days ago [-]
There's a bunch of skepticism in the replies but I ran over 100 tasks against DeepSeek 4.1 Flash and Sol (among others) and I can confirm, it is in fact a little smarter than Sol and a little more expensive than Luna. https://slopcop.com/power-ranking?pricing=api
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
jambutters 21 hours ago [-]
So if money isn't a problem, opus is the best?
jbellis 17 hours ago [-]
Yes. And on subscription pricing it's even cheaper per task solved than DeepSeek.
ffsm8 15 hours ago [-]
wouldnt taht be technically no? Fable should be best if money _really_ isnt a factor?
tbf, i barely feel the difference with opus 5.5 anymore either.
cpursley 1 days ago [-]
Is slopcop your domain, because that is awesome. Wishing you great success with it.
jbellis 24 hours ago [-]
It is. TY!
taylorfinley 1 days ago [-]
Were you using OpenRouter? I've used 1.8bn tokens in the past week from DeepSeek themselves and 99.2% were cache hits. Total cost was $18.13 usd.
faitswulff 1 days ago [-]
For readers wondering, OpenRouter isn’t capable of caching as effectively as DeepSeek is because they will, for instance, switch inference providers in the middle of a session.
Computer0 1 days ago [-]
The way I use openrouter is I find a model/provider combination I like then pin all requests for that model to that single provider.
swingboy 1 days ago [-]
If you disable all other providers but DeepSeek in your OpenRouter guardrails, is that effectively the same thing?
hyprwave 22 hours ago [-]
Sure but you still pay only OpenRouter :)
faitswulff 20 hours ago [-]
Interestingly, OpenRouter will serve inference at cost for a lot of models.
bogdan 16 hours ago [-]
No? They charge 5.5% (I believe) on every top-up. Which is the equivalent of charging you 5.5% on every request.
faitswulff 9 hours ago [-]
Ah, so that's where the profit comes from. Very payment-processor like move, makes sense that Stripe picked them up.
What harnesses are capable of doing this provider pinning when doing requests on openrouter? Can DeepSeek harness (dsh) do it? Or opencode etc
I stopped using opencode because it has some issues with caching, so I suppose it doesn't do this
pests 6 hours ago [-]
Deepseek plugins* all the way down. It should be easy to add support for this.
Using this interesting framework called Cordis I've recent discovered.
hetspookjee 13 hours ago [-]
Open code and pi can do it. Unironically, I let Claude configure these things for me.
pwython 1 days ago [-]
What pests said. And you can make a preset and pass "model": "@preset/deep-seek"
gigatexal 22 hours ago [-]
How can they? If you use openrouter Deepseek fine but can’t you select the model direct from Deepseek you’re basically getting api access models that way.
rapind 24 hours ago [-]
Fireworks directly. At the time they were the best value of cost, speed, ZDR. They got slower on me though, but I think they are retooling, so maybe things have or will get better again. I think fireworks is primarily for when you want to do your own training on top, which I wasn't doing.
ampdot 23 hours ago [-]
DeepSeek, DigitalOcean, GMICloud, and NovitaAI were the only OpenRouter providers that didn't lead to major performance degradation for me
christophilus 1 days ago [-]
This is my experience, too. It's a great model, but it burns tokens if you use it heavily for work on complex domains.
Edit: others have noted the provider and harness matters. My experience is with opencode.
InsideOutSanta 24 hours ago [-]
Same experience. I often see people say how little they spend on DeepSeek v4.1 flash, but when I put 60 bucks into my account, it was gone in a few days of non-exclusive use. I'm actually curious what the difference is. I used it through pi and opencode, but the harness seemed to have no obvious impact on usage.
esafak 23 hours ago [-]
Maybe they just use it less. If you code all day you can go through a billion tokens.
slopinthebag 21 hours ago [-]
check your cache its maybe?
pimeys 1 days ago [-]
How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.
What harness you are using?
rapind 1 days ago [-]
pi. I wasn't even going that hard. I checked the logs for Sep 18 and I did just shy of 3b input with approx 98.5% cache and 5.8m output, which cost around $35. Most of the was a Rust code review exercise with 1 driving agent and a varying number of subagents (up to 6 some times). I do the same with with Sol med/high driving and Luna x-high reviewing and get at least as much done if not more in a day, but I'd use up two x20 weekly allowances for the week. Worth noting that token cost isn't super meaningfull on it's own, because DS is super token heavy (but also great at caching) compared to Sol. (my stats show DS uses 3x the tokens as Sol)
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
ricardobeat 24 hours ago [-]
You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.
rapind 23 hours ago [-]
I agree, and I started sending well defined review packages to Luna x-high instead, which is how I have this setup when using Codex (Sol drives, Luna async reviews, and Sol keeps moving). Same process with DS really dropped my DS usage by a lot and I think that might actually be the secret sauce. Especially if you use Luna through a subscription (probably the entry pro level would be enough). I'm not sure what a Luna equivalent open model is though. 5.6 Luna max was catching a lot of issues while my implementer keeps rolling. Last couple of days I had 3-4 Sol Mediums running with Luna x-high reviewers (async reviews) and one Sol medium orchestrator and Astra X-high to plan everything out.
Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.
kolinko 14 hours ago [-]
Such a high cache may be a sign that your agents are using tools inefficiently, thus taking too many turns. Or, like someone else said - you may use tools inefficiently many agents that rediscover stuff etc.
Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.
What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.
If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.
askingforafrien 12 hours ago [-]
[dead]
the__alchemist 1 days ago [-]
What are those agents doing? I am out of the loop. Bitcoin mining? Blogging? Reddit bots?
drewnick 24 hours ago [-]
I have deepseek agents doing email responses, with real tools (think running quotes, gathering info, scheduling things) and running business processes that used to be done by $35/hr administrative type people. And the capabilities are expanding every day as I learn how to build scaffolding around the model.
esperent 23 hours ago [-]
> I have deepseek agents doing email responses
So, spam?
eru 22 hours ago [-]
Spam is unsolicited, unwanted email.
If someone is asking you for a quote, and your tool replies, it's not spam.
It might be slop for all I know (or might not be!), but it's not spam.
Similarly for arranging meetings with parties that you already have a relationship with.
abustamam 23 hours ago [-]
Not all communications done by a non-human is necessarily spam.
pjc50 13 hours ago [-]
You have an AI handing out potentially legally binding quotes for you??
ducktoysleftout 23 hours ago [-]
I am a tech lead for a useful but non-essential PaaS my company subscribes to. Their product manager occasionally sends me clearly AI-authored emails. I ignore them. I am a believer in the usefulness of AI, but nothing says I don’t care more clearly than sending me slop.
I can’t overstate how bad of an idea I think using an AI for customer interaction is.
teiferer 14 hours ago [-]
Optimum is probably using a middle ground: Use AI to do all preparations like suggesting quotes, gathering background about the customers query etc but then have the human write the email for personal touch. Still much faster than doing it all by hand.
abustamam 23 hours ago [-]
I think it depends on the industry. In my company (insurtech) a large percentage of our sales come from AI engagements (phone and text). The customers know they're talking to an AI and they can elect to talk to a human at any time but many times they don't.
But yeah, send me a non solicited AI slop email or worse, political ad, and you dont get the dignity of me saying stop to unsubscribe. Straight to spam for you.
Kyo91 22 hours ago [-]
People have been trained to "press or say" a number while going through an AI voice menu on phones for over a decade. So a chatbot looks great by comparison in that context.
abustamam 20 hours ago [-]
I wonder if various companies have recordings of me screaming "OPERATOR!!!!" or pressing 0000000000 ****
Mistletoe 1 days ago [-]
I’m lost when I read these sort of comment chains. Free Gemini works just fine for me. Maybe it’s because I don’t use it for programming? How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race.
internetter 24 hours ago [-]
First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. Humanities will stake it out a little bit longer because AI isn’t human, but AI companies would be dammed if they don’t try. Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible. Not to mention all the roles like tech support and customer service. Once they have gotten as far as they think they can go, they will try to turn up the prices. However, they might struggle to do so as models are becoming a commodity. This is why they are arguing for regulation and stating that only they can tame these beasts.
teiferer 14 hours ago [-]
Biologists? Doctors? I'd like to see AI dig through the rain forest to investigate some species and do surgery on your appendix. Before tech support?
Sorry, your comment is just far out of touch with reality.
abustamam 23 hours ago [-]
I think as long as we have open weighted models there will always be competition. Sure Anthropic could raise their prices but it wont take long for someone else to undercut them. The quality may not be as good but people would probably prefer to pay 1% of the cost for 75% of the quality.
eru 21 hours ago [-]
There's plenty of competition. Of course, Anthropic can raise their own prices, but people will just switch. They've done that in the past, when other companies had better offers.
> First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. [...] Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible.
You seem to think raising productivity is a bad thing?
> Humanities will stake it out a little bit longer because AI isn’t human, [...]
This is really silly. Many languages other than English don't use related words to describe 'humans' and 'humanities'. Will it be easier in those languages? Should we rename mathematics to 'humathematics' or so, to make it harder for AI to take over?
skulk 19 hours ago [-]
> You seem to think raising productivity is a bad thing?
Higher productivity means very little to people whose labor plummets in value over the course of a couple of years. One day they won't be needed anymore, maybe that's good for you.
eru 19 hours ago [-]
Ever since at least the industrial revolution we had this same song and dance play out again and again. Yet, unemployment rates around the world are the world are fairly steady (and vary mostly in line with how business-friendly local regulations and taxes are, and so far not much with the current state of technology).
I'm a software developer by trade. I welcome the coming brave new world in which machines can do all the software.
At the moment, they ain't quite there yet, alas.
teiferer 14 hours ago [-]
With that argument we'd all be unemployed ever since agriculture (which 90% of the populace was working in a few hundred years ago) was invaded by machines. Turns out, people get other stuff to do and are as overworked as ever.
kolinko 13 hours ago [-]
US has 1.4-4.4M programmers and devops, but considering they are more costly than many other professions, it’s quite meaningful to economy. Plus now you have non programmers doing agentic programming to speedup their work - i know nondev project managers, financial advisors, marketers and lawyers rolling their own mini apps now.
Also even if your Gemini is giving you nonprogramming output, underneath the model is most likely generating code for certain tasks.
ac29 23 hours ago [-]
> How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race
I dont consider myself a programmer but use LLMs almost exclusively for coding.
The number of people able to create useful software today is much much larger than it used to be and arguably a minority of these people are/were "programmers"
kolinko 13 hours ago [-]
Why did this get downvotes? :o
selimthegrim 19 hours ago [-]
Free Gemini has at least one subtle mistake on every scientific or programming task I ask it. It's usually in the right ballpark but I guess I do stuff outside its training locus.
pimeys 24 hours ago [-]
Work for my company. Research, code, analysis.
MisterMunchkin 1 days ago [-]
How have you spent hundreds of dollars? I’ve only spent 11 and I’ve been using it for four months!
maxdo 21 hours ago [-]
claude 5.5 with price discounts is a step up in terms of quotas, as well as gpt 6.1 sol, i still feel like for complex things claude is way better vs anything else. also for the topic started, opus 5.5 and sonnet 5.5 are much faster too. I was using grok for that reasons just to do faster changes, but now i'm considering to cancel my cursor subscription, clade is much better and it's also as fast as anything else especially with workflows.
teiferer 15 hours ago [-]
> if you don't mind feeding your data to the Meta machine (spoiler: I won't).
How is that different from feeding into the OpenAI or Anthropic machines?
CamperBob2 22 hours ago [-]
Note that $400 a month is a typical car payment. For the same money -- well, OK, probably a little more -- you could be paying off a set of 4 RTX 6000 cards that will run 4.1 Flash all you want, all day and all night, at hundreds of tokens per second.
sbrother 20 hours ago [-]
Ok, but I have two of those cards plugged into fiber internet, and they consistently rent out for $1500+/mo on vast.ai. So your example would be $3000, minus a couple hundred for business internet and electricity… that still buys a LOT of $200/mo frontier model subscriptions.
d675 19 hours ago [-]
which year are they from?
sbrother 16 hours ago [-]
I bought them direct from PNY the week the RTX Pro 6000 was released. Prices are almost double now what they were then.
CamperBob2 20 hours ago [-]
Seriously? People are paying you $1500/month to rent your 2x RTX 6000 Blackwell cards?
sbrother 16 hours ago [-]
Yes, they have been near 100% utilization on vast.ai at $1 to $1.25 depending on demand for the past two months. They're in two separate workstations, which I purchased for around $12k each before GPU and RAM prices went nuts this year.
Now, I have absolutely no clue how long this situation is going to last! But the economics don't really work out for local models while it does.
selimthegrim 19 hours ago [-]
I think the point was this was the spot price he/she would be paying if he/she had to rent and didn't own them
CamperBob2 19 hours ago [-]
You can buy 2 of them from Central Computer any day of the week for $30K US, and that's not the lowest price I've heard lately.
So if you can spend $30K and immediately start mining $1500/month out of thin air, that's a pretty nice investment even if the electricity costs $200/month. Two years later the cards will have paid for themselves entirely and (I suspect) will still be pretty useful.
It does argue in favor of just paying OpenAI or Anthropic for tokens, though.
sbrother 16 hours ago [-]
You do need fairly high end workstations to put them in, adequate cooling, power and very solid business internet as well. I've put more work into keeping them online than I had expected. But overall the economics are fantastic for it right now, the big question is how long it will continue.
teiferer 14 hours ago [-]
If you break down the pure profit per work hour you put in (after all the expenses), how much is that?
CamperBob2 5 hours ago [-]
How important is it that you maintain a good availability record? I have a few RTX 6000s in a server that I use occasionally for NDA'ed contract work and for experimenting with open models, and now I feel stupid for just letting it idle at ~1 kWh when it could be earning revenue. I didn't realize that anyone would care to rent prosumer-grade GPUs.
At the same time, when I need to use the hardware for something, whoever is renting it from me at the moment is going to get unceremoniously booted, and I imagine they are not going to be happy about that. I assume that vast.ai's providers get uptime ratings that drive their work allocation, right?
teiferer 14 hours ago [-]
Doesn't a RTX 6000 run at about $100/month for electricity? Do that x4 and I don't see where your savings would be coming from.
eru 21 hours ago [-]
Electricity costs money.
And there's a lot of labour and effort involved in setting these things up and maintaining them. People don't even run their own email servers, even though the hardware side of that is trivial.
CamperBob2 21 hours ago [-]
Electricity costs money
Not $400 a month, it doesn't. At least not around here (we average $0.12/kWh and I don't personally run my cards over 300W.)
And there's a lot of labour and effort involved...
Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.
People don't even run their own email servers...
People don't run their own email servers because a convenient coalition of spammers, standards bodies, and large email providers have done their best to make running one's own email server almost impossible.
qwerpy 20 hours ago [-]
4x RTX 6000 running at full load with relatively cheap electricity ($0.25/kWh) is around $450/month.
I did the math and decided it’d better to pay for tokens than to buy the hardware and generate them myself.
eru 19 hours ago [-]
> Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.
'Capable' doesn't mean your labour has no opportunity costs.
I think my argument is easier to attack by noting that you can use AI to substitute for much of that labour.
CamperBob2 19 hours ago [-]
That's what I did personally, not being a Linux guru. But to your point, I still had to spend a lot of time and effort dorking around with the hardware.
It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.
eru 16 hours ago [-]
> It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.
At the moment, it's still very easy to switch from one open weights model to another, and even between the closed models. So the 'no rug pull' property is nice to have, but not as big of a deal.
quacktopia 18 hours ago [-]
Power averages $0.30 to $0.45 for a business in the UK if you convert it to dollars. Smaller businesses without load shedding agreements are at the middle / top of that. (there are lots of people complaining about it). Maybe I just have expensive power, but $0.12 per kWh is pretty cheap.
teiferer 14 hours ago [-]
> Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.
Theoretically, the people who hang around HN have immense opportunity cost when doing this. Being capable does not mean it takes no time. Time they could be using for something more profitable (and fun?).
CamperBob2 5 hours ago [-]
That's one way to look at it. Another way to look at it is to point out that the potential cost of delegating cognitive resources to companies like OpenAI and Anthropic is unbounded.
The AI models we currently have still maintain their working state with a finite and laughably-small context window, but that is already starting to change. My .claude directory contains over 300 .md files that I didn't put there myself. Their contents are eye-opening. The question of who owns, stores, maintains, and can access that data is going to become insanely important over the next couple of years.
If you thought LLMs themselves were disruptive and contentious, just wait until the fight over object permanence gets under way. That's when owning your own box full of graphics cards is going to become important. My own bet is that I won't care too much about the electric bill or my opportunity cost when we all find out what the AI labs really have in mind, and what they're going to have to do in order to justify the valuations they're seeking.
TL,DR: it's not about the tokens, IMHO.
thiago_fm 14 hours ago [-]
I think you may have a problem with caching, can you check how much cache % you have?
Whats your harness?
guluarte 1 days ago [-]
same, used a wrapper around cc and i was spending up to $30 a day with basic stuff
taylorfinley 1 days ago [-]
Maybe CC does something that breaks the cache? I cannot recommend Oh My Pi enough. Every default is galaxy brained, and it plays incredibly well with deepseek flash 4.1. My favorite coding harness rn for sure.
rapind 1 days ago [-]
I found omp used quite a bit more tokens than my fairly basic pi setup... but most of those tokens would be cached with DS V4.1, so maybe worth if there are gains elsewhere.
skeptic_ai 19 hours ago [-]
[dead]
sillysaurusx 1 days ago [-]
It’s easy to hit your quota. “Speed up the compilation time of this C++ codebase. Feel free to use several subagents to search through the files in parallel.” That’ll cost you about $200 for a codebase of ~1,000 files.
Subagents are like trading derivatives. You can lose as much as you want.
abixb 1 days ago [-]
What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
MintsJohn 1 days ago [-]
And it's all measured in "intelligence", a completely meaningless term. For coding i'd be much more interested in how much context actually works, what the complexity of algorithms it can understand and create is, for what languages. How much it manages to follow existing structures or that is just adds ad-hoc machinery to pass the test, etc etc.
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
16 hours ago [-]
KeplerBoy 1 days ago [-]
It's still insane that they stopped showing you all the tokens you pay for. They could inflate the billed reasoning token amount by a lot before it would raise any eyebrows.
ldng 1 days ago [-]
Internet Ad business has been like that for a long time with bot click & Co.
mlmonkey 1 days ago [-]
Ain't that the truth.
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
kruipen 1 days ago [-]
Totally hand-wavy and non-objective ... just like how employees are billing their employers.
zer00eyz 1 days ago [-]
It's all Gacha for business.
davedigerati 17 hours ago [-]
you REALLY need to start working on your own local models, no not because your 4080 on a pi harness with qwen 3.8 will be mediocre today, but because in a year when OpenAI and Anthropic have fully enshittified there will be NO regulations and they'll have your chestnuts in a vice and you'll be glad you can tell them to rm -rf
catigula 1 days ago [-]
Yeah, you might get away with a singular 'wrt' with some consternation, but two?
gbacon 1 days ago [-]
> Subagents are like trading derivatives. You can lose as much as you want.
Excellent pithy warning.
sheepscreek 1 days ago [-]
Sadly this is true - for individual folks on the lower end of the spend spectrum.
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
ramraj07 17 hours ago [-]
I have done something similar on a larger repo on claude code max and it didnt even fully exhaust the 5 hour limit.
deadbabe 1 days ago [-]
Why use subagents at all
FearNotDaniel 1 days ago [-]
Preserve context in the lead chat - let the subagents fill up their own contexts then only return the necessary information.
deadbabe 23 hours ago [-]
Why not just have an agent that can branch its context?
gf000 23 hours ago [-]
Forking conversations have been a thing for a long time, and it's not the same thing (e.g. you may have 50% of your context used up at fork time that is carried into "both agents" afterwards)
eru 21 hours ago [-]
Also sometimes _not_ having the context is what you want.
Eg when I have the AI do a self-code-review before bothering a human, I also want to the AI to give the draft PR to a sub-agent that only has the publicly available context that a viewer of the final PR would have; and not all the accumulated reasoning that lead to the writing of the code in the PR.
deadbabe 7 hours ago [-]
It is exactly the same, you’re just doing it wrong.
23 hours ago [-]
nater5000 1 days ago [-]
Why hire a junior developer if you have a perfectly competent senior developer already on your team?
giancarlostoro 1 days ago [-]
You can do a code review on a "less capable" model that costs less, and the key model gets its output / summary, then you can have that model build a plan, and feed it to cheaper models. It's a more efficient approach than just running everything through Opus, and now that Sonnet is a lot better I'll probably use them more frequently, one thing to note is don't ask it to spin up endless subagents, I'd cap it to 2 or 3 at a time, otherwise, yeah you'll hit your limit extremely quickly.
josephg 23 hours ago [-]
Code reviews also work better in sub agents because the reviewer agent didn’t write the code being reviewed.
speedstyle 20 hours ago [-]
(More precisely, doesn't have the reasoning from when the code was written)
d675 19 hours ago [-]
here's a simple task. Collect all info from unstructured email signatures, 3 years of small business emails.
Opus directs and starts haiku/sonnet subagents. Much more efficient than opus reading it all.
Without this explicit direction, Opus burned through usage to create complicated regex parsers & 8 excel sheets.
qarl 1 days ago [-]
Because two agents are faster than one.
phyalow 1 days ago [-]
Time is money. Parallelism is very helpful optimising one to get the other.
knollimar 1 days ago [-]
Money is money too. Increasing contexts costs non zero money, even with cache hits. Also context rot is a problem that subagents help with
apsurd 1 days ago [-]
This truism is intuitive to everyone but always funny to me how everyone never has any time, needs to save time, needs to hire staff workers for every mundane job and robots can't come soon enough… all so we can binge watch Game of Thrones and 90 day Fiancé.
And watch 10 hours of football on Sunday for our DraftKings bets.
apitman 1 days ago [-]
Money is also money, which anyone who makes heavy use of parallel subagents will quickly learn.
georgemcbay 1 days ago [-]
> Time is money. Parallelism is very helpful optimising one to get the other.
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
HDThoreaun 1 days ago [-]
The models get dumb as context fills. Subagents allow them to accomplish a task with minimal context rot. You can also use cheaper models for subagent tasks
proxysna 1 days ago [-]
Afaik there is just pay-as-use with Deepseek
pnw 1 days ago [-]
I spent the weekend trying Deepseek 4 Pro on a Linux porting project and it led me down a complete rabbit hole where Linux wouldn't even boot by the end of the weekend. Waste of $120. Switched back to GPT 6 on Monday and Linux is booting again and I'm making progress.
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate.
Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance.
Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff.
Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established.
Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation.
Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration.
Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers.
Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
forsalebypwner 24 hours ago [-]
> Deepseek 4 Pro
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
rapind 24 hours ago [-]
That's an unfortunate experience. Think of v4.1 flash as actually v5.0 flash. It's night and day compared to the 4.0 flash (and 4.0 flash was unintuitively better than 4.0 pro). I would re-evaluate with v4.1 flash. I'm not saying it better than Sol or anything, but it's in the ballpark.
pimeys 1 days ago [-]
You mean 4.1 Flash which is the first great Deepseek? The one that actually surpasses Opus in my books now.
jrflo 1 days ago [-]
There's a difference between "write this function for me" coding agents and "build this prototype from end-to-end". If you're doing the former, deepseek is fine. If you're doing the latter, it's not gonna work, and that's where the extra intelligence is most valuable.
polytely 23 hours ago [-]
I'm mostly using deepseek 4.1 flash via openrouter, the way I'm doing stuff is:
1. write a sketch of a spec by hand
2. have the llm review the document and question me until it can generate a spec
3. review the spec and revise where needed
4. have it write an implementation plan
5. another round or revision/review
6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.
I've been working for a couple of hours tonight, the total cost of the session is €0.6.
it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.
naikon 7 hours ago [-]
I do the same but with gpt-6-sol (mostly via github copilot cli).
The one-shot capabilities of frontier models are nice for demonstration purposes, but not actually all that useful to me - the result is often an amalgamation of ad-hoc ideas the agent came up with, poor UI and a hodge-podge of a data model, but at least demonstrating clearly where the specs are lacking.
I agree that in practice, iterative design with lots of review and hand-holding are needed to get quality results. Generated code (I mostly program C++) is often bloated and not succinct or "simple" enough to my tastes, so it requires multiple rounds of cleanup.
For core logic, I find having the agent review code is often faster and more valuable than having it implement it in the first place - it can find and fill in the things I missed.
For my own sanity it is very important that I remain in control and understand the generated code when quality and maintainability are goals of the project.
josephg 23 hours ago [-]
One of the big differences using the better models is that you don’t need to hold their hand anywhere near as much. Fable is crazy expensive, but I’ve seen it just one-shot some remarkably complex projects. The code it produces is much better, too.
polytely 22 hours ago [-]
yeah i just think that if I gave the initial pitch to Fable it would have successfully built something that wasn't exactly what I wanted. maybe it is because i went to school for design, but how something works is so important to me and it is hard to get to that point without actually thinking about it step by step and visualizing how interactions would work. I just cant imagine a single one-shot prompt containing enough information to for the model to know what to build. I guess if you have Frontier model money like these silicon valley freaks you could just iterate by asking it to tweak what you didn't like until it is right.
huflungdung 23 hours ago [-]
[dead]
throooooo 1 days ago [-]
This was my experience 3 months ago. I had an Android app that interacted with a Bluetooth device that I wanted to reverse engineer and build my own Linux app for it. DeepSeek was struggling really hard. Claude did it end to end after 3 or 4 prompts. To be fair, I was using a web interface for DeepSeek and the CLI for Claude; maybe that makes a large difference.
sreekanth850 1 days ago [-]
build this prototype from end-to-end, is this how people build serious software with AI?
jrflo 1 days ago [-]
You're not gonna get something that's ready to ship, but as a first pass to get something running yes. Let's you explore far more ideas with only a few hours of agent time.
OrangeDelonge 22 hours ago [-]
And you can’t use deepseek with a loop approach for creating such PoCs?
jfaat 21 hours ago [-]
You absolutely can. On a large codebase. Written in multiple languages. I also use gpt models and for most tasks I prefer ds. It can orchestrate really well, and it stays on task without going off to reinvent the w̵h̵e̵e̵l̵ state machine.
thangalin 1 days ago [-]
I started KeenLore (an emotive audiobook creator) that way. I gave it software specifications, languages, JSON schema definitions, container requirements, hardware configuration (8GB NVIDIA T1000 GPU, 96GB RAM), and zero user interface mockups. For the second round, I asked it to build a completely independent, re-entrant, and data isolated demo system on top of the web application. The demo application included voice generation using one of its voice designs. Here's the output:
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
poilcn 1 days ago [-]
This is how ChatGPT, Cursor apps are built. They spent so much money on PR stunts, but "thousands of agents" can't make an app that doesn't freeze on each keystroke. Not even talking about user-friendly ui
sreekanth850 16 hours ago [-]
We use codex in vscode. Use agentic development but, agents only do things that We agree upon. Like We create impact study, proposals, feasibility and then implementation plan with full regresison, test coverage and end to end smoke test. A new fetaure will be given to front end only if it satisfy all our requirement in real API smoke test. It is not fast. but still we think we get 10X productivity gains and clean code, consistant quality and maintainability. And we dont use AI for front end.
We have also found C# is more effective for agentic development. With Roslyn APIs, strong typing, structured namespaces, and compiler level semantic information, agents can navigate and reason about the codebase with relatively little context. In our experience, C# is surprisingly token efficient for large codebases cmpared to Go or rust.
marknutter 1 days ago [-]
Weird, ChatGPT has always worked really well for me.
sunaurus 22 hours ago [-]
Such issues are pretty common among people I talk to at least. I hear comments about it pretty often, both when it comes to ChatGPT and Claude.
For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.
Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png
I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.
SOLAR_FIELDS 1 days ago [-]
Early on their software was like this, the original ChatGPT desktop app was basically unusable and full of memory leaks that would tank the software. They’ve long since fixed that though
jurgenburgen 1 days ago [-]
The new ChatGPT desktop app is a dumpster fire though, it can’t even scroll properly when streaming the answer.
mitthrowaway2 1 days ago [-]
I don't think most prototypes are serious software.
outside1234 1 days ago [-]
For a prototype or a 1 or 2 use tool, yes, this is exactly how serious people are building software.
piterrro 1 days ago [-]
Oh buddy you have no idea what a good plan and agent harness can do with deepseek…
proxysna 1 days ago [-]
I am doing mostly hardware drivers recently, it works fine for complex work.
logicchains 1 days ago [-]
"build this prototype from end-to-end" works fine with DeekSeek V4.1 Flash, the problem occurs if you're not only building a prototype but want a finished product.
apitman 1 days ago [-]
I've been trying to use DeepSeek V4.1 Flash more and been very impressed. My current (very rough) rule of thumb is that an Artificial Analysis score of ~40 is the crossover point for "good enough" for most of the things I need to do with coding agents.
mchusma 1 days ago [-]
47 is my crossover for serious things (e.g. Grok 4.7 is below the line and GPT 6 Sol is above the line). I mean, Opus 5.5 is way better, but GPT 6 Sol still gets the job done for anything that doesn't require design thinking.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
apitman 1 days ago [-]
What's the most important task you would/wouldn't trust with an agent below the line?
hmontazeri 1 days ago [-]
Had the same experience. I got rid of my pro sub of OpenAI. I’m really freaking impressed.
dzink 1 days ago [-]
Where are you inferencing DS4.1 flash reliably ?
apitman 1 days ago [-]
I use OpenCode Go. I've also used the DeepSeek provider through OpenRouter a bit in the past and it seemed solid.
nullbyte 1 days ago [-]
The intelligence difference between models like DS4.1 and Sol/Opus is NOT negligible.
bel8 1 days ago [-]
The premium price is only worth fo the hardest problems.
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
linuxftw 1 days ago [-]
The issue is once you solve the hard problems, the lower models start messing things up that were working and reverting all fixes for the hard problems. They'll just go off and do dumb stuff.
tripleee 1 days ago [-]
If both DS4.1 and Opus can complete the tasks you throw at it at good enough quality the differences are negligible.
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
ctolsen 23 hours ago [-]
Opus 5.5: $4/$20
Deepseek 4.1: $0.02/$0.60
Just to illustrate how cheap the Corolla is in your analogy. Also Opus output would be $50 without competition.
usef- 22 hours ago [-]
I suspect most people in this position would be using the subscriptions, though.
A $20 Anthropic subscription is $500+ equivalent of API credit.
You can build quite a lot on the $20 plan and get to use the best model.
Their API pricing has healthy margins built in.
LimitExperience 18 hours ago [-]
[dead]
nkjoep 1 days ago [-]
Or just a cheaper way to reach 0-60
_benj 1 days ago [-]
Specially if using the 200mph car when you need it is just a /model away.
fragmede 1 days ago [-]
A car that feels safe to be driving at 200 mph is going to feel more comfortable at 60 mph, compared to one for which 60 mph is at the very limits of its abilities. Analogies only go so far so I'm not sure there's anything to be learned from that though.
usef- 22 hours ago [-]
Yes, even for crud apps there's a huge variability in how nicely you can make them, and how many mistakes or footguns models make along the way. I'm very curious to what level these people are building to.
zozbot234 1 days ago [-]
In the Artificial Analysis index, MiMo 2.6 Pro is smarter than GPT-Sol 6.1 Low at the same cost, and only slightly dumber than Medium. MiMo 2.6 Flash is marginally cheaper and smarter than GPT-Luna 6 Max. (There is no GPT 6+ Terra, which would otherwise be in that range.) These are not negligible or trivial results.
thiht 1 days ago [-]
I've been using Claude Code at work and OpenCode for side projects for a few months. Every OpenCode model I've tried always felt subpar compared to Claude, but good enough. But it changed with DeepSeek 4.1 Flash, I've been using it for the past few days and I've come to forget I was not using Claude, it's a really good model and it's basically free for my usage (I used it almost all the weekend and spent ~$5)
rmaxdev 1 days ago [-]
What do you do? I’m 200 bucks deepseek flash in about 2 months and it’s increasing
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
iammrpayments 1 days ago [-]
I put 5 dollars at deepseek a long time ago, and somehow it has never been fully spent, can’t imagine anyway to spend 200$ on that thing
proxysna 1 days ago [-]
Recently, drivers for a bunch of obscure hardware. Lots of c\c++, that i am ok with but not enough to make hardware drivers (i am just impatient). Just using pi agent with a few plugins.
marknutter 1 days ago [-]
How complex are they? Drivers vs an entire application would be a big difference in token usage, and it could also depend on the type of work being done.
proxysna 23 hours ago [-]
Complex enough where performance matters. I've done entire applications, client, server, infra and ci as an experiment with deepseek v4 pro a few months ago. Works for that too.
LarsDu88 1 days ago [-]
I've spent $200+ on deepseek and this is for making a multiplayer FPS game. Trust me there are use-cases.
And no it did not deliver. A lot of it was re-done by Astra
abroszka33 23 hours ago [-]
> I've spent $200+ on deepseek and this is for making a multiplayer FPS game.
Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.
LarsDu88 23 hours ago [-]
Deepseek v4 pro was able to create procedural dungeons, placeholder 3d assets, 12 weapons, and menus. Struggled with ragdoll effects, and gun modeling in blender.
Astra was able to model low poly enemies, rig them, do simple animations, and greatly improve procedural generation. I have it modeling assets in blender every day, which I often have to go in and fix
abroszka33 23 hours ago [-]
Nice, that's all single player stuff.
LarsDu88 15 hours ago [-]
I have a roadmap for server authoritative coop with rollback and most physics disabled. I can barely think of things to throw at 6.1 Sol to use up my weekly credits to be honest.
ksec 16 hours ago [-]
>At this point i am OK with being 6 months behind the "frontier",
Indeed. A few people [1] pointed on twitter that current best of Open Source Open Weight models are better than all the frontier model 6 - 8 months ago. And we heard Deepseek v5 and MiMO v3 are quite a leap.
I am not understanding the value proposition of these frontier model. Or at least not to a point they are looking at $1T or $2T market cap. Especially MiMO which may be ( correct me if I am wrong ) the most transparent model we currently have?
Or did I overlook or missing something? It is quite scary if true.
I think the only answer to this is you're just not using agents enough, because even with the very cheap pricing, it's still easy to rack up a large bill.
anonzzzies 13 hours ago [-]
> Not sure what kind of usage can justify $200/month of either openai or anthropic,
Making more than $200/month with the results?
some-guy 20 hours ago [-]
Sonnet 5.5 on low is leagues ahead of deepseek v4.1 flash in price per _result_ in my testing. I say this as someone who has the same attitude as you and really wants open weight models to be competitive and succeed.
pmoriarty 20 hours ago [-]
The biggest problem with DeepSeek 4.1 Flash is all that it hallucinates a lot.
Everyone uses it differently. Not sure what the intent is of these kinds of posts.
Some people are able to take advantage of the increased capabilities that frontier models give us. For others, working with open weight models (that are extremely good in their own right) is good enough.
Deepseek models are very impressive. No one's debating that.
But frontier models are definitively ahead. Perhaps not by much in some areas, but there are clear gaps.
And that is kind of the point I think. Open source models add a much needed source of competition to the world of LLMs and are very much an important counterweight to the dominance of Anthropic and OpenAI.
jwpapi 1 days ago [-]
A lot of people having different pricing experience. I think it’s important to understand that caching can differ, than if the agents spend waiting on code, or consume a lot of content. It depends on how you structure you codebase and how explorable it is, how much effort you set and probably some other issues.
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
zzleeper 1 days ago [-]
A bit tired of spending $200 out-of-pocket for openai. What do you use as harness? (for me the harness if half of the benefit... controlling my PC, working from phone, etc.)
pimeys 1 days ago [-]
https://omp.sh/ has amazing defaults and it sips tokens. Works really well with DeepSeek V4.1 Flash.
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
kleinishere 1 days ago [-]
Did you ever try pi by itself? For those new to the pi ecosystem - any rationale to go with pi vs omp?
pimeys 24 hours ago [-]
It's like choosing between vim and helix. I started my career with vim in the early 2000's, customized the whole thing and had my config in a version control.
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
simlevesque 1 days ago [-]
I use Claude Code + eternal terminal + tailscale + tmux + some custom skills to get notifications through nfty.sh.
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
What makes it run well on low bandwidth environments? Is it eternal terminal? (I am not familiar with that).
I found that ssh is pretty bad in these scenarios, so my usual herdr over a phone doesn't always work properly.
Now I'm using paseo which in principle solves my issues properly, but unfortunately it's "reconnection" and state sync is pretty slow (probably going over their servers).
simlevesque 7 hours ago [-]
Yes Eternal Terminal has been a godsend for bad connectivity and switching networks.
The other part is that since SSH only sends the changes on the screen, it uses very little bandwidth.
Try Eternal Terminal it's very easy to setup and is just a simple layer over SSH.
gf000 6 hours ago [-]
My main gripe is latency. Typing out letters with bad connectivity is very error prone - a native text box is all I would need, but I will give it a try.
IOT_Apprentice 1 days ago [-]
I’m curious about your use of Tailscale, is that for you to reach a local LLM remotely from anywhere?
simlevesque 1 days ago [-]
I have a big beefy desktop at home which I ssh (using EternalTerminal instead of raw port 22) into. I built it last summer right before the prices got very expensive. It's headless so I use a cheap Macbook Air to connect to it at home and I use Termux on my phone to continue working from everywhere.
cruffle_duffle 1 days ago [-]
Dude prompt your agent to set up always on remote connections via systemd… then you can drive from Claude or codex mobile apps natively over their native hookup. Works great.
simlevesque 1 days ago [-]
I don't want to rely on Claude of Codex mobile features at all. Also none of this works when you use third party LLMs which is what the current comment tree is about.
windexh8er 1 days ago [-]
This. There's so many better ways to drive a fleet of agents this way than using "remote" features of which I don't trust, anyway.
phyalow 1 days ago [-]
I have a server living in my home office, always on. I have a tmux session on it with vanilla Codex and Claude Code CLI, I can via my Ubiquiti network stack wiregaurd in to this box anywhere on the globe with just my laptop. Works super well for me. I also have some cheap shelley power plugs that I can use to cycle my PC’s power state if needed.
marknutter 1 days ago [-]
This is basically my setup but I'm using tailscale and zellij. I don't have any contingency plan in place for my power or home internet going down though..
stavros 21 hours ago [-]
I used to do the same but being unable to copy/paste easily, follow links, upload files, etc led me to write a small bridge so I can turn Github issues into my harness frontend:
curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git am
steve-atx-7600 17 hours ago [-]
You must not be a software engineer shipping features that span multiple systems constantly as fast as possible in order to differentiate your company from competitors.
byzantinegene 16 hours ago [-]
90% of software engineers are not, so they don't need a premium model.
23 hours ago [-]
ApolloFortyNine 1 days ago [-]
The speed of deepseek is insane to experience after using claude code with opus for so long. Not only is the tps roughly 3x faster, but the round trip times are magnitudes faster.
mrbonner 24 hours ago [-]
$200/month is for Navier-Stoker grade problem.
laurels-marts 15 hours ago [-]
> which shadowy government
A bit of false equivalency there
dzink 1 days ago [-]
Where do you do your inference?
redanddead 17 hours ago [-]
Abundance is a value all its own
giancarlostoro 1 days ago [-]
Are you just using it directly from them?
proxysna 10 hours ago [-]
yeah
m3kw9 1 days ago [-]
Kind of ambiguous without saying token amounts and cost.
paulddraper 23 hours ago [-]
Deepseek isn't as good as Kimi, let alone the frontier models.
It's a very noticeable hit. It's like programming with mid-2025-era models: ignored instructions, dead code, mistakes.
It works...but anyone going from Sol to Deepseek is going to have a rough transition.
FailMore 1 days ago [-]
API pricing?
UltraSane 1 days ago [-]
Opus 5.5.is crazy good. I ask it to do things and it just writes the code to do it.
pdntspa 1 days ago [-]
I just ran a huge text/image extraction grudgematch against all the current inexpensive models except gpt-5.5/5.6/6 (due to some issues with openrouter and bugs in my code) and DS4 ranked very poorly. Accuracy winner was Gemini 3.8 flash with minimax M3 and qwen 3.8 placing, and the chinese models beat the incumbent (Gemini 2.5 Flash) on cost whilst keeping like 95% of the accuracy.
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
alfalfasprout 1 days ago [-]
It's trivial to hit that kind of quota if you're trying to execute on major projects. Especially as you start having dozens or hundreds of subagents investigating, prototyping, and working on different things.
case540 23 hours ago [-]
You clearly haven’t used opus or astra or only had simple tasks. Such a difference maker
fraywing 1 days ago [-]
> GPT‑6.1 Sol matches GPT‑6 Astra at roughly one-fifth of the cost
Astra is a pretty impressive model. Excited to try this.
gobdovan 1 days ago [-]
They have also cut allowances for subscriptions in half. So even in the best case scenario it's about 2.5 times cheaper for Codex users. They just seem to have matched Claude Sonnet 5.5 *API pricing*, but from what I see online, it seems Claude Code now has a much more generous subscription allowance.
Tadpole9181 1 days ago [-]
Only for the $100 subscription, correct?
gobdovan 1 days ago [-]
Only for the $200 one. The $100 one was already pretty poor value for allowance/$. Without the old $200 sub, I wouldn't have used Codex.
pazimzadeh 1 days ago [-]
Can someone explain to me why on these benchmarks like these a higher effort level often has a lower score?
For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive
Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..
zamadatix 1 days ago [-]
More thought can cause the important info to leave the context or hallucinated info to be enshrined in the context and later acted upon, especially in long horizon benchmarks like DeepSWE.
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
heaney-555 22 hours ago [-]
Some of the benchmarks subtract score for doing tasks that are out-of-scope or beyond what the user explicitly asked for, and the higher effort levels often result in this.
intenex 1 days ago [-]
They released GPT 6 Sol literally 6 days ago. We've accelerated to a weekly model release cadence. That seems like...a big deal.
zarzavat 1 days ago [-]
It's more like they released GPT 6 Sol too early because they were under pressure and now they are releasing the real version. You cannot do anything more than minor post-training in a week.
thoughtpeddler 15 hours ago [-]
Are you saying that 6.1 is based on a new pre-train base model? My intuition would lead me to believe that Astra (GPT-6) was the new pre-trained base model, and 6.1 was a post-trained fine-tune that came out a short time after, but curious to hear if I'm wrong about that..
zarzavat 8 hours ago [-]
I'm saying that the gains were not a result of 1 week of post-training 6.0 Sol.
It seems that 6.1 Sol is closely related to Astra. As for whatever 6.0 Sol was, it's anyone's guess. I'd speculate that 6.0 was just 5.6 with post training. It's possible however that it is the same model but with severe performance issues resolved so they bumped the version number.
toasty228 1 days ago [-]
Implying they don't have like 3 or 4 "models" (different quants, post training, plain renaming) on the back burner at any point in time to do exactly that
ychnd 1 days ago [-]
I think this was supposed to be new Astra, but it came out shit. So instead of shelving it, they found a way to make it still seem like progress.
tripleee 1 days ago [-]
It's a competitive strategy with Anthropic, not necessarily releasing as soon as they're done
nater5000 1 days ago [-]
>That seems like...a big deal.
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
harshityadav95 51 minutes ago [-]
Far More mileage in Codex but noticing slower speed in gpt 6.1 sol medium
modeless 1 days ago [-]
GPT 6 Sol is obsolete after only one week! I am glad that they are not afraid to update the models more frequently. The Navier-Stokes thing revealed that it took them only a week or two to train a model more capable than Astra, and I want the pace of public releases to keep up with that.
pmdr 1 days ago [-]
I had it write some code the other day, boy was it awful-looking compared to 5.6. Worked perfectly, but ugly nonetheless.
samuelknight 1 days ago [-]
Sol 6 was a flop. Nobody would have cared if it was called Terra 6.
aabajian 1 days ago [-]
Opus 5.5 is on another level, especially when it comes to mathematics implementations. You can drop it a PhD-level physical simulation (for example, a contrast-injection simulation for angiography in my case), and it just...implements it. With full-on WebGL rendering in the browser, from scratch (or using an existing library, if you prefer).
jdprgm 1 days ago [-]
6.0 Sol was literally a week ago... Basically continuous integration for model releases at this point.
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
alright2565 24 hours ago [-]
They do, when I was on the $20 plan, I got shown a Luna Reserve model which had its own dedicated quota.
bayesianbot 1 days ago [-]
Interesting idea, but at the same time it is just so cheap that you can just run it with API pricing. I sometimes do even if I have available usage that I'm going to cap so I'll save it for bigger models
poisonborz 1 days ago [-]
As many have surmised, this may have been a panic rename of Astra 6.1
machomaster 23 hours ago [-]
Astra 6.1 light, not the full-fledged version.
iamdelirium 1 days ago [-]
I wonder if releasing this soon sort of validates the rumor that Sol 6 was just the Terra model they bumped up and slashed the price.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
squidbeak 1 days ago [-]
Whether it was or wasn't, Terra's absence shows Sol has replaced it as the new middle model.
hyperpape 1 days ago [-]
If that were true, they’d have axed their margins.
TomGarden 1 days ago [-]
Impressive improvements, but GPT 6 Sol came out 7 days ago, and this one will behave differently. The panicked pace is becoming a liability, maybe they should have waited and released this as the 6.0 release
enraged_camel 1 days ago [-]
They are behind, hence the panic. On top of that, Sam has been trying to do another funding round, so he's desperate to make the company look good.
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
atonse 1 days ago [-]
People said Astra was a gut punch and that Anthropic was reeling. (Opus 5 was almost universally panned)
The best thing is that we benefit from these constant back and forth gut punches :)
JacobAsmuth 1 days ago [-]
I don't know why anyone was saying that when Anthropic clearly knew Opus 5.5 significantly outperformed Astra at the time of Astra's launch. I think it might be a good exercise to go back and find out who called Astra a "gut punch" and lower your credence in their future claims.
atonse 22 hours ago [-]
Random people on twitter, sorry, that is the best source I have, not having any visibility inside these companies :)
So I'm happy to accept your view just as easily, stranger on the internet, since it logically makes more sense.
jhonof 1 days ago [-]
I actually heard the opposite, the folks at Anthropic felt pretty confident they were ahead after the Astra release because it wasn't as good as they were expecting.
ylsilva 1 days ago [-]
For most sane people, OpenAI is the way to go... A lot of usage with very good models, but you know that Anthropic is laughing all the way to the bank with Opus 5.5 being "the best" model right now... There are a ton of people (and companies) that will just refuse to use anything else than the highest benchmarking model in existence.
xyzzy123 1 days ago [-]
Whats interesting right now is that those people also see opus 5.5 as being CHEAP because it's something like half the price of fable or astra.
user43928 1 days ago [-]
It is at least 3x cheaper than Astra in the 20x subscription.
So yes, it is clearly cheap in comparison.
jfrbfbreudh 1 days ago [-]
I have personal Max subscriptions for both. Nothing OpenAI offers touches Opus 5.5.
Stevvo 1 days ago [-]
I'm on both $20 plans. Got more out of Anthropic these last few weeks. But the value proposition shifts constantly.
mark_l_watson 4 hours ago [-]
Because of the drastic cached input costs, they used the published DeepSeek KV caching techniques, right? I didn’t see any public acknowledgement.
thunfischtoast 3 hours ago [-]
Reminds me of the old "I made this" meme
jjcm 1 days ago [-]
Here's a comparison of a image->html flow for GPT 6.1 Sol vs Opus 5.5.
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
Low key ad instantly sold me. Awesome looking tool, bookmarked
jjcm 20 hours ago [-]
Thanks dude, it's been fun making it.
I am trying to make these webpage buildouts more legitimate benchmarks though, as I do think image->html flows are going to become more popular as people realize how good images as a starting point are. Hit me up if you have any feedback!
jumploops 1 days ago [-]
If the Terminal Bench 4.0 scores are to be believed[0] GPT-6.1 is an incredibly efficient model.
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
This is great. But maybe part of the motivation is that 6-Sol wasn't as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.
mkaic 1 days ago [-]
I got a popup in my Codex just now saying "Try out 6.1 Sol!" and so I clicked the button to try it, and intriguingly, it set my model selector to "GPT-6 Astra Light" which makes me think 6.1 Sol may be in some way just a lighter/distilled version of Astra? defo interesting, not sure if I should read too much into it though. I see no option for directly selecting 6.1 Sol in my Codex Desktop UI.
6 hours ago [-]
solarkraft 1 days ago [-]
I wouldn't read too much into it. I've gotten this popup for a model I hadn't had access to yet before and it resulted in what you describe.
slekker 1 days ago [-]
Astra Light is the default option in the UI, so likely a bug
skerit 1 days ago [-]
So their original plan was to axe Terra, but then introduce an "Astra Light" model a week later? They had a nice lineup named for a whole 3 months, and they're already messing with it.
recursive 1 days ago [-]
Wait, so is coding not solved?
alirezaxdehghan 10 hours ago [-]
more like testing is not solved
welder 5 hours ago [-]
Is anyone else over all the metaphorical model names? Sol, Opus, Sonnet, Astra... Just give them versions instead of trying to be the next Apple and name every release.
rpdillon 5 hours ago [-]
Its easier to understand when you see the progression, IMO:
Haiku -> Sonnet -> Opus -> Mythos/Fable
Luna -> Terra -> Sol -> Astra
Weirdly it seems like OpenAI originally just released by version numbers for their models and have switched to Anthropic's naming approach of late.
becquerel 5 hours ago [-]
Don't you remember how awful they were at using versions? When o1 and 4o were both products available at the same time, or how they skipped from o1 to o3?
jrecursive 5 hours ago [-]
6.1 is usable on/around max, but it's no Astra, and it's currently very, very slow. OpenAI needs to get it together and throw their Pro 200 subs a bone with better performance. Still feeling screwed on the Pro 200 sub dilution after a good sleep. I understand the reasons, but it still stings. A lot of goodwill evaporating.
nzoschke 1 days ago [-]
OpenAI is feeling really competitive again.
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
I'm surprised by the crazy pace at which these models have been coming out over the past few weeks. The cheaper cache here is really nice, but I think it depends entirely on your workflow. If you're constantly compacting or burning through your context, you're probably going to have a bad experience regardless of which model you're using
codewithcheese 1 days ago [-]
sol-6 is terra-6. They figure that no one was using terra and they could bring the speed and cost saving of terra distilled on astra, but rebranded as the more popular sol.
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
However, that doesn't say much. You can just run a smaller model at a larger batch size to get higher throughput but lower interactivity.
5555watch 1 days ago [-]
It's not that terrible then if that's the case. 5.6 Sol was superb in my eyes, and while I've tried a million things, there hasn't been anything that 6 Astra improved or did better than 5.6 Sol, while costing a ton of time and resources. Including research topics, where it should have excelled. So even if 6.1 Sol is no worse than 5.6 but cheaper and faster, the gutted Pro200 might still make sense.
neosat 1 days ago [-]
Do these benchmarks have any meaning anymore? And do the announcements seem less exciting now? (Not taking anything away from the advances we are making but it seems more incremental now?) The reliable way to tell if you'll like a model is reliable collage/X reviews to gauge a model's capability and then trying it out to see if you like the style.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
CharlieDigital 1 days ago [-]
I fully believe that these models perform better in benchmarks versus their predecessors, but in real world usage inside of real, production codebases? They feel just as flawed as ever. I honestly have not seen any significant improvement in a few months. The last thing where I felt "wow" was `/fast` mode and Deepseek.
paskejl 1 days ago [-]
Wholeheartedly agree. Astra was some improvement over 5.6-sol in the sense that I'd "argue less" with it, but still frustrating and still sloppy. I'm starting to feel people are not honest about their experiences, they do very simple things or have very low standards. The biggest improvement i've seen from Astra so far is speed.
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
CharlieDigital 1 days ago [-]
> I'm starting to feel people are not honest about their experiences...
I think it falls under:
1. They don't actually look at/care what the agent is producing as long as it works (not planning on maintaining/ops yet).
2. They are using 3rd party benchmarks (which is fair given how widely real-world workloads change from day-to-day, feature-to-feature, making it difficult to really know how well the models would perform).
3. They are doing greenfield work where there is no scaffolding, no existing code, no legacy code, nothing to guide the agents along. I believe in these cases, new models can possible do better from a blank slate. But in existing codebases, I feel like the agents are more likely to simply follow existing patterns and existing guidance to begin with so things are a wash and more reliant on harness and existing code hygiene.
dom96 1 days ago [-]
Surprisingly (or maybe not) it matches the performance of Astra on my benchmark[1], but is much cheaper. It is also head to head with Opus 5.5 on both the price and pass rate, but edges it out slightly.
Your benchmarks are really weird though. Opus 5.5 claimed to be trash, Gemini Flash claimed to be the best, Sol 6 claimed to be better than Astra 6? How are you testing? Your findings don't match anyone else's.
XCSme 21 hours ago [-]
I am also surprised by some results, but I also don't favour any model, so all models are tested exactly the same, and some simply fail some tests: incorect answer, requests timing out, not respecting output format requirements, writing code that doesn't generate the correct response, etc.
The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.
XCSme 21 hours ago [-]
The benchmark is a bit sarurated at the top, there it's more relevant to see cost/response times/efficiency.
Opus 5.5 is not trash, still in the top models, but Anthropic models always have struggled with instructions following and refusing to answer questions. That being said, most tests are basically questions or simple tasks, and models could either do them ok or not. Nowadays the models and how good they are in practice is given more by the harness, than the model itself. I do think I should probably find a way to test the models including their harness, and to do so for a complex, long-running task, the so called "agentic" use-case.
Also, Gemini models have the best all-around knowledge, they are way above other models in general knowledge and domain-specific knowledge. No tests have web search enabled, and most modern models are indeed optimized for that use-case nowadays.
tl;dr: the leaderboard simply shows, given any simple question or programming task, which model is most likely to get the answer right.
killerstorm 11 hours ago [-]
I got some extremely long responses from 'ChatGPT' yesterday - e.g. 10+ pages long essay whereas other models reply with 1-2 pages to the same question.
I wonder if it's a regression of 6.x Sol. They removed model indicator from 'chat' section, so it's now mystery meat model.
Also, I'm generally very angry that we don't have "this response was generated by XXX" for each response because that's pretty damn important.
seaal 1 days ago [-]
Just got access in Codex, looking forward to trying it out. Opus 5.5 has blown me away with what it's capable of doing, hopefully 6.1 will actually be a worthwhile contender.
Excited to tryout Decisions API as well.
moinism 1 days ago [-]
Ok, but we need 6.1 Luna soon. 6 feels worse than 5.6 in our agentic use case.
lampcord 12 hours ago [-]
If the performance holds up, that price point is a game-changer. My current dev environment costs are getting ridiculous.
lukehandcool 1 days ago [-]
What happened to "we urge you to urge us to stop moving AI so fast"?
Aboutplants 1 days ago [-]
So when does Anthropic answer? Tomorrow?
iosjunkie 1 days ago [-]
hopefully the answer doesn't include an increase in cost/decrease in usage.
Alifatisk 1 days ago [-]
You live in a ping-pong.
t-sauer 1 days ago [-]
Wasn't 6 released like last week? I can't keep up anymore.
algoth1 1 days ago [-]
It was so underwhelming that it didn't even make it to chatgpt chat interface
oh_no 1 days ago [-]
Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
RugnirViking 10 hours ago [-]
we have it in our enterprise. The admins have to individually approve new model releases. Ours don't enable astra :(
nsingh2 1 days ago [-]
Yes but it was underwhelming, so they seem to have rushed 6.1 Sol out. Also Opus 5.5 may have spooked them too.
SirMaster 1 days ago [-]
Do you need to? Do you always keep up with all the version bumps on the software you use?
samayashar 11 hours ago [-]
The Sol series might be OpenAI's best release for everyday tasks. The model provides great all-round performance at a nominal cost.
I've been actively using it since the past couple of months and it has rarely disappointed me.
hlynurd 1 days ago [-]
Weren't there headlines just yesterday that they weren't releasing this due to safety concerns?
pkulak 1 days ago [-]
That was 6.1 Astra. And I'm assuming it's being tabled because it still doesn't match Opus 5.5.
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
cmrdporcupine 1 days ago [-]
6 Sol was worse than 5.6 Sol from my own experiences. Far worse.
Will see if this remedies things.
pkulak 1 days ago [-]
Yup, same experience here. I used it for one day, spent the next day fixing its lousy code, then went back to 5.6.
tedsanders 1 days ago [-]
GPT-6.1 Astra is what those headlines referred to. This is GPT-6.1 Sol.
Are most comments on HN now just astroturfed $LLM_COMPANY promotions?
ChromeUltron 22 hours ago [-]
yes
rldjbpin 11 hours ago [-]
this release could be an email or a website post. they happen to do a devday so i understand the formal announcement i suppose.
meanwhile the downgrade to the 200 plan, while not as outrageous as github copilot, is just not good pr fwiw. the base number, i.e. what plus users get, is arbitrary as it is, so there was no need to break the appearance of "20x". using the efficiency excuse is not helping here lol. same as further segmentation of pro plans with ultrafast mode.
baxtiyor7 16 hours ago [-]
Didn't Anthropic and OpenAI agree on slowing down? But they jumped into releasing new models every week
sakompella 16 hours ago [-]
"pacing the frontier" is about the frontier in fairness. Astra, Fable level models. Sol / Opus / Sonnet etc are all worse than "the frontier".
16 hours ago [-]
dzogchen 1 days ago [-]
I feel a little salty about the plan changes. I wanted to upgrade to the $200 plan a day after it was blocked. Now it only includes half the usage unless for those that got grandfathered into the x20 usage.
InsideOutSanta 1 days ago [-]
Grandfathered for a whole month. You're not missing much.
SirMaster 1 days ago [-]
Guys, are we slowing down yet?
scottyah 1 days ago [-]
Yes, obviously. They're both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.
coopykins 15 hours ago [-]
I had the feeling that GPT 6 Sol was actually GTP 6 Terra, it didn't felt as good as Sol, and the benchmarks also showed it was a bit worse than the previous Sol release.
Now this feels like its the actual Sol 6 that was meant to be.
cregy 9 hours ago [-]
$10 output is a good sweet spot. Use that for problem solving and glm5.3 (at 40¢) for bulk work
xkcd-sucks 1 days ago [-]
> Not sure what kind of usage can justify $200/month of either openai or anthropic, i'm not even talking about $500
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
holbrad 24 hours ago [-]
It seems pretty clear that this is a much larger model than Sol 6, and you can see this in the much lower generation times. I think this is also the main explanation for the $200 plan being cut in terms of API usage.
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
aetherspawn 23 hours ago [-]
Astra requires multiple turns and fresh refactoring agents to produce good code.
Fable 5.1/Opus 5.5 isn’t different, but the first cut is better quality.
Astra is a whole order of magnitude cheaper than Fable, and the Anthropic usage limits are ridiculous. Layers on layers of limits that constantly trip.
We don’t really use Sol because Astra X High is cheap. Some have mentioned regressions but we haven’t noticed any with Astra.
alvis 1 days ago [-]
Cache is priced at $0.1/M, 50% as sol 6 and sonnet 5.5.
arctic-true 1 days ago [-]
It’ll be interesting to see what happens to the economics of this business if we hit a wall on peak intelligence but keep finding cool ways to lower prices.
RugnirViking 10 hours ago [-]
that sounds good for literally everyone that isn't an anthropic/openai shareholder
resters 24 hours ago [-]
Frontier models are being used to obtain training data from users. We burn tokens teaching OpenAI how to make a cheaper model that is almost as good. I think the new $500/month pricing strategy is a significant misstep by someone who has clearly not tried Gemini 3.8 Flash or Deepseek 4.1 Flash.
killingtime74 16 hours ago [-]
Have you heard of zero data retention?
resters 5 hours ago [-]
Not everyone opts for that.
KingOfMyRoom 1 days ago [-]
The main issue I have is how they nerf their models and the quality difference between API users and their subscribers.
medler 1 days ago [-]
What is the quality difference between the API and subscriptions?
najmbajwa123 14 hours ago [-]
Has anyone measured their real hit rate in agent loops? The API reports it on every call cached_tokens under input_token_details.
kenzic 1 days ago [-]
Is it really 1/5 of the price if most people who use it are also losing 1/2 of their credits?
thefounder 1 days ago [-]
They need to fix Astra first. My main issue is with GPT in general is that unless steered it goes into building AI “sloppiness”/machinery that is not “needed”.
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
23 hours ago [-]
diego_sandoval 23 hours ago [-]
I find GPT 6 to be lacking in common sense when it comes to interpreting my prompts.
I have to be more literal with it than with GPT 5.x, otherwise, it sometimes does something totally different than what I want.
absoluteunit1 21 hours ago [-]
I am convinced the comments after a every model release announcement is just paid posts shilling whatever model provider paid them lol
nicce 1 days ago [-]
It is hard to trust these scores. GPP 6 Sol has been so bad for few days.
mekpro 1 days ago [-]
Why they are not even benchmark model against Anthropic or anybody ?
1 days ago [-]
meowers1 22 hours ago [-]
Crazy fast turnaround from being on top of the game with Astra release to being so far behind almost instantly.
thimble_io 11 hours ago [-]
Near-Astra intelligence? More like near-miss on JSON schema output. Still, that price makes it hard to ignore for hobby projects.
gcanyon 20 hours ago [-]
It has to be the case that OpenAI, Anthropic, and Google are doing internally what they accuse others of doing: using expensive advanced versions of their models to train lighter-weight more efficient models. Have they explicitly acknowledged that Astra was used to train Sol?
tensegrist 1 days ago [-]
noting that input:output:cached is 10:50:1 for astra and luna but 10:50:0.5 for sol. this doesn't mean a lot without "tokens per task" information but it's still interesting for there to be a "dip" like this instead of a monotonic change in one direction or the other
dangoodmanUT 1 days ago [-]
I can appreciate they show that Opus 5.5 is objectively better intel. and cost wise on multiple benchmarks
1 days ago [-]
AmazingTurtle 1 days ago [-]
regarding the new pro max subscription btw:
it's 500$ for 25x the plus usage, thats pro (max).
this implies that the old 200$ 20x pro (more) is now more like 10x the usage of plus.
they are slashing our subscriptions in half and make it "but we're more efficient!"
user43928 1 days ago [-]
Which they are, so another way to see it would be that they make the API/enterprise cheaper.
However, I don't know how future larger models such as the cancelled 6.1 Astra will be priced.
If the price stays high, this would indeed be quite bad for the $200 subscription..
Lapalux 1 days ago [-]
It's relentless, isn't it?
m4rtink 1 days ago [-]
Price wars, those always end well, especially if you are loosing billions per day!
zf00002 1 days ago [-]
I can blow through my weekly on astra in a few hours; hopefully this really is as good.
skybrian 1 days ago [-]
I get decent results telling it to use Luna subagents for implementing commits.
unsupp0rted 1 days ago [-]
Excited for Gemini 4 at equal or better coding and a fifth of the price of 6.1 Sol
oh_no 1 days ago [-]
3.8 flash is more expensive than sol 6.0, google uses a LOT of reasoning tokens
demibabs 1 days ago [-]
Is OpenAI’s Cloudflare turnstile new? Never gotten it before.
Either way, a little ironic…
MisterMunchkin 1 days ago [-]
Still too expensive. Needs to come down to $1 or less to compete with china.
AspireOne 23 hours ago [-]
It did. And 6.1 now outcompetes China in most cases.
Chinese models are cheap and fast by token, but they generate oceans of thinking tokens in order to accomplish the same result GPT-6.1-Sol accomplishes in, comparatively, two drops of thinking tokens, making the Chinese models come out costlier and slower per actual task performed end-to-end.
API pricing. With Sol having SUBSTANTIALLY better performance.
sanex 1 days ago [-]
If it's so good why is Dots, also released today, based on Astra?
ChaseRensberger 1 days ago [-]
very surprised by the sentiment against GPT 6.0 Sol, I've been using it exclusively since release and it feels like a cheaper astra to me. admittedly I haven't tried any anthropic models in a while other than small tests since i can't use my anthropic subscription in other harnesses (like OpenAI has supported natively for a long time).
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
unsupp0rted 1 days ago [-]
This is the first time I've seen praise for GPT 6.0 Sol: it's widely disparaged on Reddit and here in the HN comments too. My own experience likewise shows 6.0 making loads of silly mistakes, both for things 5.6 Sol is good at and things 5.6 Luna Xhigh is good at.
ChaseRensberger 1 days ago [-]
well i could certainly be in the wrong; i'm just speaking from my personal and likely flawed experience but i feel like i've noticed silly mistakes in every (llm) model that has been released (and that i've sufficiently used) and it hasn't felt like 6 Sol was much of a regression from 6 Astra (more than reported in both model cards), both of which ive very extensively.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
mpweiher 14 hours ago [-]
Hmmm...is it just me or are pretty much all of the recent announcements from the AI companies about lower prices for comparable performance?
So we've already reached the commodification phase, due to either the frontier models not getting usefully better or the models already being good enough, or both.
In either case, there appears to be a plateau in performance. With local open-source models catching up.
Interesting times.
nilslindemann 1 days ago [-]
Thanks, but I need the money for the next laptop.
gitghxst 11 hours ago [-]
excited to read some reviews about it
gavin_gee 1 days ago [-]
the race to the bottom on models is well underway. huge IPO's only really make sense for DC/HW lockups, and going vertical.
christkv 10 hours ago [-]
They just informed me they are gimping my 200 plan to 10x instead of 20x the normal plan size and threw me some extra tokens to compensate. So the gimping is starting the new 200 plan will be 500.
cannonpalms 1 days ago [-]
OpenAI is cutting their subscription's token value in half. Half. I don't think this is enough to help them compete with Opus 5.5.
jrflo 1 days ago [-]
Just for the $200 tier, which was formerly the 20x weekly usage of the $20 tier. Now it's 10x the weekly usage, just like Anthropic's $200 tier.
aaronbrethorst 1 days ago [-]
Shots fired, half the price of Opus 5.5.
barrenko 1 days ago [-]
A night of no sleep this one will be.
ghm2180 1 days ago [-]
> At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
par 1 days ago [-]
Ok when can i get this in codex?
harlan_pdx 11 hours ago [-]
Another week, another "near-Astra" claim. Need to see some real-world performance metrics before that price makes sense.
1 days ago [-]
epolanski 1 days ago [-]
DeepSeek and GLM made it impossible for "sota" to price any way they want.
israrkhan 18 hours ago [-]
so this is what pacing down looks like
tultra 1 days ago [-]
Still not available to me
itzikkatz 1 days ago [-]
I’m a bit disappointed with Sol 6.1. I suspect they didn't show the benchmarks and test results because it would have been embarrassing to reveal that their flagship model can't compete with the capabilities of Sonnet 5.5. That said, I still think this model is useful for a great many things, but it looks like Anthropic has the upper hand this time.
cmiles8 1 days ago [-]
The only real question that matters at this point is what do the economics look like for OpenAI. If they make good money on this with sane accounting principles then great. If this is just throwing more gasoline on the pile of burning cash to avoid losing more inference business then this bubble can’t pop soon enough.
throwitaway222 1 days ago [-]
So yesterday we were consumed with how this was being delayed because of safety, yada yada.
Guess not?
nimonian 1 days ago [-]
I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don't think this release in particular contradicts it too much.
minimaxir 1 days ago [-]
That model was implied to be GPT 6.1 Astra, not Sol.
murbard2 1 days ago [-]
Astra 6.1, this is Sol 6.1
anotha_one 1 days ago [-]
[dead]
1 days ago [-]
1 days ago [-]
Alifatisk 1 days ago [-]
I guess they released this because GPT-6 Sol was underwhelming, they didn't even release it to ChatGPT. It was basically GPT-5.6 Terra for the price of Sol. However, who doesn't like price cuts? Astra for the fifth of the price? Wow, OpenAI have been quite generous recently, I still have not forgotten their 90% price cut with GPT-5.6 Luna, and now this? Astra was truly a milestone, and now they are offering similar "intelligence" for cheaper price. Incredible.
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
andsoitis 1 days ago [-]
Sol, Astra.
Eventually: black hole.
MetaverseClub 1 days ago [-]
Sam Assman needs money to buy a new super car or private jet?
hacker_88 1 days ago [-]
AGI is here
EugeneOZ 1 days ago [-]
Can't trust a company which can halve a subscription any moment they want.
1 days ago [-]
24 hours ago [-]
spwa4 1 days ago [-]
I was wondering why GPT-6 astra has been performing so incredibly bad on codex for the last week. This seems to be a repeating pattern, to dial the settings on the current models to the idiot setting, and then release a new model about a week later.
varispeed 1 days ago [-]
I wonder if this is the reason GPT barely works now and some chats have become not accessible.
sergiotapia 1 days ago [-]
Typically how long does codex take to update with the right model metadata for the release of a new model?
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
IshKebab 1 days ago [-]
I wish they'd list the environmental cost. My employer has an unlimited AI budget so I don't care about using Astra if it's just more profit for OpenAI. I care more if it actually uses 5x more energy.
otterley 1 days ago [-]
Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.
lp92 1 days ago [-]
Did you care about this when it came to your other computing needs? What PC/laptop are ypu running and how efficient is that?
IshKebab 1 days ago [-]
Yes I do. I've got a spare desktop that isn't too efficient (probably ~100W idle but annoyingly I've lost my power meter) so I don't leave it on even though I would like to use it as a server.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
codehorses 1 days ago [-]
Likely orders of magnitude different, this is a weak whataboutism.
Bolwin 1 days ago [-]
Look at Neuralwatt. They report energy usage with every call as well as aggregate statistics.
sergiotapia 1 days ago [-]
I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar/battery/nuclear.
rs_rs_rs_rs_rs 1 days ago [-]
I don't understand the point of this, why just now when it comes to llms. Why wasn't anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.
Edit: for context, just Steam alone has ~200million monthly active users.
JDups 1 days ago [-]
Because people find video games fun, though I suppose there's some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
paulryanrogers 1 days ago [-]
Considering Nvidia's hard shift to crypto and now AI, I doubt videogames are even in the same ballpark.
How many DCs are devoted solely to gaming?
lp92 1 days ago [-]
Considering many games make use of cloud computing for online play and similar functions, they probably make up a pretty goot bit of global cloud compute capacity. Likely quite a lot less than the big AI players, but not an insignificant amount.
HelloMcFly 1 days ago [-]
The energy costs of the cloud computing required for gaming are substantially less in power - not to mention overall demand - than LLMs. Come on, we're not in the same energy ballpark here.
empthought 1 days ago [-]
You're not including the physical supply chain energy consumption of distributing video game equipment in this analysis. Nobody ships LLMs to big box stores and tries to sell them to consumers.
paulryanrogers 21 hours ago [-]
Historically all those boxes and discs were significant, at least if you sum them all for all time.
Today physical discs are rapidly becoming a niche that newer consoles just won't have. PC games are almost never sold physically anymore.
rs_rs_rs_rs_rs 1 days ago [-]
>The energy costs of the cloud computing required for gaming are substantially less in power
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
paulryanrogers 21 hours ago [-]
Steam games can run on less than 4GB of RAM, and per person are likely played only an hour or so a day. AI hyper scalars use far more to serve a single customer.
1 days ago [-]
rs_rs_rs_rs_rs 1 days ago [-]
>How many DCs are devoted solely to gaming?
An entire planet. Just Steam alone has one or two hundres million monthly active users.
paulryanrogers 21 hours ago [-]
The whole planet's DCs are not solely serving games. Apparently games take about 300 TWh whilst AI is around 500 TWh.
All of gaming is in the 300B USD range while just the CapEx of hyper scalars is already over 200B.
lbrito 1 days ago [-]
If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.
rs_rs_rs_rs_rs 1 days ago [-]
> you will remember that people have indeed been enraged with the environmental costs of things for a long time
Yeah? Show me the big movements against computer gaming.
lbrito 1 days ago [-]
There are movements against consumerism and the environmental impacts of industry in general. Greenpeace is over half a century old.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
IshKebab 1 days ago [-]
I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
rs_rs_rs_rs_rs 1 days ago [-]
> non-AI things
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
IshKebab 1 days ago [-]
> But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
I dunno what you're not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.
Also gamers tend not to spin up dozens of other machines to also game for them.
prometheus1992 1 days ago [-]
where's the pelican?
zerof1l 1 days ago [-]
Am I the only one who is struggling to keep up with all these GPT model versions and which one to use and when?
joduplessis 1 days ago [-]
Sol is so good, honestly - the sweet spot for me. I've only ever found it stumbles when you don't give enough direction. But for idea execution - Sol is the GOAT.
jennnnx 21 hours ago [-]
I'm about to give up on OpenAI. I can't keep up with all these version numbers and sub-names like Sol, Terra, Astra, etc. Wasn't 6 released very recently? What the heck is the difference between 5.6, 6, 6.1, etc? Is 5.6 Sol better or worse than 6 Luna?
formvoltron 1 days ago [-]
didn't 6 sol just come out a couple weeks ago?
Alifatisk 1 days ago [-]
Other comments have already addressed this.
lynx97 1 days ago [-]
Came here only to check if the pelican spam has made it to the top again.
slopinthebag 1 days ago [-]
$2/10 is pretty cheap for a frontier model...
dcchambers 1 days ago [-]
GPT-6 Sol released a week ago. Shortest model life ever?
cmrdporcupine 1 days ago [-]
Taking GPT-6 "Sol" outside behind the shed and giving it a merciful end is about the best outcome possible.
Huge misstep releasing it.
tandr 1 days ago [-]
The misstep was naming it Sol - it was Terra-level all along.
algoth1 1 days ago [-]
Bro, I'm a visual thinker
cmrdporcupine 1 days ago [-]
Sorry, I'm GenX. Growing up they showed us "Old Yeller" in the school gym every year like that was some kind of treat.
cmrdporcupine 1 days ago [-]
Aka "we made an oopsie last week and released what should have been called GPT 6 Terra with the name GPT 6 Sol"
jdw64 1 days ago [-]
GPT 6.0 Sol was so terrible—I wonder if 6.1 Sol will be good?
gxs 1 days ago [-]
I have used both extensively and I don’t care what the bench marks say - Claude has been way, way better and most importantly, predictable
I can handle issues much better if they are predictable even if the model makes mistakes — much more frustrating when the model is erratic
I find codex wanders off road more often and fails to see the “bigger picture” (as much as LLMs can see the bigger picture at least)
And tbh when it was first released Astral felt even worse
I’m being forced to use it right now and at the end of the day I’m making do so it’s fine, but Claude makes for a smoother experience
sehw 1 days ago [-]
I stopped using LLMs. I shit you not. My life got better.
Starlevel004 1 days ago [-]
Okay, now price cut 6 Sol (and rename it to Terra again).
dools 24 hours ago [-]
All of a sudden getting competitive on token pricing over the past couple of releases tells me they’re about to kill subscription pricing big time. The subsidised tokens aren’t going to survive the IPOs but if they can capture baseline dev tasks at a cost competitive with open weight models through Luna then capture the frontier token spend as well they could be pretty well placed. The Jarvis bros aren’t going to be able to afford their dashboards though.
aitoolcrux 9 hours ago [-]
[flagged]
ariwilson 1 days ago [-]
[dead]
nicolamanzini 1 days ago [-]
[dead]
jp1016 1 days ago [-]
[dead]
1 days ago [-]
FpUser 22 hours ago [-]
[dead]
amelius 1 days ago [-]
If these models are so smart, can't _they_ select the right model for each task?
skulk 1 days ago [-]
the right model for the task is the one that transfers the maximum amount of USD from your pocket to the provider's bank account.
amelius 1 days ago [-]
No because then I'll go to the competition.
mholm 1 days ago [-]
Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.
I guess the question is, does the Dunning Krueger effect apply to models? The dumb ones might think they're up to the task.
aleph_minus_one 1 days ago [-]
Why don't you simply ask the respective model which model is best for a specific task? :-)
amelius 1 days ago [-]
Because it is more work?
SkyBelow 1 days ago [-]
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
godwinson__4-8 1 days ago [-]
Let's all boycott and move to Claude until they release 6.1 Astra. I don't like to be teased.
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
ColonelPhantom 1 days ago [-]
Isn't Anthropic doing the same, with Opus 5.5 being out while Fable/Mythos is still on 5.1?
wren6991 1 days ago [-]
It's just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
godwinson__4-8 1 days ago [-]
Is this due to a similar safety concern or just because it's not ready yet for one (or more) of a myriad of possible reasons?
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
Rendered at 21:09:53 GMT+0000 (UTC) with Wasmer Edge.
Data at https://gertlabs.com/rankings
This is a plus in my opinion.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?
1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?
How is 2 better in any way for anyone involved? Why would you, as a paying customer, holding it wrong, want other people to keep that information from you?
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
https://community.openai.com/t/why-does-codex-report-a-258-4...
yes, exactly
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
I barely compact at work in a very complex monorepo (neither with Fable 5.1 nor Opus 5.5), and yet in my personal greenfield project Astra keeps compacting all the time, to the point of it being unusable.
~/.codex/config.toml
I’ve also found compaction not to be a problem when it does happen.
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I've always seen them be prompted, but this is something high priority on my list of unknowns to understand better...
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.
Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
However it's less willing to obey your instruction so it's less usable for general runtine flows.
Looks like 5.5 is the new 4.6
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
Going on a slight tangent, I find that I get the best results when I force Codex models (Astra/Sol) and Claude models (Opus/Fable) to consult each other (just have them build a simple skill). There are tasks that neither one can fully solve on their own, but their differences are large enough to make a difference when they collaborate.
My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now, but either way having 2 "smart" models available from 2 different companies is a boon.
Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
1 - https://bench.killswitch-lang.org
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.
Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.
I have seen you make these immature accusations that users here are OpenAI employees multiple times today.
This is not even remotely true.
If you're running 2-3 parallel agent session with a few sub agents and waiting for you to prompt them, you'll have a very different experience!
This is a tiny minority of people, not “most people”.
(Usage limits are entirely dependent on what you're doing with them. If you're not running it on 1 million LoC codebases you can get a lot of mileage out of even a 5x account particularly with the recent cheap models)
My main worry is that we get to a point where they have something much, much smarter than anything public and access is gated by extraordinarily high costs.
That probably can't happen though right? Inference is surprisingly cheap compared to the training.
This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.
It's also used to describe the SFT bootstrapping for posttraining, which is what people generally refer to as Chinese labs "distilling".
I would almost guarantee that smaller US frontier models (ex Luna/Sonnet) are distilled from their respective large models.
Have they, actually? A lots of speculation and claims but what is the level of admittance?
- Training data, harnesses, and processes, that drives the quality of the models. Several companies have those. Some of those companies are in China.
- Infrastructure and funding for running those. That's a scarce commodity currently, training the latest frontier models cost billions of dollars apparently. And if you don't have the infrastructure already and don't have the suppliers on speed dial, good luck getting anything.
- Infrastructure for running inference for running what comes out of those. Several of the key providers of this infrastructure are using their own in house chips for this now. At scale this means huge cost savings.
If you start from scratch without infrastructure, there are a bunch of open source things you can find. But beyond that, you'll have a lot of catching up to do. That's the very definition of a moat.
Infrastructure is readily available. Anthropic and OpenAI don’t own anything, they’re just paying for compute. You might not be able to buy 100k GPUs right now but you can rent it.
You only need to look at how Jev immediately became one of the highest volume models after launch on OpenRouter to see how this industry is moatless.
OpenAI and Anthropic employees frequently leave to start their own labs and raise hundreds of millions for it which they can use to immediately pay for infrastructure and training data. At most OpenAI and Anthropic have… brand and talent.
If anything, OpenAI and Anthropic are heavily disadvantaged because they have huge long term financial commitments that have backed them into a corner, likewise the regulatory pressure… startups have none of that.
And I'm going to keep telling people that it is not unreasonable for individuals to spend $30-40K on an AI rig if they are as productive and useful as many people imagine, my friends and family can take advantage of my rig when I'm not using it, and I'm independent of some snooping corporation - totally private.
There is little reason for these things to be concentrated in far away data centers other than that pooled usage is more elastic (but if I'm sharing my rig with many other people, AI is agentic now, and I can lay off some of my compute to others and they can lay it off on me I'm pretty elastic on my own.) The real reason is because we moved to a rentier economy before AI with the cloud and locked-down phones, and it is a very profitable model that doesn't care if it's rational.
You could subscribe to training material like you would subscribe to a newspaper. The most important training material would be your system learning off you, so training will have to be done in place (we can't rely forever on injecting specifics into contexts.)
These things should be in people's houses, and in the winter you should be using them as radiators. Don't think the power grid can handle it? Put solar on the roof and have batteries at home; helps both you and the grid.
This depends on what governments do. Crony capitalism in the US can lead to bans on home AI even if they have to root all general purpose computers to do it. The same is inevitable in Europe due to both total elite disdain of democracy and their slavish service to US interests.
Even if Western governments just lock down and antitrust-exempt the frontier labs without rooting personal computers in order to keep out Chinese/open models, I'm in far from a comfortable position when I'm depending on the Chinese government to protect my freedom.
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
What distinction are you drawing?
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
do you think it will be exponential forever?
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)
What a time to be alive.
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
Replacing white collar work would be worth dozens of trillions of dollars at minimum. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B Annual Revenue Run Rate in 200 days. And Anthropic is growing so fast they're on pace to hit $100B in Annual Revenue.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic and OpenAI are still growing enterprise usage heavily. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit about what qwen does. And specialized models often perform worse than generalized ones.
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
Edit: removed a comment that was uncharitable and rude, for which I apologize.
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
We've gone from 80% in some places to 80% in some more places.
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
Have you used recent ones on any large projects or coding issues? They've improved tremendously lately in my experience, in terms of implementing non-trivial things, debugging, handling legacy/complex codebases, etc. I usually keep a text document of things that models failed to do/fix and recent ones have wiped it all away.
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
People were talking about plateau for years already.
Meanwhile, in the US, companies like Boeing, GM, and Intel will never be allowed to experience more than minor financial inconvenience before the government bails them out with protectionism, guaranteed loans, and outright subsidies.
I don't see a material difference between how the Chinese government treats their strategically-important industries and the way we do here in the West. Terms like "socialism," "Communism," and "capitalism" are just fodder for Fox News camp-followers.
It’s not happening and the imminent bust is coming. Strap in while the music gets turned up (dots, ipo etc) and people decide to leave the partayyy!
It's not even anything controversial..
it's also why there have been so many calls for regulation and slowdowns.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
I remember when bandwidth was super expensive and now it’s dirt cheap.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
I'll just leave this here: https://marginlab.ai/trackers/codex/
I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).
Then when opus 5.5 came out, again same thing + far cheaper and faster.
I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.
I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.
Anthropic models are very thoughtful and have so much taste all around.
And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.
FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.
One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
I haven't used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it's not just a tune.
Astra is clearly far better than anything prior though, so I'm not sure what you mean really.
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.
I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.
I've never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
But the origin is size inflation. The original Starbucks sizes were short and tall. Shorts are still available, in at least a few outlets, despite having not been on the menu for at least 20 years. (For a while a few years ago here in the UK, I even saw short as a column on the menu at one, but that's very rare.) Then they added grande, literally meaning large but maybe meaning more like "extra large". Then when they added an even bigger size, since large was already taken, they just stated the actual size "venti" = 20 (fluid ounces).
(The first time I ever went to a Starbucks had the same confusion as you. I didn't know what an espresso was but at least its sizes of "single" and "double" were easier to understand, so I got a double and was a bit surprised at how small my drink was!)
Like who can figure out the ordering? With Luna < Terra < Sol < Astra it's obvious at a glance.
I propose for some third company to name after monsters: Cyclops < Minotaur < Ettin < Cerberus < Hydra < Kraken < Nyarlathotep (the AGI singularity stage)
It.. is?
When was the last time you heard anybody say "opus" or "sonnet" before this?
Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
Fable kind of screwed that up, though.
Not a lot, but occasionally. I was familiar with the words.
> Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
It really doesn't.
The one real weakness of Anthropic's naming is that Fable doesn't fit. A fable is typically a fairly short story whereas I think of an opus (if a written work) as being at least a long book and maybe a whole collection of books.
Shall I compare thee to a summer’s day?
You are serious?
terra, sol == luna, astra
I feel like people who don't get it immediately are just being deliberately obtuse.
The more obvious, to me, interpretation was distance from us so I would've expected Terra -> Luna -> Sol -> Astra.
When I think of space and astronomical scales, I think of distance as much more common concern.
Okay. I don't believe that you got it immediately.
Many of which are thiccer than our beta ass sun
Pretty badass name tbh
clear would be something like
piss-cheap - it’s-alright-i guess - okay-relax - ouch-my-wallet
Honestly, the old boring business lingo of adding Pro and Max suffixes would have been clearer for both companies.
Also, cyclops is assuredly bigger than minotaur in almost any reference. And what's "ettin"?
I don't think it gets better if you swap "fable" for "myth" either, although when you add "-os" to the end it suddenly sounds grander so there's that. Mythos isn't a model I can use, though.
Then they could probably move to proper nouns: Sagittarius, Andromeda..
Bug < Ticket < Story < Epic < Theme
Not sure that's why they did it. But that was my experience.
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
It's not frontier pelican without the back leg behind the bike frame IMO.
GPT 6.1 Sol — 91, ~9 min, $0.51 https://jonclegg.github.io/pacman-bakeoff/#gpt-6.1-sol
Opus 5.5 — 99, ~9 min, $2.00 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
GPT 6 Astra — 87, ~10 min, $2.42 https://jonclegg.github.io/pacman-bakeoff/#gpt-6-astra
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
Full gallery: https://jonclegg.github.io/pacman-bakeoff/
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
Our VC-backed subscription days are numbered
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
$500 for the old $200 is definitely a fumble.
Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.
There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.
The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.
RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.
I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.
> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
Now you can use it in coding harnesses that call the API.
Why? You can use in codex, right?
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
The vast majority of their revenue comes from large businesses buying for their teams, which are almost certainly not going to juggle lower tiers of different subscriptions to save a few bucks.
Being grandfathered by OAI and happy is not the same as having both, and noticing "hmm maybe Claude is much better for my case, Ill suggest that to our manager"
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
Qwen3.8-Flash-Next is better still if you can run fast enough. If you have a dual R9700 setup you certainly can. That model is even better.
Qwen4-27B has been announced but not released yet. I'm super pumped for it because I already use 3.8 as my daily driver at home for all my personal stuff, so I'm definitely happy to take an increase in capability.
There is clearly still room for improvement in local models on consumer hardware. With the Qwen 27B models, If you have at least a 5070 Ti I think you can get away with running a small Q4 quant if you use KV cache streaming. The 24GB cards can run Q4 comfortably. If you have a 32B card you can run Q6 comfortably. If you have 48GB ~ 64GB of VRAM you can Q8 comfortably. Using llama.cpp Vulkan let's you pool VRAM across cards (even AMD and NVIDIA etc), so my machine has a 5060 Ti and an R9700.
A dual R9700 rig is really the sweet spot right now with the vLLM-radiance fork. If you can swing a 5070 Ti in there as well to retain some CUDA access, then all the better. That's basically the equivalent to spending 2 years on a subscription, but gets you a system that can run Qwen-3.8-Flash-Next and of course the even more capable Qwen4-Flash when it releases. At the end of the two years it'll run even better models I'm sure.
I'm all in on local now.
I can justify $200/mo but more than double is not appealing to me.
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
I think this is still true provided you're not using Astra.
You have to do a lot of things in parallel.
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
The only way is for prices to go up. Way up.
The cost of providing the tokens for a heavy user (and let's be frank, the people paying $200 are likely heavy users) is many, many times more than the $200 recurring revenue they generate.
Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
You can actually check the approx. financials of OpenAI and Anthropic. The growth is insane.
There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
> Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
DeepSeek increased prices substantially not long ago. I find their profit margins hard to inspect considering I have very little idea what sort of environment they may get in China (from cheaper energy to government subsidies). I honestly doubt you have any insight here as well.
> You can actually check the approx. financials of OpenAI and Anthropic.
No you can't. They are not publicly traded, and they constantly and selectively leak bullshit metrics, from extremely unclear ARR, to extremely deceiving EBITDA. You willingly eat their bullshit and call me a picky eater in return.
> There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
I see no reason to believe (much less any actual evidence) that OAI or Anthropic have any path to profitability.
If inference (particularly for subscriptions) was in anyway as profitable as you claim today, they wouldn't need private investment rounds like crazy nor they would be desperate to offload this hot potato in an IPO.
82% margins lol. Are you telling me that if you created a machine that turns 1 dollar in 5 what you would do is dillute your ownership of the machine instead of using these fabulous profits to expand the business?
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
tbf, i barely feel the difference with opus 5.5 anymore either.
https://openrouter.ai/docs/guides/routing/provider-selection
I stopped using opencode because it has some issues with caching, so I suppose it doesn't do this
Using this interesting framework called Cordis I've recent discovered.
Edit: others have noted the provider and harness matters. My experience is with opencode.
What harness you are using?
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.
Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.
What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.
If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.
So, spam?
If someone is asking you for a quote, and your tool replies, it's not spam.
It might be slop for all I know (or might not be!), but it's not spam.
Similarly for arranging meetings with parties that you already have a relationship with.
I can’t overstate how bad of an idea I think using an AI for customer interaction is.
But yeah, send me a non solicited AI slop email or worse, political ad, and you dont get the dignity of me saying stop to unsubscribe. Straight to spam for you.
Sorry, your comment is just far out of touch with reality.
> First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. [...] Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible.
You seem to think raising productivity is a bad thing?
> Humanities will stake it out a little bit longer because AI isn’t human, [...]
This is really silly. Many languages other than English don't use related words to describe 'humans' and 'humanities'. Will it be easier in those languages? Should we rename mathematics to 'humathematics' or so, to make it harder for AI to take over?
Higher productivity means very little to people whose labor plummets in value over the course of a couple of years. One day they won't be needed anymore, maybe that's good for you.
I'm a software developer by trade. I welcome the coming brave new world in which machines can do all the software.
At the moment, they ain't quite there yet, alas.
Also even if your Gemini is giving you nonprogramming output, underneath the model is most likely generating code for certain tasks.
I dont consider myself a programmer but use LLMs almost exclusively for coding.
The number of people able to create useful software today is much much larger than it used to be and arguably a minority of these people are/were "programmers"
How is that different from feeding into the OpenAI or Anthropic machines?
Now, I have absolutely no clue how long this situation is going to last! But the economics don't really work out for local models while it does.
So if you can spend $30K and immediately start mining $1500/month out of thin air, that's a pretty nice investment even if the electricity costs $200/month. Two years later the cards will have paid for themselves entirely and (I suspect) will still be pretty useful.
It does argue in favor of just paying OpenAI or Anthropic for tokens, though.
At the same time, when I need to use the hardware for something, whoever is renting it from me at the moment is going to get unceremoniously booted, and I imagine they are not going to be happy about that. I assume that vast.ai's providers get uptime ratings that drive their work allocation, right?
And there's a lot of labour and effort involved in setting these things up and maintaining them. People don't even run their own email servers, even though the hardware side of that is trivial.
Not $400 a month, it doesn't. At least not around here (we average $0.12/kWh and I don't personally run my cards over 300W.)
And there's a lot of labour and effort involved...
Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.
People don't even run their own email servers...
People don't run their own email servers because a convenient coalition of spammers, standards bodies, and large email providers have done their best to make running one's own email server almost impossible.
I did the math and decided it’d better to pay for tokens than to buy the hardware and generate them myself.
'Capable' doesn't mean your labour has no opportunity costs.
I think my argument is easier to attack by noting that you can use AI to substitute for much of that labour.
It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.
At the moment, it's still very easy to switch from one open weights model to another, and even between the closed models. So the 'no rug pull' property is nice to have, but not as big of a deal.
Theoretically, the people who hang around HN have immense opportunity cost when doing this. Being capable does not mean it takes no time. Time they could be using for something more profitable (and fun?).
The AI models we currently have still maintain their working state with a finite and laughably-small context window, but that is already starting to change. My .claude directory contains over 300 .md files that I didn't put there myself. Their contents are eye-opening. The question of who owns, stores, maintains, and can access that data is going to become insanely important over the next couple of years.
If you thought LLMs themselves were disruptive and contentious, just wait until the fight over object permanence gets under way. That's when owning your own box full of graphics cards is going to become important. My own bet is that I won't care too much about the electric bill or my opportunity cost when we all find out what the AI labs really have in mind, and what they're going to have to do in order to justify the valuations they're seeking.
TL,DR: it's not about the tokens, IMHO.
Whats your harness?
Subagents are like trading derivatives. You can lose as much as you want.
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
Excellent pithy warning.
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
Eg when I have the AI do a self-code-review before bothering a human, I also want to the AI to give the draft PR to a sub-agent that only has the publicly available context that a viewer of the final PR would have; and not all the accumulated reasoning that lead to the writing of the code in the PR.
Opus directs and starts haiku/sonnet subagents. Much more efficient than opus reading it all.
Without this explicit direction, Opus burned through usage to create complicated regex parsers & 8 excel sheets.
And watch 10 hours of football on Sunday for our DraftKings bets.
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
1. write a sketch of a spec by hand
2. have the llm review the document and question me until it can generate a spec
3. review the spec and revise where needed
4. have it write an implementation plan
5. another round or revision/review
6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.
I've been working for a couple of hours tonight, the total cost of the session is €0.6.
it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.
The one-shot capabilities of frontier models are nice for demonstration purposes, but not actually all that useful to me - the result is often an amalgamation of ad-hoc ideas the agent came up with, poor UI and a hodge-podge of a data model, but at least demonstrating clearly where the specs are lacking.
I agree that in practice, iterative design with lots of review and hand-holding are needed to get quality results. Generated code (I mostly program C++) is often bloated and not succinct or "simple" enough to my tastes, so it requires multiple rounds of cleanup.
For core logic, I find having the agent review code is often faster and more valuable than having it implement it in the first place - it can find and fill in the things I missed.
For my own sanity it is very important that I remain in control and understand the generated code when quality and maintainability are goals of the project.
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.
Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png
I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
Deepseek 4.1: $0.02/$0.60
Just to illustrate how cheap the Corolla is in your analogy. Also Opus output would be $50 without competition.
A $20 Anthropic subscription is $500+ equivalent of API credit. You can build quite a lot on the $20 plan and get to use the best model.
Their API pricing has healthy margins built in.
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
And no it did not deliver. A lot of it was re-done by Astra
Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.
Astra was able to model low poly enemies, rig them, do simple animations, and greatly improve procedural generation. I have it modeling assets in blender every day, which I often have to go in and fix
Indeed. A few people [1] pointed on twitter that current best of Open Source Open Weight models are better than all the frontier model 6 - 8 months ago. And we heard Deepseek v5 and MiMO v3 are quite a leap.
I am not understanding the value proposition of these frontier model. Or at least not to a point they are looking at $1T or $2T market cap. Especially MiMO which may be ( correct me if I am wrong ) the most transparent model we currently have?
Or did I overlook or missing something? It is quite scary if true.
[1] https://x.com/samsaffron/status/2104903300298469877
Making more than $200/month with the results?
[1] - https://www.youtube.com/watch?v=P4dTq4X8bqk
Some people are able to take advantage of the increased capabilities that frontier models give us. For others, working with open weight models (that are extremely good in their own right) is good enough.
Deepseek models are very impressive. No one's debating that.
But frontier models are definitively ahead. Perhaps not by much in some areas, but there are clear gaps.
And that is kind of the point I think. Open source models add a much needed source of competition to the world of LLMs and are very much an important counterweight to the dominance of Anthropic and OpenAI.
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
I found that ssh is pretty bad in these scenarios, so my usual herdr over a phone doesn't always work properly.
Now I'm using paseo which in principle solves my issues properly, but unfortunately it's "reconnection" and state sync is pretty slow (probably going over their servers).
The other part is that since SSH only sends the changes on the screen, it uses very little bandwidth.
Try Eternal Terminal it's very easy to setup and is just a simple layer over SSH.
https://github.com/skorokithakis/symphony
Now I chat with OpenCode/Pi over Github issues, which is a miles better experience.
A bit of false equivalency there
It's a very noticeable hit. It's like programming with mid-2025-era models: ignored instructions, dead code, mistakes.
It works...but anyone going from Sol to Deepseek is going to have a rough transition.
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
Astra is a pretty impressive model. Excited to try this.
For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive
https://openai.com/index/introducing-gpt-6-1-sol/#deepswe
Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
It seems that 6.1 Sol is closely related to Astra. As for whatever 6.0 Sol was, it's anyone's guess. I'd speculate that 6.0 was just 5.6 with post training. It's possible however that it is the same model but with severe performance issues resolved so they bumped the version number.
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
The best thing is that we benefit from these constant back and forth gut punches :)
So I'm happy to accept your view just as easily, stranger on the internet, since it logically makes more sense.
So yes, it is clearly cheap in comparison.
GPT 6.1 Sol: https://html.non.io/lcars-gpt-6.1-sol
Opus 5.5: https://html.non.io/lcars-opus-5.5
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
I am trying to make these webpage buildouts more legitimate benchmarks though, as I do think image->html flows are going to become more popular as people realize how good images as a starting point are. Hit me up if you have any feedback!
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
[0]https://x.com/thsottiaux/status/2105007628460109953
Haiku -> Sonnet -> Opus -> Mythos/Fable
Luna -> Terra -> Sol -> Astra
Weirdly it seems like OpenAI originally just released by version numbers for their models and have switched to Anthropic's naming approach of late.
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
A few more thoughts here https://housecat.com/blog/introducing-housecat-agent
Back fired because of opus 5.5.
So now we get the real sol-6 as sol-6.1, and OpenAI will eat the cost to stay competitive.
This could be invalidated if sol-6.1 is the same speed as sol-6.
However, that doesn't say much. You can just run a smaller model at a larger batch size to get higher throughput but lower interactivity.
The last time a model announcement felt like a leap in capability beyond other things out there was Fable - which was promptly taken away. Sol and recently Opus 5.5 were strong because they approach that capability with a lot more efficiency and don't blabber incoherently (looking at you Opus 5.1).
Deepseek is a workhorse for those who prefer open and API usage. Other than that the model announcements all just seem like a blur and quite interchangeable but I wonder if that's just me tuning out or do others feel the same way?
My experience with agentic coding on projects I care about (because my responsibility in my firm is to care about these things, at least for now) has not changed a lot in the past few months, and I have kept up with every single model update / experimented with harness a great deal.
1. They don't actually look at/care what the agent is producing as long as it works (not planning on maintaining/ops yet).
2. They are using 3rd party benchmarks (which is fair given how widely real-world workloads change from day-to-day, feature-to-feature, making it difficult to really know how well the models would perform).
3. They are doing greenfield work where there is no scaffolding, no existing code, no legacy code, nothing to guide the agents along. I believe in these cases, new models can possible do better from a blank slate. But in existing codebases, I feel like the agents are more likely to simply follow existing patterns and existing guidance to begin with so things are a wash and more reliant on harness and existing code hygiene.
1 - https://bench.killswitch-lang.org/
https://artificialanalysis.ai/?models=gpt-5-6-luna-low%2Ccla...
According to this, at Max it's better and cheaper than 5.5 Medium, but worse than 5.5 High. At Medium, it's better and cheaper than 5.5 Low.
Actually, it is so close to Astra, that I'm starting to beleive it's a slightly quantized version of Astra.
[0]: https://aibenchy.com/compare/openai-gpt-6-1-sol-xhigh/openai...
The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.
Opus 5.5 is not trash, still in the top models, but Anthropic models always have struggled with instructions following and refusing to answer questions. That being said, most tests are basically questions or simple tasks, and models could either do them ok or not. Nowadays the models and how good they are in practice is given more by the harness, than the model itself. I do think I should probably find a way to test the models including their harness, and to do so for a complex, long-running task, the so called "agentic" use-case.
Also, Gemini models have the best all-around knowledge, they are way above other models in general knowledge and domain-specific knowledge. No tests have web search enabled, and most modern models are indeed optimized for that use-case nowadays.
tl;dr: the leaderboard simply shows, given any simple question or programming task, which model is most likely to get the answer right.
I wonder if it's a regression of 6.x Sol. They removed model indicator from 'chat' section, so it's now mystery meat model.
Also, I'm generally very angry that we don't have "this response was generated by XXX" for each response because that's pretty damn important.
Excited to tryout Decisions API as well.
I've been actively using it since the past couple of months and it has rarely disappointed me.
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
Will see if this remedies things.
https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
These moves all make sense when you take into account the enterprise market.
https://news.ycombinator.com/item?id=49889873
meanwhile the downgrade to the 200 plan, while not as outrageous as github copilot, is just not good pr fwiw. the base number, i.e. what plus users get, is arbitrary as it is, so there was no need to break the appearance of "20x". using the efficiency excuse is not helping here lol. same as further segmentation of pro plans with ultrafast mode.
Now this feels like its the actual Sol 6 that was meant to be.
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
Fable 5.1/Opus 5.5 isn’t different, but the first cut is better quality.
Astra is a whole order of magnitude cheaper than Fable, and the Anthropic usage limits are ridiculous. Layers on layers of limits that constantly trip.
We don’t really use Sol because Astra X High is cheap. Some have mentioned regressions but we haven’t noticed any with Astra.
The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable/Claude just cannot get/fix even when you point it.
I have to be more literal with it than with GPT 5.x, otherwise, it sometimes does something totally different than what I want.
it's 500$ for 25x the plus usage, thats pro (max).
this implies that the old 200$ 20x pro (more) is now more like 10x the usage of plus.
they are slashing our subscriptions in half and make it "but we're more efficient!"
However, I don't know how future larger models such as the cancelled 6.1 Astra will be priced.
If the price stays high, this would indeed be quite bad for the $200 subscription..
Either way, a little ironic…
Chinese models are cheap and fast by token, but they generate oceans of thinking tokens in order to accomplish the same result GPT-6.1-Sol accomplishes in, comparatively, two drops of thinking tokens, making the Chinese models come out costlier and slower per actual task performed end-to-end.
Here's from artificialanalysis.ai, per task:
* Deepseek-v4.1-flash (max): 0.27$, 5.5 minutes, 89k tokens generated.
* GPT-6.1-Sol (medium) : 0.21$, 2.2 minutes, 8k tokens generated.
API pricing. With Sol having SUBSTANTIALLY better performance.
If OpenAI cuts alternative harness support it will be a weird day trying to figure out what to do next, it's been so clearly the best bang for your buck (imo) for a while. maybe id finally have to give smaller models a try.
anything to avoid using the dogwater codex & claude code tuis.
anyways this seems like a nice cost improvement over GPT 6 Sol and I expect this will be my new daily driver.
not saying this is the case here but it does feel a bit like wine tasting sometimes, everyone claims to be an expert that can taste a few tokens and tell you exactly what region and vineyard its from.
So we've already reached the commodification phase, due to either the frontier models not getting usefully better or the models already being good enough, or both.
In either case, there appears to be a plateau in performance. With local open-source models catching up.
Interesting times.
Fuck altruism, ammi right? lets make money, gobs of it by screwing the middle users as much as we can to push them into just two tiers: Ones that use it for recreation and others that pay through their noses.
Guess not?
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
Eventually: black hole.
{"type":"item.completed","item":{"id":"item_0","type":"error","message":"Model metadata for `gpt-6.1-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues."}}
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Edit: for context, just Steam alone has ~200million monthly active users.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
How many DCs are devoted solely to gaming?
Today physical discs are rapidly becoming a niche that newer consoles just won't have. PC games are almost never sold physically anymore.
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
An entire planet. Just Steam alone has one or two hundres million monthly active users.
All of gaming is in the 300B USD range while just the CapEx of hyper scalars is already over 200B.
Yeah? Show me the big movements against computer gaming.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
I dunno what you're not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.
Also gamers tend not to spin up dozens of other machines to also game for them.
Huge misstep releasing it.
I can handle issues much better if they are predictable even if the model makes mistakes — much more frustrating when the model is erratic
I find codex wanders off road more often and fails to see the “bigger picture” (as much as LLMs can see the bigger picture at least)
And tbh when it was first released Astral felt even worse
I’m being forced to use it right now and at the end of the day I’m making do so it’s fine, but Claude makes for a smoother experience
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
After what DeepSeek pulled with V4.1 Flash I've given up on trying to map LLM versions to semver.
The coverage around 6.1 Astra seems deliberately playing into the dubious, recently headline "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves back over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the "safety" concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.