NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
GLM Built Its Own Inference Infrastructure (z.ai)
dada216 18 minutes ago [-]
We built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
axiosgunnar 12 minutes ago [-]
[dead]
tefkah 8 minutes ago [-]
> Today, GLM-5.3 has become an indispensable daily coding partner for everyone on the team, and it is moving steadily toward replacing us. If this trend continues, given enough compute and enough time, its endpoint is a system that can design and train its own successor entirely autonomously. This is known as Recursive Self-Improvement, or RSI.

Statements dreamed up by the utterly deranged.

embedding-shape 15 minutes ago [-]
I was gonna ask how people found their coding plans, and realize, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

They must have hit really hard scaling limits if the prices were hiked so much so quickly.

asp_hornet 2 minutes ago [-]
The way I look at it, their coding plan doesn’t retain data or use it for training making it one of the cheaper plans for me.

https://docs.z.ai/legal-agreement/privacy-policy

Daviey 3 minutes ago [-]
I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
broodbucket 13 minutes ago [-]
Yeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground
bbor 4 minutes ago [-]
It's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered.

For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.

bbor 8 minutes ago [-]
Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
jensb1 5 minutes ago [-]
What is "illegal" about it?
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 10:25:32 GMT+0000 (UTC) with Wasmer Edge.