NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Show HN: TurboGPT: train 22KiB transformer in 13s (github.com)
_345 7 hours ago [-]
Trying to be a little less negative than the guy that got flagged, I too don't understand the motivation behind these projects. I've seen a hundred of them at this point- and each of them is probably worse and has less learning value than the one Andrej Karpathy made to teach people the building blocks involved in a GPT

So, why do people keep making these? Asking genuinely

Axodouble 4 hours ago [-]
Fun? I am unsure why it keeps getting shared though, I too have reinvented the wheel a few times for fun.

* Edit, I did however just notice literally all of this is just another vibeslopped banger, so I am not quite sure how much fun there really is, feels more like coding for the sake of keeping the wheel turning.

alightsoul 5 hours ago [-]
to add it to their resume and "prove" they know how machine learning works
jazzpush2 5 hours ago [-]
You mean pushing everything up in one giant commit with Claude isn't learning!?
lostmsu 4 hours ago [-]
> So, why do people keep making these? Asking genuinely

I extensively used minGPT for home experiments on transformer architecture. It is great for learning!

However, if you want to scale the experiments up at home you need to go faster. Karpathy made optimized https://github.com/karpathy/nanoGPT, but it is tuned for "8XA100 40GB node in about 4 days of training".

13s is a bit overkill here (my machine builds that project in 30s). But it gives some space for experimentation with architectures that don't have optimized primitives.

vjsrinivas 7 hours ago [-]
Interesting trend of labeling models after how big they are on disk vs how many learnable parameters they have.
jey 9 hours ago [-]
At that scale it seems like you could just solve the KKT conditions directly. Exaggerating but only a little
w4yai 7 hours ago [-]
"hn1g.txt" => not available
lostmsu 4 hours ago [-]
Just uploaded to https://huggingface.co/datasets/lostmsu/hn1g/tree/main

But it is a byte predictor. You can train it on any file.

pane_poker_11 7 hours ago [-]
[dead]
fleshmonad 7 hours ago [-]
[flagged]
kadoban 7 hours ago [-]
Jesus. Would hate to see the work you _do_ want to shit on.
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 05:22:12 GMT+0000 (UTC) with Wasmer Edge.