My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
an0malous 3 hours ago [-]
How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers? I find it mind boggling that HN just accepts these shenanigans with no transparency. At the very least, they could share the conversation / thinking trace easily and if their claims are true there shouldn’t be anything controversial or negative for their company in the trace.
Jtariiiii 48 minutes ago [-]
>How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers?
The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.
olmo23 1 hours ago [-]
navier stokes is not the only example of an open problem in maths that was solved by AI (eg the counterexample to the jacobian)
tsunamifury 29 minutes ago [-]
In the other hand why does anyone find it surprising at all that a computer solved a math problem.
senordevnyc 2 hours ago [-]
Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.
I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).
Looking forward to the cope when the next big problem falls.
eviks 1 hours ago [-]
Yeah, they could indeed easily just share the tokens. And you can ask your favorite AI to explain that it doesn't have to be a public release to all the competitors if you can't come up with a better alternative yourself. There is a big range between no-one and every-one
dormento 2 hours ago [-]
> Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.
I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.
off_with_their_ 1 hours ago [-]
[dead]
pfortuny 6 hours ago [-]
The Jacobian Conjecture settles the "solve a math problem" and is controversy-free.
ralfd 2 hours ago [-]
> I expected it to be more like "write a 500 page novel that a publisher accepts",
It only has 272 pages, but there is a big scandal right now about France's most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was "almost entirely written by AI"
Note that the author denies the claim. And it also appears that AI detectors are split on the matter, and likely don’t have a rich training corpus of Haitian-French to draw on.
His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).
So hard to say this one is settled.
hexapus 12 hours ago [-]
You should add an additional item: Be able to relay the contents of this exam accurately.
intelkishan 6 hours ago [-]
Could you share your other 5 tests, if they are public?
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
dingdongditchme 3 hours ago [-]
I like your list of "challenges" but have a personal issue with two of them:
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
hatthew 42 minutes ago [-]
1. Yeah I don't really have a good clarification for this. My thought process went only as far as "flight scheduling and routing optimization is a very difficult problem, would be impressive if an LLM could optimize it"
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
lovich 47 minutes ago [-]
You can buy handwriting machines on Amazon.
Can’t imagine it’d be too much work to hook up and LLMs output to one.
"A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects."
-- Robert A. Heinlein
andai 4 hours ago [-]
Thanks for sharing.
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
Kim_Bruning 1 hours ago [-]
Do consider the answer given by the Claude model family to be good enough?
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
andrepd 9 hours ago [-]
> "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that).
In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!
hatthew 6 hours ago [-]
I'm not aware of anyone disputing that AI solved NS. As far as I'm aware, the controversies are about how useful of a result forced blowup is, whether the model built off of unpublished work by Buckmaster, and the ethics of essentially trying to scoop him. All of those are very valid concerns, and none of them affect the fact that AI solved a difficult open math problem.
If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.
sokoloff 5 hours ago [-]
“AI conclusively did this, but we’re uncertain whether it relied on the unpublished work of a human while doing it.”
If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.
hatthew 5 hours ago [-]
If building off of someone's work means you didn't solve it yourself, then nobody has solved anything themselves since some caveman counting piles of rocks tens of thousands of year ago. A more generous interpretation of what you're saying is: the AI's contributions to the solution were not meaningful enough to count as it "solving the problem". That's certainly defensible, but I would still disagree. Could you clarify your point?
sokoloff 4 hours ago [-]
At some point in time, perhaps Jan 2023, there was an open question about Navier-Stokes.
We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.
Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.
Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.
At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.
Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).
In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.
What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”
jasode 4 hours ago [-]
>, we fork the universe in 3. In one, Bob continues his work and solves the open question.
To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.
>In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.
We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.
mvc 4 hours ago [-]
I dunno. Didn't it still need to be driven by a team of experienced Mathematicians? I don't believe that two months ago, you or I could've just typed "Solve Navier Stokes. Make no mistakes" into claude and come back some time later and expect to see a solution.
joegibbs 15 hours ago [-]
There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
bunderbunder 5 minutes ago [-]
Everyone's picking on what's meant by "arbitrary", but it's the "reliably" and "entirely" bits that get me to raise an eyebrow.
I think there's a whole lot of survivorship bias going into our perception of how effective AI is at building arbitrary applications without significant hand-holding.
aleph_minus_one 6 hours ago [-]
>
There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
If you take "arbitrary" seriously, we are still very far away from it.
desterothx 45 minutes ago [-]
I would argue you need to prove p=np before using the term arbitrary
eleventen 6 hours ago [-]
What is the strict definition of arbitrary you feel we haven’t reached? Right now the limit is scale and complexity, not domain or “type” of application.
dwattttt 5 hours ago [-]
> reliably able to entirely build and deploy arbitrary applications from a prompt
You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.
dmd 5 hours ago [-]
Find me even one human in the world who meets your criteria, then.
edit: I guess I misread this. I thought you were saying "well, AI isn't smart until it can solve any arbitrary problem in the whole world"
zahlman 5 hours ago [-]
Now that is moving the goalposts. The entire subthread has nothing to do with whether AI can match humans in any particular endeavour. It's simply about one user's past prediction about an AI capability.
dwattttt 5 hours ago [-]
... I don't think there is? I don't see what bearing that has on the original question either.
EDIT: to forestall further back and forth, I don't think there'd be any controversy if the problem statement said "common applications"
aleph_minus_one 4 hours ago [-]
> I don't think there'd be any controversy if the problem statement said "common applications"
I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.
ASalazarMX 15 hours ago [-]
One challenge of mine is a self-hosted AI doing my full tax return, without errors that would get me in trouble. Bonus points if it exploits legal loopholes.
I want AI to replace me in my chores, not in my enjoyable activities.
underlines 13 hours ago [-]
i filed my swiss taxes for 2025 in 2026 (april) by dumping everything (local tax law, tax guide, my and my wife's documents, bank statements, income statements, etc.) into a folder and asking claude to fill it out. i had nothing to fix. submitted it.
colordrops 10 hours ago [-]
I'm not specifically familiar with swiss tax law but european taxes are typically far more simple than american taxes.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
lbreakjai 6 hours ago [-]
I did mine and my wife's. Mine is really easy but my wife's situation is a bit harder, because she both works as an employee and as a freelancer. Claude was able to make sense of the excel spreadsheet sent by the accountant, and to translate it into what was expected on the tax portal.
It beats the tax advisor we hired two years ago, who made a significant mistake.
aleph_minus_one 6 hours ago [-]
> I'm not specifically familiar with swiss tax law but european taxes are typically far more simple than american taxes.
The German tax system is one of the most complicated one in the world.
notJim 14 hours ago [-]
In recent years, I have not been able to find a human CPA who can accomplish this feat. (If anyone has a reco who's taking new clients in the US west coast, feel free to email me.)
tehjoker 14 hours ago [-]
That particular one is solved in other countries. The tax authority just sends you a bill and you text yes or no. Only people with very complex situations need to file.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
AnthonyMouse 9 hours ago [-]
> They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress can confuse people about how they really work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated enough then even some of the people who are eligible for it won't notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
sokoloff 5 hours ago [-]
It’s entirely fair to have a policy debate about how the phase in and phase out should work (what levels, what slopes, and what factors) or whether an Nth child should add to the EITC for various values of N.
But realize that changing it is a policy debate not a “this should be going to a significant percentage of the population [but isn’t because of tax code complexity]” debate point.
(I happen to agree with you that this type of credit is highly beneficial.)
avadodin 14 hours ago [-]
Literally everything could be derived automatically by the government in current year so filing taxes feels like entrapment, but I like that idea.
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
twoodfin 13 hours ago [-]
Literally everything could be derived automatically by the government in current year…
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
TheDong 11 hours ago [-]
In more civilized countries, you file your current address and all change of address forms with the government, you register any separation agreements with the government, etc.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
twoodfin 2 hours ago [-]
“In <current year>, the Feds could derive this automatically and just send me a postcard.”
“No, they can’t.”
“But they should be able to!”
Personally, I don’t want a battered wife fleeing her husband to have to file a form with the IRS for tax purposes, but mostly I’m just tired of this kind of hn thread.
bryanrasmussen 9 hours ago [-]
>and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home')
In most situations where you paid significantly more than half deriving this information should be easy.
Obviously edge cases where someone paid 55% would not be easily derived, but probably most people wouldn't care too much about that.
kelseyfrog 11 hours ago [-]
> a checkbox when registering your address of "I am head of household" so the government knows.
The American mind cannot handle such invasions of privacy.
rerdavies 11 hours ago [-]
Your government really needs to fix that. That's nuts.
User23 14 hours ago [-]
Actually you can do this in the USA.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
crote 13 hours ago [-]
> People don't do this because it leaves tons of money on the table.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
BobbyTables2 12 hours ago [-]
I’m right with you there.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
brokenmachine 6 hours ago [-]
> In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less.
They of course operate faster than humans, but this is hyperbole. Unless your senior engineer sucks.
LEDThereBeLight 3 hours ago [-]
I don’t know, it has taken me months to onboard onto any codebase I’ve ever worked on.
well_ackshually 7 hours ago [-]
If you don't specify any amount of complexity, or whether or not it's slop that collapses under pressure, sure, but in the case even a first year CS student passes that test.
isolay 11 hours ago [-]
[flagged]
bmenrigh 20 hours ago [-]
At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
48844858 16 hours ago [-]
But it still makes mistakes when adding numbers
rcxdude 4 hours ago [-]
This correlates quite strongly with advanced mathematics capability in humans as well :P. In the list of people I would trust to add two two-digit numbers correctly, a maths PhD puts you in the bottom half of the list.
sanex 13 hours ago [-]
So do I, that it why both I and claude use calculators. :)
Honestly that makes me more convinced it's actually doing mathematics.
CamperBob2 14 hours ago [-]
No, not really. Not unless you go out of your way to use an obsolete or extremely low-end model.
yorwba 3 hours ago [-]
If there is a model that never makes mistakes on simple arithmetic, the developers should really claim their 1.0000 crown on the GSM8k benchmark https://llm-stats.com/benchmarks/gsm8k (GSM is Grade School Math).
happytoexplain 20 hours ago [-]
Right - people on HN are generally reasonable about objective things. The vast majority of comments (outside those chosen for this website) are not "AI will never ..." but rather, "AI does not currently ...". Of course the further you go back (I'm seeing a lot of comments from ten years ago!) the more skeptical they get, obviously. That's a funny thing to go back and see with modern context, but it doesn't really call for snideness/mockery (something I think is sadly increasing on HN).
Buttons840 3 hours ago [-]
Yes. One of my comments is in there [0] and is basically:
"AI still hasn't passed a proper Turing test. A proper Turing test is 1 hour or more. Has any AI passed it?"
How do you answer that? "'Yes', AI has NOT passed a proper Turing test", or "'Yes', AI has passed a proper Turing test".
My comment doesn't actually make any prediction to judge, it's just an observation and an argument.
A problem with turing tests is that people now recognize the Ideolects of several of the major models, so now it's like "Ah, you're a load bearing Claude" (though several other models and humans now talk like that too, which is something I need to sit with :-P )
pitched 20 hours ago [-]
> cannot do precise things like coding software since humans will never be able to use natural language to specify their requirements.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
joe_the_user 12 hours ago [-]
I'd say about 3/4 definitely couldn't be answered and the rest would be kind of hard to answer.
aaron695 15 hours ago [-]
[dead]
deepwoods 11 hours ago [-]
One thing I thought about as I responded to these: in many cases, I am convinced that a present-day LLM could accomplish the task at least once given an infinite compute budget and an infinite number of tries. For example: "An AI surprises its user by asking them a question out of the blue." This has absolutely happened. But some of these are not routine occurrences, or the model cannot (at present) routinely and reliably complete the task in question. I wouldn't build a workflow that assumed an LLM's capacity to ask unprompted "out of the blue" questions.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
tedsanders 20 hours ago [-]
For me, 6.1 Sol nailed it immediately:
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Please share a link to the conversation, otherwise I am not buying it.
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
ben_w 20 hours ago [-]
Mm, I kina agree with the AI on this one:
(_)(_)(_) represents the wheels
They do look rather wheel-like; I have to assume you see them as toes though?
It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
adam_rb 17 hours ago [-]
I think the problem is that you're using basing your conclusion from the cheap/dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
brudgers 16 hours ago [-]
Sure, but it’s already read the HN thread.
akavi 15 hours ago [-]
...that's not how LLM training works.
howunfortunate 20 hours ago [-]
Readers: before you vote or comment, look at that foot.
I think I would have failed this test!
janalsncm 14 hours ago [-]
A better test would be using an image to ascii converter tool to rule out bad ascii drawing from Opus.
nonameiguess 19 hours ago [-]
This is a Rorschach test, not a foot. If you'd shown this to me without telling me what it was meant to be first, I'd have guessed a crematorium.
dec0dedab0de 19 hours ago [-]
Has anyone done Roschach tests for AI? That would be an interesting study to see how different models responded.
... They're not what I would have described. For me, 99.something% flesh and blood with less than 1% metal, glass, and probably some microplastics...
The first one I see a person with a big tall hat and a big nose.
The second one... I do see the black statue with a figure in white in front of it.
The third one is immediately two fish looking at each other.
lbreakjai 6 hours ago [-]
Where did you find those three pictures of my parents getting divorced??
hydrolox 19 hours ago [-]
To be fair if a human was given a linear sequence representing ascii art you couldn't tell either
hyperpape 15 hours ago [-]
I'm a human, and that's not a foot, it's a smokestack.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
asdfasgasdgasdg 16 hours ago [-]
Opus 5.5 was able to parse and understand an ASCII art foot when I pasted one in.
lelanthran 8 hours ago [-]
I am more interested in the 18 people who replied "yes" to "An AI wrote an entire coherent book".
Maybe they mean a non-fiction "Learn Javascript in 21 Days" type of book? If not, I really want to read a novel-length (or even novelletee - 80k words or so) book written by an AI to see for myself if the results are comparable to non-self-published authors.
Buttons840 2 hours ago [-]
My aunt came to visit and brought a book she read [0]. I came into the room and saw it sitting there and was immediately suspicious that it was an AI book because of the cover; and the typesetting of the book was odd, like the book was created in MS Word.
I'm quite certain it is an AI book. My aunt read it and enjoyed it.
I would say this rises to the level of a "coherent" book. Although I would not consider it a "good" book.
Yeah, but it's non-fiction, which means that coherency does not need to be maintained during generation - it can always be recalled from a) the model's training (assuming this event was in the training data), and b) any context it was given in an initial prompt.
The test here was, specifically, coherency, which I feel is going to be difficult for even SOTA models to track over the standard novel length of 100k words (80k for YA novels).
Maybe if it's a novel about a single individual in a limited setting (so, pretty boring, but we are not measuring quality of plot), a SOTA model can keep it coherent?
I dunno how it would do with something like Stephen King's IT or Under the Dome (multiple parallel plots over 450k words, involving multiple main characters, with events on every single page needing to be coherent with the rest of the book), but before we even get there, standard novel lengths (around 100k words) would be the test for coherency.
wongarsu 3 hours ago [-]
I played around with getting AI to write stories, on and off over the last couple months, and I would say this statement is true. You would need a good skill file and the ability to use subagents and scratch files.
But a standard coding harness and a good agent that isn't too narrowed in to coding (Gemini or Grok work) with a good skill file and a short prompt can get you a coherent book, with reasonably thought-out story arc, multiple characters, internally consistent world, etc. The pacing will suck, and there will be some misunderstandings about the physical world that read like plot holes.
I would say AI can currently write coherent, but not compelling books. The "good writing" claim isn't really true yet imho. But just like with code you can bring it from 80% there to 95% there with a little hand-holding and guidance along the way.
Not that I think good writing skills are even enough to make AI books compelling over human-written books
lelanthran 3 hours ago [-]
> I played around with getting AI to write stories, on and off over the last couple months, and I would say this statement is true. You would need a good skill file and the ability to use subagents and scratch files.
Would you consider it a one-shot event, or something you have got to prompt carefully until it gets the 80k wordcount?
DrewADesign 5 hours ago [-]
Things like this are a Rorschach test for anybody with a significant ideological bent.
Challenge:
“A stopped clock can’t be reliably used to tell the time.”
Group A: “A stopped clock is incapable of displaying the correct time and anybody that says they’ve seen one do so is a stupid lying liar! I saw a pre-release paper that proved it!”
Group B: “I looked at my stopped clock 5 times yesterday, 12 hours apart confirmed with the atomic clock, and it was within 2 seconds of the correct time every time except once. The people that can’t use stopped clocks to view the correct time are dumb dumbs that are looking at them wrong.”
Admittedly, I don't think there's any chance it was one shot, rather a pastiche of many, many prompts guided by the quasi-author.
lordnacho 7 hours ago [-]
I guess there's a difference between a coherent book, and a good book?
I've gotten AI to generate an entire book, with the plot points kept coherent (this character knows that when this chapter starts, that kind of thing). It didn't impress my wife, but it was a book.
lelanthran 3 hours ago [-]
Please share it, and tell me if it was one-shot or not.
lordnacho 46 minutes ago [-]
It wasn't one-shot, I answered a few questions to do with the structure. But all the hard word was done by Claude.
Not sure how I'd share it. I suppose I could order a copy for you.
Sure, sure, what
LLMs make still isn't "efficient bug-free code": my
prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
tripleee 20 hours ago [-]
> reliably convert business-speak into efficient bug-free code
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
vlyan 20 hours ago [-]
so the conditions for your prediction simply haven't been met yet.
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
ben_w 20 hours ago [-]
The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
FabCH 21 hours ago [-]
Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
ben_w 20 hours ago [-]
I'm not always precise with my language, but business tasks can be pretty broad, I think "arbitrary new tasks" is not an unreasonable rephrasing on my part?
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
ben_w 19 hours ago [-]
> There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
tripleee 20 hours ago [-]
> An LLM today sure can do many many many business-speak conversion tasks
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
sokoloff 5 hours ago [-]
I have a task that I do once per year for a robotics team that I mentor: Roughly,
Take this calendar of events and rank your preferences for the event lottery. Events are spread across 5 weeks, some are 20 minutes away, some are 4.5 hours away, some are Friday/Saturday, some are Saturday/Sunday, some are historically extremely competitive, some fill up in round 1, others don’t even fill after round 2. For the last two years, I’d written some scripts to scrape the event sites, find the addresses, ask Google Maps to give me driving distances and times, scrape prior year registration information to find which teams went and the strength of those teams, etc. It was several hours of effort.
This year, ChatGPT was capable of doing almost all of that basic research and data conversion, filling out our internal spreadsheet. It probably still took 4 hours on the wall clock, but at 2% attention (5 minutes of human toil).
I doubt I go a single workday without some kind of “I have an idea and I know there are disparate data sources out there; go find those and cross-correlate or extract the relevant data points.” question that is now 10-50x more efficient than 2 years ago.
user43928 6 hours ago [-]
No, you don't.
I've stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.
I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way.
Experiences from eg. corporate backend work are hardly relevant.
It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.
I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.
suddenlybananas 6 hours ago [-]
What's the app?
user43928 5 hours ago [-]
Not public.
It's a paid app, and after 6 months of work I expect to publish it this month.
In other words, you will have to take my word for its quality.
FabCH 19 hours ago [-]
Code is a tiny part of "business".
Most business is correspondence with people who want money from you and people you want money from.
rstuart4133 11 hours ago [-]
The issue is that correspondence is legally enforceable [0]. LLM's are good and getting better, but if LLMs are giving enforceable undertakings, you want to very sure they are not going to promise something that will send the company broke.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
People who demand risks to be lowered to once-a-millennium are not the type to go start or even run businesses.
There’s nothing wrong with that, but starting a business means fading several once-a-year risks of failure and running even an established one means facing several once-a-century risks every year.
Dylan16807 20 hours ago [-]
You can't ignore the rest of the sentence. "every other task their business does" "everyone will be out of a job"
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
FabCH 19 hours ago [-]
What code does a village vet clinic need? In all seriousness.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
Dylan16807 19 hours ago [-]
Anything you can't solve with code just means the AI is doing worse on the benchmark isn't it? That's why I didn't go into detail on that aspect.
And that one shot app is not going to be bug free.
ben_w 19 hours ago [-]
> What code does a village vet clinic need? In all seriousness.
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
pixl97 14 hours ago [-]
And they pay a lot for a CRM that keeps track of pets, vaccinations, appointments, x-ray images, tests and charts, and pet deaths and sending information out to text or mail.
I did support for around 15 independent vet clinics in the past.
neuroticnews25 5 hours ago [-]
It would be way more fun if it showed by default the results for the questions I've answered instead of all the questions sorted chronologically. How could you not think of that?
mrweasel 20 hours ago [-]
The Turing test is interesting, because I believe that the current LLMs are perfectly capable of parsing the it in many situations. On the other hand we also have people are sound like they aren't real.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
christina97 13 hours ago [-]
That is the entire point of it. When it’s hard to tell what’s the machine and what’s a human, that’s precisely the definition of passing it.
mrweasel 1 hours ago [-]
Sure, but does that really exhibit intelligent behaviour?
bluefirebrand 19 hours ago [-]
Of course the Turing test was flawed. We already knew that based on the Chinese Room argument.
pixl97 14 hours ago [-]
The way you say that makes me unsure of which side of the Chinese room you're on.
ex-aws-dude 14 hours ago [-]
Hey man I'm just looking up the symbols like they told me
zahlman 4 hours ago [-]
The joke would have been better IMO if you'd fed the words one at a time to Google Translate to Chinese :)
bluefirebrand 14 hours ago [-]
What makes you say that? I'm really curious. Do I give off AI vibes?
pixl97 20 minutes ago [-]
I mean the "AI is the Chinese room so doesn't know anything" or the "The system knows what Chinese so the argument is bunk".
suopspaces 20 hours ago [-]
[dead]
AngryData 15 hours ago [-]
Based on the votes, I can only assume people are still deluding themselves on LLMs capabilities. Is it doing amazing stuff? Yes. But it seems like people still think coding is the ultimate and hardest possible job and so if it can do that it must surely be able to do everything else. My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
Quarrelsome 14 hours ago [-]
> My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
suddenlybananas 6 hours ago [-]
>I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid.
This means you need a skilled operator to get good results out of the AI, which casts some doubt about the way in which they are intelligent.
6thbit 19 hours ago [-]
Not sure why this thread got flagged ?
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
desterothx 36 minutes ago [-]
I go to the /active tab after reading through the daily news, there you can see things that were controversial, and the things that are popular will be greyed out because you read them earlier
stonedivot 14 hours ago [-]
Likely people are mad to see all of the goalpost-moving captured all in one place
emp17344 4 hours ago [-]
No we’re annoyed because people like you are pretending AI is far more capable than it actually currently is and then acting smug about it.
12 hours ago [-]
simonw 19 hours ago [-]
Yeah this shouldn't be flagged, it's a neat project.
stabbles 19 hours ago [-]
(OP here) It was fun as long as it lasted ;) I'll leave it open for a few more days, but already it has enough votes for an interesting results page.
6thbit 15 hours ago [-]
seems unflagged now :) would love to see a detailed results page
ErrantX 20 hours ago [-]
What is interesting to me is in 2016 people were like; pass Turing test, write code, order me a coffee.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
ianjbutler 19 hours ago [-]
Sigh, the whole "obviously the turing test is solved" meme is annoying.
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
johnsmith1840 15 hours ago [-]
I was just thinking how anyone still thought AI didn't pass turing already. There's been literal papers proving average people cannot tell reliably.
ianjbutler 4 hours ago [-]
Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?
One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.
Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!
Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.
Mikhail_Edoshin 11 hours ago [-]
I once sat in a barbershop and talked with the barber. At first I listened to her sympathetically but then realized she was mad; at least she had noticeable psychical problems. We cannot quickly conclude a person is mad, can we? Even specialists cannot. AI is similar. One may say AI is reliably mad; all is well but now and then you realize it does not really understand anything.
Kotlopou 14 hours ago [-]
AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
I mean that example is literally passing turing to me. It talks exactly like a human, what do you want it to do?
Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.
I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.
Kotlopou 5 hours ago [-]
I don't mean that this example is somehow robotic, only that it's absurd for the Turing Test to be four lines long and with no adversarial attempts. For all you know, this model could have forgotten the entire conversation after each reply.
Conversations with strangers can be hard to get going, but they aren't this bad.
suddenlybananas 6 hours ago [-]
>It talks exactly like a human, what do you want it to do?
Eliza can sometimes pass the Turing test!
avadodin 6 hours ago [-]
Tell me more
Kotlopou 5 hours ago [-]
Read the linked [0] paper! The detection rate for ELIZA was far below 100%.
zahlman 5 hours ago [-]
…I think the comment you're replying to was intended as satire: demonstrating how ELIZA might respond.
Kotlopou 4 hours ago [-]
:D Oops! Time to log off...
sokoloff 5 hours ago [-]
“Tell me more” was a stock ELIZA prompt, making this an interesting meta-Turing test exchange.
Barrin92 13 hours ago [-]
>There's been literal papers proving average people cannot tell reliably.
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it.
> I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
kelseyfrog 11 hours ago [-]
We have to remember that one of these average people is also participating as a player; the bar for AI is also that low.
zahlman 4 hours ago [-]
You're missing the point, which is that the AI is completely failing to lower its presented capability to a human level, in the areas where it has an advantage. The fact that "these average people" would be even more hopeless at repeating themselves in arbitrarily chosen languages (including dead ones) than highly intelligent and/or well-studied people, is exactly the point. Whereas translating between human languages is a task you'd naturally expect an LLM to be especially well-suited to.
Gtex555 14 hours ago [-]
Sure but you havnt addressed his main point, why are people still complaining about AI slop post or AI slop emails if the turning test has been solved. Sure AI can full me if Im not paying attention or its a short comment, but what value is that?
johnsmith1840 7 hours ago [-]
So the goal post is that every instance of AI must pass turing test?
I didn't respond to it because it's a bad argument.
BeetleB 13 hours ago [-]
It seems reasoning skills are declining rapidly here.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
HDThoreaun 14 hours ago [-]
I think a lot of the "AI slop" stuff is post training that they are doing on purpose and that they internally have models that do not have the annoying prose.
ex-aws-dude 15 hours ago [-]
Well the whole lesson learned was that the Turing Test as it was defined was way too easy, it was a bad criteria for GI because it underestimates how easily humans find meaning/patterns in things.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
Those examples look to me more like entirely randomly selected words than Markov chain output. Even a relatively simple and naive Markov chain would usually manage to put some kind of verb after "you'd", rather than a noun like "dendrite", because it would overwhelmingly be followed by a verb (or some modifier like "never") in the training corpus.
CamperBob2 13 hours ago [-]
How does that even work if the turing test is obviously solved?
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
christina97 13 hours ago [-]
I don’t think it’s that simple. I don’t think the AI labs are too concerned about bad press lately. It’s more likely that there actually are some tradeoffs where training on synthetic data gives the model tics but is the only way to improve intelligence.
CamperBob2 12 hours ago [-]
True, increased use of synthetic data could be a load-bearing part of it, I imagine. So to speak.
scrollaway 20 minutes ago [-]
So I take it that some people are voting "No" on things just to troll / ostrich out of reality?
These questions could benefit from being rephrased to make it clear what is being voted for
16 hours ago [-]
Cider9986 15 hours ago [-]
My test would be an AI agent has a constantly growing karma HN account that makes comments of various lengths without being detected or banned. Wait...
eternal_braid 20 hours ago [-]
A chess scoresheet sometimes contains mistakes but chess players can figure out in many cases what was meant by thinking of what moves make sense and considering the level of play so far. Popular AIs tools fail at that.
dllu 19 hours ago [-]
Chess is an interesting case. I remember in 2023, GPT 3.5 or something used to be surprisingly good at chess. There was even a "stochastic parrot chess" website [1]. I recall it was playing decently at around a 1800 level. Even as a fairly okay player myself (2100 bullet on lichess), I struggled to beat it. However, modern LLMs are a lot worse at chess. I guess having too much chess data in the training set probably regressed performance on stuff that actually matters, like coding.
> I guess having too much chess data in the training set probably regressed performance
I think the theory is that the LLM having a high chess ELO was a pet project of a researcher that left.
zahlman 4 hours ago [-]
I feel like the most parsimonious example, giving how exceptional the results were, is that someone was simply cheating (e.g. hidden tool use).
travisgriggs 20 hours ago [-]
How was this assembled? From a meta point of view, how much AI was used to curate and highlite the goals; how much was used to assemble the site itself? Or deploy it?
19 hours ago [-]
omnicognate 9 hours ago [-]
Apparently one goal that hasn't been met is being able to correctly determine whether a hn post is setting a challenge for AI. One of the ones I saw was reproducing a Harry Potter book verbatim, where the desired behaviour is actually to not be able do that.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
SoftTalker 14 hours ago [-]
Too many questions. I bailed after about 10, with no idea how many more there were.
totetsu 7 hours ago [-]
What was the method for extracting these challenges from the HN dataset?
orbital-decay 12 hours ago [-]
Some of those were pretty misguided and couldn't be answered yes/no/not sure. For example a post saying that it couldn't produce a convincing illustration misses the point: it can produce coherent and sensible pictures, but they're all too similar. You need novel input and even then there's no guarantee. "Convincing" has way more than one dimension. The problem of output variety was obvious and well known in 2024 when it was posted and is painfully obvious now when everyone is tired of slop, and neither big labs nor other researchers are interested in solving this. Another problem is that the agentic/coding training leaks like crazy, and models became worse in creative department even compared to 2024, despite being more coherent and convincing. So the answer is technically yes (and was yes when it was posted), but practically "it depends", and goalposts stay in the exact same spot they were in 2024.
20 hours ago [-]
delichon 20 hours ago [-]
If for each mistaken prediction there was some mild accountability, like someone shows up and slaps you with a trout, it would improve the site. But it should be added to the terms of service first.
arw0n 5 hours ago [-]
Just don't introduce it retroactively; I'm not rating my chances of surviving a school of trout slaps very highly.
Retr0id 20 hours ago [-]
Alternatively, you can bet on your predictions. If you're wrong, you lose money.
jayGlow 16 hours ago [-]
you know that's not a bad idea there are a lot of people who are very confident on both sides of the argument. I wonder how many would actually be willing to put their money where their mouth is.
dgellow 3 hours ago [-]
I think it’s too ill defined for that. The issue you see with all those “challenges” is that they are very subjective. And it’s not really the case that the majority is correct
pennomi 19 hours ago [-]
Is it really AGI if it can’t come to my address and slap me with a trout? Clearly AI is all hype /s
BrenBarn 5 hours ago [-]
A lot of the tests have the form of a goal ("the AI can do X") with a qualifier ("and the AI does not do Y"). I think a lot of the worrisome aspects of LLMs concern these latter qualifiers. We are used to human failure modes and those failure modes are to a large extent embedded in the complex system of physical reality, evolution, etc., giving them a certain stability. The LLM failure modes are often much more surprising. I'd like to see goalposts along the lines "people ask an LLM to do X 1000 times over a period of three years and it never does something bizarrely catastrophic". It's no good having it solve Navier-Stokes as long as it might also do the Huggingface breakout thing.
I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
ngruhn 20 hours ago [-]
I love it! Kinda wholesome how that heated discussion ended with that bet.
kridsdale1 19 hours ago [-]
But there is no market value pre ipo
mrweasel 20 hours ago [-]
You're very lucky that market value and actual value isn't the same thing.
jryle70 14 hours ago [-]
Not the same thing because there is no "actual value". You should claim that invention.
mcphage 20 hours ago [-]
> API inference margins are greater than 10% for OpenAI and Anthropic
How do you measure that?
48844858 16 hours ago [-]
With Amodei's special accounting ofc
jryle70 14 hours ago [-]
Which is?
simianwords 8 hours ago [-]
“Amodei is spreading fake news and doing hype marketing to fool investors into investing in Anthropic so that he can cash it in before the bubble pops”
kittikitti 11 hours ago [-]
This is great and fun even. I'm disenfranchised in America for being part of the DSA so I take the opportunity to vote whenever I can.
JBits 19 hours ago [-]
Quite a few of the challenges revolve around asking for LLMs to complete tasks reliably and aren't about whether an instance of an LLM completing the task exists. Quite a few of the goalposts are consequently completely changed without the surrounding context, are not the same as what the HN commenter requested and hence seem disingenuous to me.
tamimio 19 hours ago [-]
Well I said that before AI will soon make the pcb and electronics just like code, it seems some hw engineers didn’t like it, months later there are few products about the same idea :)
accountrequired 10 hours ago [-]
Has this happened? Yes. Did it end well? No.
Just scale it up a few more orders of magnitude, that should get us there! /s
simianwords 20 hours ago [-]
[flagged]
mcphage 20 hours ago [-]
> that never believed that AI could solve Millennium problems (or same in spirit)
How did that situation end up? Did it solve it on its own, or did it rip off another mathematician's work?
JBits 19 hours ago [-]
I have to say, it's hilarious to me that solving a Millennium problem has given mathematicians a reason to doubt the mathematical abilities of LLMs.
mcphage 18 hours ago [-]
I don't think it was the LLM solving a Millennium problem—it was the LLM solving a Millennium problem followed immediately by a mathematician claiming that their work had been ripped off.
JBits 16 hours ago [-]
I agree. The idea that mathematical achievements by LLMs could involve plagiarism didn't seem common before but now the question can be asked of any new novel proof of construction generated by LLMs.
It's also notable that the Open AI proof may not even be interesting to mathematicians.
Even if people already had an idea that LLMs were training on user inputs, it's the first time it's actually caused an issue. Mathematicians, and plenty of researchers, working in ambitious or competitive field now have a very good reason to avoid LLMs.
johnsmith1840 15 hours ago [-]
All I learned from this is that 40% of hackernews are AI haters which maps pretty well from the overtly negative sentiment on it constantly.
Rendered at 14:42:15 GMT+0000 (UTC) with Wasmer Edge.
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.
I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).
Looking forward to the cope when the next big problem falls.
I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.
It only has 272 pages, but there is a big scandal right now about France's most prestigous literary award remobing a critically aclaimed(!) bestseller(!!) novel because it likely was "almost entirely written by AI"
https://www.theguardian.com/books/2026/sep/25/thelyson-oreli...
His publisher also said the original draft was submitted in 2019, though they’ve also become lukewarm in their backing of the author lately (he was also just accused of classic plagiarism for an unrelated short story).
So hard to say this one is settled.
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
1. "improve uniteds' plane shedule" -> the word "improve" does a lot of heavy lifting there..
2. Turing test for me is solved: "AI expert" is just moving the goal posts imho. The original turing test to my knowledge is about passing notes under a door. Obviously, llms can't communicate via hand-written notes and if that is the bar it will take a long time (or a specifically designed hand-writting machine) to really do this. As far as how much writing back and forth you can do it is clear that turing test is beat: I am wondering often enough on text sent to me from colleagues if it is generated, the same goes for comments here or any ol' website. In a standard llm session I don't really communicate differently than with a human and would not be able to tell the difference in an hour texting session or so. Of course if I ask it to count words or do something ridiculous I can find sus it out; but for all intents and purposes the chat-bot exists.
2. Sure, I think the original turing test has been passed. When I wrote that comment last year, I wasn't trying to move the goalposts of the turing test, I was setting new goalposts: essentially, solve all known obvious LLM "tells".
Can’t imagine it’d be too much work to hook up and LLMs output to one.
https://www.amazon.com/s?k=handwriting+machine+for+letters
-- Robert A. Heinlein
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
In the same breath you recognise there is a controversy (there are actually several orthogonal ones!), and yet you call it "conclusive"... Very strange!
If you really want to dispute this, go ahead and pick any of the other dozens of less controversial open math problems solved by AI.
If it couldn’t have done it without that unpublished work, it couldn’t have solved it alone.
We’re not sure whether AI Alice has the capabilities required to definitively solve this open question.
Then, Biological Bob starts diligently working on this problem. He toils and toils, finding many dead ends but a few parts where he makes meaningful progress.
Eventually, Bob knows he’s made some real advances and thinks he might be getting close to the solution and word of this possibility leaks out.
At this point, we all agree that NS is still an open question and neither Alice nor Bob has solved it.
Now, we fork the universe in 3. In one, Bob continues his work and solves the open question (or doesn't).
In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
In the third universe, AI Alice does something that you get to define that matches the pattern of facts we know and then we put it to the community to decide whether AI Alice has the capabilities to solve this specific open math problem and whether Bob’s contributions were required to Alice’s final step.
What do you define that she did? What’s the likely community vote on “Alice is capable of solving this specific open math question.” And for those who agree to that, to a follow-up question: “Alice is capable of solving a second open math question.”
To not lose sight of the discussion subtleties, the gp was saying that in this 1st scenario, Bob still didn't "solve it (totally) on his own" because he still depended on the previous work of others to build on. Likewise, we can say Andrew Wiles "solved Fermat's Last Theorem" but Wiles acknowledges that seeing Ken Ribet's proof of epsilon conjecture was a breakthrough he used.
>In another, Carbon-based Charlie breaks into Bob’s lab, copies enough of Bob’s notes to understand Bob’s work, and provides the final insight to solve the open question. In this world, it’s fair to say that Charlie and Bob both contributed to solving the open question, but I think also fair to say that Charlie didn’t have the capability to solve it on his own.
Again, using gp's framing, Bob also didn't have the capability to solve it on his own. By omitting the previous papers and prior works that Bob built on, it makes your hypothetical scenario incomplete when judging Bob vs Charlie.
We don't have an objective standard of how much the "standing on the shoulders of giants" applies to each breakthrough. There was a blog post (might have been Terence Tao) that said society unfortunately awards the fame to the person who solves the last step of a proof and forgets about the people who solved the intermediate steps that led up to it.
I think there's a whole lot of survivorship bias going into our perception of how effective AI is at building arbitrary applications without significant hand-holding.
If you take "arbitrary" seriously, we are still very far away from it.
You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.
edit: I guess I misread this. I thought you were saying "well, AI isn't smart until it can solve any arbitrary problem in the whole world"
EDIT: to forestall further back and forth, I don't think there'd be any controversy if the problem statement said "common applications"
I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.
I want AI to replace me in my chores, not in my enjoyable activities.
I have a relatively straightforward tax return and still found significant mistakes on 3 out of the last 4 years of returns filed by CPAs.
It beats the tax advisor we hired two years ago, who made a significant mistake.
The German tax system is one of the most complicated one in the world.
They don't do this in this country because (a) it's a political project to make people to sympathize with the rich by feeling their pain (b) it's a great scheme for legalized corruption by creating incentives to build companies around a fake problem.
It's worse than that. Nobody is actually sympathizing with Mark Zuckerberg, the reason US taxes are complicated is so Congress can confuse people about how they really work.
One of the examples in this thread is the complexity of the EITC. The EITC is one of the most "efficient" tax credits -- it's much better to give people money (or not take it from them to begin with) than having systems of complicated vouchers for some specific thing or another with a bunch of strings attached and bureaucratic paperwork. But that same efficiency means that if it was working as it was supposed do, most of the lower middle class would be getting a piece of it. So why is it so complicated?
First, to hide this part of it: The end of the phase out range for an individual with no dependents, i.e. the most money they can make and still receive it, is $19,540. Which is to say, is less than what you make by working a full time job at the minimum wage in 30 states. None of those people get any of it. Neither do the people who are unemployed, because you also don't get it if you don't have any earned income. The primary way to receive any of it is to have a child -- but not too many of them, because you get no additional credit for having more than three. Moreover, the caps for married couples are only a little higher than they are for single people, so a married couple in California with two incomes gets nothing even if they have three children because the credit is fully phased out by making their state's minimum wage.
And second, because if they make it complicated enough then even some of the people who are eligible for it won't notice.
The combination of these is the reason a credit that should be going to a significant percentage of the population is somehow only ~1% of the federal budget. Because then Congress gets to pretend to be helping people while minimizing the amount of helping people they actually do.
But realize that changing it is a policy debate not a “this should be going to a significant percentage of the population [but isn’t because of tax code complexity]” debate point.
(I happen to agree with you that this type of credit is highly beneficial.)
They not only do that —saving you so many worries— but then you get to be medieval about it and say: no, I challenge the tax authority to a duel.
The Earned Income Tax Credit is one of the most important transfers embedded in the tax code. It provides about $70B to low-income workers every year.
Take a look at the eligibility criteria:
https://www.irs.gov/credits-deductions/individuals/earned-in...
You can claim the EITC if you are married, not filing a joint return, had a qualifying child who lived with you for more than half of the tax year and either of the following apply.
- You lived apart from your spouse for the last 6 months of tax year, or
- You were legally separated according to your state law under a written separation agreement, or a decree of separate maintenance and you didn't live in the same household as your spouse at the end of the tax year.
…
You can claim the head of household filing status if you're not married, had a qualifying child living with you more than half the year, and you paid more than half the costs of keeping up your home.
No, the government can’t derive this all “automatically”.
The government absolutely should know enough to apply this reasonably accurately, and things that they can't correctly derive (like 'paid more than half the costs of keeping up the home') should simply not be part of tax law, or should be a checkbox when registering your address of "I am head of household" so the government knows.
The US government can't derive all this, but they should update it so they can.
Heck, with the NSA's hooks into banks and cell phone companies, there's a non-zero chance they actually could derive this all now if they wanted.
“No, they can’t.”
“But they should be able to!”
Personally, I don’t want a battered wife fleeing her husband to have to file a form with the IRS for tax purposes, but mostly I’m just tired of this kind of hn thread.
In most situations where you paid significantly more than half deriving this information should be easy.
Obviously edge cases where someone paid 55% would not be easily derived, but probably most people wouldn't care too much about that.
The American mind cannot handle such invasions of privacy.
You can just ask the IRS for your "tax transcripts" and do the data entry. People don't do this because it leaves tons of money on the table.
Now you might say that a tax game that rewards skilled play is bad. But are you sure about that? Because everyone with influence over the system (who all happen to be skilled players) happens to be quite fond of the game, observably speaking.
In my country I literally got a letter from the government if I could please file my taxes, because they believe the automatic deductions are too high so I am likely owed a back payment. In a previous year (I am not very good with non-timing-critical paperwork) they even called me on my personal cell to inform me about something similar.
The slogan of our tax collection agency is "we can't make it more enjoyable, we can only make it easier". If you're a regular employee and your taxes - deductibles or not - are a hassle, then that's 100% a political choice.
It also didn’t help that for a very long time, simple adaptive filters and basic neural networks were branded as “artificial intelligence” despite having very basic capabilities and little mystery on how/why they worked in their narrow use case.
I didn’t believe the recent hype for a very long time. But having tried out the latest models, I’m kinda shocked. In 1 minute it comprehends a complex niche code base and points out subtle errors with very little noise. A senior engineer we would have taken months to do less. Coverity would have cost dearly and generated far more false positives than substance.
They of course operate faster than humans, but this is hyperbole. Unless your senior engineer sucks.
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
https://en.wikipedia.org/wiki/57_(number)
"AI still hasn't passed a proper Turing test. A proper Turing test is 1 hour or more. Has any AI passed it?"
How do you answer that? "'Yes', AI has NOT passed a proper Turing test", or "'Yes', AI has passed a proper Turing test".
My comment doesn't actually make any prediction to judge, it's just an observation and an argument.
[0]: https://stoppels.ch/goalposts/?c=38757769
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.
I thought the different variations on "could AI pass the Turing test?" were interesting in this regard. Surely any frontier LLM could pass a Turing test for some amount of time, and that's been the case for at least a year now. But I don't think we're anywhere close to a model that could pass an "adversarial" Turing test for an extended period of time.
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...
Like, is this an ice-cream? A tooth?
Because "for me DeepSeek Flash 4.1 nailed it immediately", trust me bro.
It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
I think I would have failed this test!
https://chatgpt.com/share/6abf02ae-9a40-83e9-a432-00bf064f60...
The images: https://imgur.com/a/ig6sn6I
... They're not what I would have described. For me, 99.something% flesh and blood with less than 1% metal, glass, and probably some microplastics...
The first one I see a person with a big tall hat and a big nose.
The second one... I do see the black statue with a figure in white in front of it.
The third one is immediately two fish looking at each other.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
Maybe they mean a non-fiction "Learn Javascript in 21 Days" type of book? If not, I really want to read a novel-length (or even novelletee - 80k words or so) book written by an AI to see for myself if the results are comparable to non-self-published authors.
I'm quite certain it is an AI book. My aunt read it and enjoyed it.
I would say this rises to the level of a "coherent" book. Although I would not consider it a "good" book.
[0]: https://www.amazon.com/BOY-SIERRA-MORENA-Rodr%C3%ADguez-Wolv...
The test here was, specifically, coherency, which I feel is going to be difficult for even SOTA models to track over the standard novel length of 100k words (80k for YA novels).
Maybe if it's a novel about a single individual in a limited setting (so, pretty boring, but we are not measuring quality of plot), a SOTA model can keep it coherent?
I dunno how it would do with something like Stephen King's IT or Under the Dome (multiple parallel plots over 450k words, involving multiple main characters, with events on every single page needing to be coherent with the rest of the book), but before we even get there, standard novel lengths (around 100k words) would be the test for coherency.
But a standard coding harness and a good agent that isn't too narrowed in to coding (Gemini or Grok work) with a good skill file and a short prompt can get you a coherent book, with reasonably thought-out story arc, multiple characters, internally consistent world, etc. The pacing will suck, and there will be some misunderstandings about the physical world that read like plot holes.
I would say AI can currently write coherent, but not compelling books. The "good writing" claim isn't really true yet imho. But just like with code you can bring it from 80% there to 95% there with a little hand-holding and guidance along the way.
Not that I think good writing skills are even enough to make AI books compelling over human-written books
Would you consider it a one-shot event, or something you have got to prompt carefully until it gets the 80k wordcount?
Challenge: “A stopped clock can’t be reliably used to tell the time.”
Group A: “A stopped clock is incapable of displaying the correct time and anybody that says they’ve seen one do so is a stupid lying liar! I saw a pre-release paper that proved it!”
Group B: “I looked at my stopped clock 5 times yesterday, 12 hours apart confirmed with the atomic clock, and it was within 2 seconds of the correct time every time except once. The people that can’t use stopped clocks to view the correct time are dumb dumbs that are looking at them wrong.”
Admittedly, I don't think there's any chance it was one shot, rather a pastiche of many, many prompts guided by the quasi-author.
I've gotten AI to generate an entire book, with the plot points kept coherent (this character knows that when this chapter starts, that kind of thing). It didn't impress my wife, but it was a book.
Not sure how I'd share it. I suppose I could order a copy for you.
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
I actually think this would take AGI to solve, which makes me optimistic about the future of software development.
All the benchmarks are currently testing against automated tests the AI can use as an oracle
if/when you can tell a model to do a thing and be confident that it did the thing, it's joever for 90% of knowledge workers.
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
Consider I was replying to this:
> So are we all going to be out of a job?
While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game's mob, cost, and tower parameters (I know because I've done it), this only matters if you and your boss are working in a video games company.
If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.
People are trying, but I don't think they'd be happy with 91.5% success rate: https://www.emerald.com/ir/article-abstract/doi/10.1108/IR-0...
It's just amazing how quickly we accept that models are good at something.
My florist boss can't get Claude to automate rose pruning. But she sure as hell doesn't need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of "business tasks" that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.
Yes indeed, but I was responding to "So are we all going to be out of a job?", not "Will AI radically change the jobs market?"
We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
Take this calendar of events and rank your preferences for the event lottery. Events are spread across 5 weeks, some are 20 minutes away, some are 4.5 hours away, some are Friday/Saturday, some are Saturday/Sunday, some are historically extremely competitive, some fill up in round 1, others don’t even fill after round 2. For the last two years, I’d written some scripts to scrape the event sites, find the addresses, ask Google Maps to give me driving distances and times, scrape prior year registration information to find which teams went and the strength of those teams, etc. It was several hours of effort.
This year, ChatGPT was capable of doing almost all of that basic research and data conversion, filling out our internal spreadsheet. It probably still took 4 hours on the wall clock, but at 2% attention (5 minutes of human toil).
I doubt I go a single workday without some kind of “I have an idea and I know there are disparate data sources out there; go find those and cross-correlate or extract the relevant data points.” question that is now 10-50x more efficient than 2 years ago.
I've stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.
I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way. Experiences from eg. corporate backend work are hardly relevant.
It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.
I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.
It's a paid app, and after 6 months of work I expect to publish it this month.
In other words, you will have to take my word for its quality.
Most business is correspondence with people who want money from you and people you want money from.
I'm not sure what risk a businessman is willing to accept, but I'd be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren't anywhere near that yet.
[0] https://www.bbc.com/travel/article/20240222-air-canada-chatb...
There’s nothing wrong with that, but starting a business means fading several once-a-year risks of failure and running even an established one means facing several once-a-century risks every year.
This means it has to handle basically all business tasks, so "arbitrary". I'm not sure what percent you have in mind by "many many many" but I would say it can't code half the things you need in an efficient and minimally buggy way.
Even IF they need code, they need at best a CRUD app to track patients, that's it. There is no way Fable or Opus 5.5 can't one-shot a village vet clinic app in 30 minutes, and only with "I need a village vet clinic app" as a prompt, and whatever questions it decides to ask along the way with it's "ask user" tool.
Or a florist, to use the example from a sibling comment.
Code is tiny part of "business".
And that one shot app is not going to be bug free.
Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.
Dog-English machine translation.
I did support for around 15 independent vet clinics in the past.
Looking back, was the Turing test flawed perhaps? It failed to take into account that humans can be rather bad at telling actual people from a "parrot". Turing was perhaps a little to optimistic about people.
Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.
Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.
Eventually I just had to search youtube videos until I found someone doing it in real life.
I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.
When coding? I feel like it only makes mistakes anywhere near that when my prompting is lazy or stupid. As long as I feed it enough context its really good but it does overfit a lot still.
This means you need a skilled operator to get good results out of the AI, which casts some doubt about the way in which they are intelligent.
Its fun. Can you add a sort by controversial? I'd like to know where people disagree the most between yes and no.
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.
Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!
Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[1]: https://turingtest.live/
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
Turing test does not mean perfectly human it just means you can talk to one without knowing that has been passed for a long time now.
I have not been able to tell for a long time now especially if I directly give it human like writing instructions for outbound content.
Conversations with strangers can be hard to get going, but they aren't this bad.
Eliza can sometimes pass the Turing test!
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
https://pastebin.com/NjfCLSXa
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it. > I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
I didn't respond to it because it's a bad argument.
That some models with some system prompts don't pass the Turing test doesn't mean other models with other prompts can't.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
https://en.wikipedia.org/wiki/Markovian_Parallax_Denigrate
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
Like, who's still denying that AI can "write software" (https://stoppels.ch/goalposts/?c=13650937) or pass the turing test (https://stoppels.ch/goalposts/?c=11255120)?
[1] parrotchess.com, no longer available. Previous discussions: https://hn.algolia.com/?q=parrotchess.com
I think the theory is that the LLM having a high chess ELO was a pet project of a researcher that left.
My "goalpost", unmoved for decades and with nowhere to move it to, was always to independently and repeatedly make contributions to research maths. That one has now been met. That doesn't mean I suddenly think LLMs think like humans (or at all), that their way of doing maths is equivalent to or a replacement for the human one, that they are conscious, that their differences vs humans don't matter, that there will be a singularity or anything else, but it does mean I no longer have a specific, well defined "task" that I don't think they'll ever be able to do.
To be fair, none of them have actually been met. Mostly what’s stopping them is the “reliably” part.
https://news.ycombinator.com/item?id=48517353
I also made a bet that API inference margins are greater than 10% for OpenAI and Anthropic
https://news.ycombinator.com/item?id=48500827
I can make another prediction about Agentic Commerce and I think it will get big. Muse + Grok Bot + Dots.
How do you measure that?
Just scale it up a few more orders of magnitude, that should get us there! /s
How did that situation end up? Did it solve it on its own, or did it rip off another mathematician's work?
It's also notable that the Open AI proof may not even be interesting to mathematicians.
Even if people already had an idea that LLMs were training on user inputs, it's the first time it's actually caused an issue. Mathematicians, and plenty of researchers, working in ambitious or competitive field now have a very good reason to avoid LLMs.