NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Jeeves. Reasoning improves Jev-like decision models (github.com)
sharih 19 hours ago [-]
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
zihotki 19 hours ago [-]
I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
nico 18 hours ago [-]
For email you can use a classifier

One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier

With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)

Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)

janalsncm 13 hours ago [-]
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.

You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).

nico 13 hours ago [-]
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
zihotki 16 hours ago [-]
I wonder what numbers you'd get using another system one model - Contrastive Language Model https://contrastive-lm.notion.site/

That model scales very well with quantities of requests.

atombender 18 hours ago [-]
> hold your horses to paint it as dirt cheap

For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.

idiotsecant 17 hours ago [-]
Darmok and Jalad, at Tanagra
calebhwin 17 hours ago [-]
How are you benefiting from prompt caching for simple classification?
zihotki 17 hours ago [-]
There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.

With Jev you each time pay for your prompt, you can't cache it.

sarkarghya 15 hours ago [-]
I mean, it sounds like it's only ideal for cases with significant system prompt overhead. I don't think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
jedberg 16 hours ago [-]
Are you getting better performance from an LLM than a Bayesian classifier?
StarlaAtNight 6 hours ago [-]
BLASPHEMY! OUT WITH YOU!
catlifeonmars 5 hours ago [-]
Could you not just copycat jev and run a fast, small local model?
HawtAds 15 hours ago [-]
How many requests per second do you have for spam that you are reliably hitting the Luna cache?
simplisticelk 14 hours ago [-]
Is that just because the Jev implementation is less mature? Couldn't it also implement prompt caching?
tyre 18 hours ago [-]
What are the costs compared to an ML model?
olgava 18 hours ago [-]
[dead]
amelius 16 hours ago [-]
Next step: make it classify the next word.
esafak 19 hours ago [-]
Jev ought to offer a flex mode that uses their spare capacity for a discount.
TN1ck 18 hours ago [-]
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.

It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.

[1] https://tn1ck.com/blog/jevdit

TN1ck 14 hours ago [-]
Update: Jeeves took about 2 hours to moderate 394 data points and performed really well. It’s not as good as Jev, but it’s super close! In general, it’s super cool that you can tune how strict you want content moderation to be with these models.
nicowaltz 14 hours ago [-]
cool to see!
thm 19 hours ago [-]
Ask Jeeves - Only took us 30 years to come full circle.
rsingel 16 hours ago [-]
Too true. I worked there.

Ask Jeeves hired hundreds of cheap liberal arts majors to classify data, some users thought Jeeves was real, the stock spiked when big companies hired Jeeves to automate support thinking it was a silver bullet, and the whole thing collapsed when a better model came along, and it degenerated into ripping off rubes with bottom of the barrel ads.

kridsdale1 16 hours ago [-]
Sounds like the story of OpenAI in 6 years
NetOpWibby 16 hours ago [-]
Damn, what a way to go.
kkukshtel 16 hours ago [-]
You found the joke!
victordmor 18 hours ago [-]
I met one of the founders once in Oakland. Amazing fella.
tmnstr85 19 hours ago [-]
this was the comment i came here for
onaclov2000 19 hours ago [-]
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
aftbit 17 hours ago [-]
I used Altavista
davedigerati 17 hours ago [-]
lol was thinking the exact same, named some ML projects Jeeves along the way...
itzikkatz 15 hours ago [-]
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
theanonymousone 17 hours ago [-]
This reminds me of "on-premise cloud".
teravor 17 hours ago [-]
you don't need to post-train anything for this.

just get an LLM to think and then force it to output a specific json with prefill post-think.

make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.

betenoire 17 hours ago [-]
A classifier is a subset of generative text, so I think responses like this miss the point. Jev is cheap enough and fast enough to sprinkle across your app in ways that LLM would be infuriatingly laggy and unnecessarily expensive, and it's never going to be injected to provide a sorting algorithm in python.

The point isn't that new type of problem has been unlocked, rather a new approach that can unlock new use cases.

teravor 17 hours ago [-]
this is exactly why Jev doesn't have thinking.

when you want a machine to reason about the prompt and generate a structured output not using an actual LLM makes no sense. I have been doing it since the first chain of thought open models became available.

perhaps there may be a way to get a Jev-type model to think for a very specific number of steps to gain control over its latency, if so that would be the next step. truncating LLM thinking like this does not work well, and its thinking isn't efficient anyway.

alienbaby 20 hours ago [-]
Just curious, where has this term 'noul' come from for yes/no ansers?

/a bit more digging and..

A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.

I hate it :)

LudwigNagasena 19 hours ago [-]
In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term.
finding_alfred 13 hours ago [-]
Jeeves pretended to understand. Is there data showing Jev's actually calibrated?
doginasuit 19 hours ago [-]
I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.
k__ 19 hours ago [-]
The whole "no hallucinations" premise is based on that.

Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.

kjs3 19 hours ago [-]
force the user to decide in the end

And that's...bad?

k__ 17 hours ago [-]
Not entirely.

I think, it's a bit much to call this "no hallucinations".

Technically true, but in practice you could still choose the wrong result or the probabilities can be off.

doginasuit 19 hours ago [-]
That seems like the only possible way to eliminate hallucination, short of a model that is never wrong.
rusk 19 hours ago [-]
Wait til you hear about how digital circuits work at die level
keepitwiel 20 hours ago [-]
Bernoulli
user3939382 20 hours ago [-]
If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this
kevindamm 20 hours ago [-]
Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.

In an analog circuit, maybe.

user3939382 6 hours ago [-]
Both the voltage and the threshold are probabilistic. Within a tight envelope sure. This is true even of the physical fields that comprise the transistor itself never mind the signal it’s designed to discretize. A more extreme example is a transistor hit with a stray gamma ray. When you dig deep enough there isn’t technically a Boolean anywhere, that’s a platonic construct that makes discussion of engineering convenient.
trencedamp 11 hours ago [-]
Jev noob here. I'm seeing all this jev talk and I understand the difference between this and normal models, but what are some actual use cases for jev?
swader999 20 hours ago [-]
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
winddude 15 hours ago [-]
That completely defeats the point of something to make decisions faster.
zerop 20 hours ago [-]
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
swingboy 17 hours ago [-]
Any good classifiers like this or Jev that support image input?
rgbrgb 14 hours ago [-]
openai's new Decisions API looks to be targeting that https://openai.com/index/devday-2026-recap/
quantized_state 17 hours ago [-]
I'd assume this would work with Qwen's image encoder probably better after a bit of tuning
RamblingCTO 19 hours ago [-]
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.

But funny that jev is getting its lunch eaten apparently in under two weeks?

danieltanfh95 18 hours ago [-]
it just a classifier. I guess we have to thank typesafe for spending VC money on marketing classifiers as decision models instead.
RamblingCTO 27 minutes ago [-]
it's a generalistic classifier w/training needed tho. as far as I'm concerned that is somehwat novel and very practical. you get an ok classifier without working out training data and whatnot. sure, you can do classic ML experiments, but the closest that comes to mind is auto ml. not sure how far we got there, it's been a minute for me on this front.

so I don't agree on the premise that it's "just a classifier". it's not something earth shattering, but practical nonetheless

/e: there's another post from sebastian raschka on this: https://magazine.sebastianraschka.com/p/classifier-history-a...

santadays 17 hours ago [-]
Doesn't the fact that it's general purpose warrant a new term? It's partly that it doesn't need to be trained, but it's also able to play games based on game state, I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states. The general purpose llm world understanding underneath it allows for this.

I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.

I like the term decision model and I think it's warranted.

danieltanfh95 6 hours ago [-]
outside of ML, for normies i mean, classifier models are supposed to be general purpose. Like I don't think we say segmentation model to do X. In ML, we know of their constraints so we usually say classifier models trained to do X.

Not complaining that decision models are a better term though.

nico 16 hours ago [-]
> I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states

Yes, one general classifier would be very hard to train. However, you can create a sort of ensemble of classifiers, each trained in different tasks

I’m currently experimenting with this. So far I’ve combined classifiers for 13 different datasets, my target is 95 (the ones Laya used for training)

svachalek 15 hours ago [-]
This isn't eating Jev's lunch. This is someone who doesn't understand the entire use case of Jev replacing it with something that doesn't handle it at all.
pavlov 19 hours ago [-]
It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
woadwarrior01 20 hours ago [-]
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
HarHarVeryFunny 15 hours ago [-]
The number of people who feel the need to try and argue that you don't need Jev can only be astroturfing by those with something to lose - Anthropic and OpenAI employees.

Like it or not, companies are going to use Jev unless you can offer something just as cheap and fast.

I wonder just how much of the business automation market, previously held by LLMs, is at risk here?

druskacik 17 hours ago [-]
How's the performance compared to ordinary 9B LLM with structured outputs? Both accuracy and speed?
loclol101 18 hours ago [-]
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
Naitik88 19 hours ago [-]
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
AnodicElegy 19 hours ago [-]
I'm surprised we haven't seen a "Jehovah" yet.
jadar 19 hours ago [-]
With the amount of talk about "inventing god", I'm surprised too.
mxkuzn 19 hours ago [-]
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
captainbland 19 hours ago [-]
See if it can beat Jev's Pokémon benchmark
mxkuzn 18 hours ago [-]
what benchmark this one or it's just fun?
captainbland 18 hours ago [-]
https://news.ycombinator.com/item?id=49845172

I guess it's not really a benchmark but you could say if it can do it faster it sort of could be taken as one.

quantized_state 17 hours ago [-]
The diffusion drafter adaptation is nice
raverbashing 20 hours ago [-]
Jeeves, that's a name I haven't heard in a long time...
gizajob 19 hours ago [-]
Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended.
fishfasell 20 hours ago [-]
If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia
grokkedit 20 hours ago [-]
jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves.

soon I'll make sure that my home assistant pod answers to "Hey jeeves"

kjs3 19 hours ago [-]
We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door???
lherron 19 hours ago [-]
…a long time.
singularity2001 17 hours ago [-]
In my experience, Jev is only faster because it's a small shitty model. Any objections?
esafak 18 hours ago [-]
Jev-like models give calibrated decision probabilities, but at low accuracy.

So why didn't they show both??

phplovesong 19 hours ago [-]
So "askjeeves" has been resurrected?
14 hours ago [-]
speedping 13 hours ago [-]
Meh. Wake me up when it reasons in latent space and answers in less than a second
SV_BubbleTime 14 hours ago [-]
“Somehow Jeeves returned…”
hjun1052 20 hours ago [-]
If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?
keel_dev 16 hours ago [-]
[flagged]
mabini 18 hours ago [-]
[dead]
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 07:51:46 GMT+0000 (UTC) with Wasmer Edge.