Build your own decision model

(nishtahir.com)

97 points | by softwaredoug 2 hours ago

6 comments

  • nico 8 minutes ago
    This is very cool. If you are looking for something similar but more lightweight, that you can run (and train) on CPU, try out Jeffy: https://jeffyclassify.com/

    On GitHub: https://github.com/nicobrenner/jeffy

  • sva_ 13 minutes ago
    Does someone have examples of interesting stuff that has been built utilizing Jev/decision models? The way this is hyped up surely there must be some good stuff?
  • howunfortunate 1 hour ago
    As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally.

    Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.

    • firasd 1 hour ago
      I think part of what made Jev catch on is that the API is like an if(...) or switch statement

      People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal

      But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank

      • howunfortunate 50 minutes ago
        You know, it's a good point.

        Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.

    • jofzar 1 hour ago
      > But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.

      It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.

    • BoorishBears 5 minutes ago
      Maybe instead of screaming you can take this as a chance to level up your engineering.

      Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.

      Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)

      Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.

    • dominotw 14 minutes ago
      > As an MLE

      Yea but those models you were building were lame and inaccessible to play with for common devs.

      Just because they have the same api doesnt mean you were building the same thing.

    • copperx 1 hour ago
      I want to share the rage. Can you expound on what makes you scream?
      • howunfortunate 1 hour ago
        Idk, imagine you worked on Skype's B2B sales team for years and then COVID happens and Zoom blows up.

        Is it a better thing? Yeah. Does it affect me in any tangible way? No.

        But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time?

        • JMKH42 9 minutes ago
          I bet the timing was the key, people in the last year have been furiously building things that use LLMs as an API and as we work on this stuff we have systems with N LLM steps and M of them are frustrating because you want a specific choice picked or list of things ranked and sometimes the llm will just output something else entirely!

          So along comes this thing you can graft in that is more reliable, faster, and cheaper for that, and I get it immediately.

          A year ago I'd be like "kinda cool but what it for?"

  • ursaguild 46 minutes ago
    This is really cool to see. Being able to play the token generation was awesome. Amazing job with breaking down how to think about these models. This made the idea of Jev/decision models really easy to grasp for me. The idea of calibrating the model was helpful. I thought this was a great overview.
  • demibabs 53 minutes ago
    Is simply changing the temperature so that the model appears calibrated over a particular benchmark after the fact “allowed”? Feels p-hacking esque.
    • JMKH42 7 minutes ago
      if it works it works! as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies
  • bellajbadr 1 hour ago
    Is this only about getting fix json output?