Contrastive Language Models

(contrastive-lm.notion.site)

55 points | by erichocean 4 hours ago

8 comments

  • mugul 2 hours ago
    Very interesting insight on the training process, it's pretty cool to have some experimental justification for why they took these exact steps, what they tried and did not work, etc. Feels a bit less like dark magic.

    However I agree the latency argument doesn't hold much value with Jev because it runs on a remote server. Seeing how many open Jev-like models came out recently it would be much more interesting to have a comparison with them.

  • sdan 29 minutes ago
    I tried running this on a H100 and got 190ms compared to Jev's 170ms. Maybe I set it up wrong?
  • fxwin 50 minutes ago
    I really hope that "System One" won't stick around as a new buzzword simply meaning "fast".
    • TeMPOraL 39 minutes ago
      • fxwin 12 minutes ago
        i'm well aware of the origin and meaning of the term

        > "System 1" is fast, instinctive and emotional

        this implies more than just "fast", which is precisely why i don't like its present usage

        • tancop 4 minutes ago
          Instinctive is a good way to describe it compared to generative LLMs. Jev gives you one instant answer, fast and usually correct but without nuance or any explanation. Human instincts work the same way.
  • r0x0r007 2 hours ago
    what's with that dino run? Jev is slow but it jumps correctly, their model always touches the cactus or whatever it is...I am guessing it doesn't matter? Or does it?
  • eadwu 2 hours ago
    Does the latency even matter?

    You're comparing a local GPU to network hops? Wouldn't be surprised if Jev was actually similar in runtime and their is just a great deal of network latency.

    The evaluation is quite interesting though - I'd actually say the raw answer is correct in the absence of detail and prior knowledge (Who wrote the play Romeo and Juliet).

    • brookman64k 1 hour ago
      I tried TypeSafe’s Jev playground. It outputs the model latency and network latency separately. The model latency was 100-200ms in my tests.
  • in-silico 4 hours ago
    I wonder when work started on this project, and how the public release of Jev played into their timing.
  • 0x4139 54 minutes ago
    [dead]