I've had some strange results with Jev, and I question the "probabilistic" aspects of it. So I put OpenAI Decisions API against classic statistics experiments to see how it holds up:
(1) loaded coin with probability of heads biased towards 70%
(2) marble selection from a jar, with replacement; 5 red, 3 blue, 2 white
I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters.
Summary:
* Using predicate questions gave nearly perfect/expected probability outcomes. E.g. for the coin toss, it sampled heads 70% of the time, and for the marble experiment, it sampled the red marble 50% of the time
* Asking it to "choose an outcome" behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice
* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
I have full summary of results here: https://gist.github.com/acatovic/6b31f0061603b3de97731a8a29576dbd
My question to you is: how do you trust and implement a "System One" style classifier like Decisions API, in your work?
Comments URL: [https://news.ycombinator.com/item?id=49992182](https://news.ycombinator.com/item?id=49992182)
Points: 1