Hacker Newsnew | past | comments | ask | show | jobs | submit | jwpapi's commentslogin

I don’t want it to talk like me, I want it to talk exactly.

I don’t care to look up terms as long as they are correct.


how do you get 2 cgpt pro?

You can have multiple accounts w/ OpenAI, attached to different emails - just log out of one and log into the other.

And fine to do it in same folder same local laptop?

Yep, no problem at all. The only drawback is sessions cannot be shared directly between accounts, so if you're in the middle of something you'll have to do some extra work. To that end I have a handoff skill to persist state to a local markdown and a resume skill to load that state into a new session.

Wasn’t there a website that was tracking that?

Damn that page took down my Chromebook, never happened before..

Text under the "Examples" section...

"Try examples while GPT-2 model is being downloaded (600MB)"

That's a hefty chunk of download and likely compute too.


Would be curious if https://spreadsheets-are-all-you-need.ai/gpt2/ works on your Chromebook. You need to download the weights and then drag and drop them back into the page but maximum compatibility was one of the design goals.

All good I’ll check on my main!

Thanks!

twice...

In the chart they use "Pareto Line", which I think is wrong. Pareto is 20% effort leading to 80% results. Which could be interpreted as models costing 20% having 80% of peak intelligence, but that’s not what it looks like to me.

It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.

I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.


There are two (or more) concepts named after the same person:

- Pareto efficiency/Pareto curves: Basically the convex hull of points along the edge of a graph, indicating the best tradeoff between the axes. This is what the post is talking about.

- Pareto principle: this is the 80/20 rule you're talking about


This is the Pareto Front [1], rather than the Pareto principle. It's the idea that anything that's more intelligent is more expensive and anything that's less expensive is less intelligent.

[1]: https://en.wikipedia.org/wiki/Pareto_front


No, Pareto refers to Pareto efficiency https://en.wikipedia.org/wiki/Pareto_efficiency

What you call "frontier line" is also called "Pareto frontier" https://en.wikipedia.org/wiki/Pareto_front

Your description of it is basically correct though


"Pareto" is many things, but here it does indeed refer to the frontier: https://en.wikipedia.org/wiki/Pareto_front

Thank you guys. I learned something new.

and economically viable

I think it can make a lot of sense to make a set with jev and then train laya on that set and do the rest with laya for saving $

I did that on a little macbook m4 last night on my model of the innate immune system—fine tuning took 15m or so. Just wish it had a larger context window

How about runpod

No need—small and light enough was able to do everything offline, though runpod would work fine, though it‘d be be a quick job

With everybody bashing OP here, how is he supposed to even make money ? It’s n open-source model you self-host right? So he doesnt seem to be just greedy? He might genuinely feel like stolen. I hope he doesnt take the comments personally and is able to find motivation in it.

you may reconsider OP's primary motivation based on another of their HN submissions

https://news.ycombinator.com/item?id=49674396

my hunch is that an Ai has been validating their biases


Okay I understand.. he’s up for the internet fame.

I’ve tried it versus Jev and I got significantly worse decisions. I hosted it on runpod nvidia t4.

I wanted to classify business b2b vs b2c and business model. Am I holding it wrong?


Jev says you should restate state in the question and I tried it:

{ "decision": { "type": "noul", "instructions": "Is the rolled number in state odd?" }, "question": { "type": "noul", "instructions": "Is the number odd?" }, "question-3": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the number odd?" }, "question-4": { "type": "noul", "instructions": "a 6 sided dice rolled a 3 Is the rolled number odd?" } }

=>

decision,0.168,0.83 question,0.141,0.86 question-3,0.029,0.97 question-4,0.021,0.98

so im confused too..

A weakness with numbers?


Breaking news: small language models struggle with math

But didn’t you hear?

> Jev is neither small nor an LLM


It's either a small language model or a large language model (LLM). It's not a generative model, but neither is BERT, which is also a language model.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: