Hacker Newsnew | past | comments | ask | show | jobs | submit | more KerrAvon's commentslogin

What's amazing is that all of these are true at once. If you allow for some significant slack in what "superintelligence" means.


> the 24GB card that's been sitting in gaming PCs since 2022.

I'd wager most people have less. In 2022 a 3080 might have 12 GB if you were lucky, 10 if you weren't -- and you paid for the privilege. A current RTX 5080 is only 16GB.


The most recent Steam Hardware Survey[0] lists the most popular VRAM at 16GB in 25.9% of users. Second most popular build is 8GB in 25.3%. >=24GB is ~7%, which is quite a bit higher than I expected.

[0] https://store.steampowered.com/hwsurvey/En


As RAM prices are so nuts, I picked up a used gaming laptop early this year with an RTX 3070 (i.e. 8GB) to match the specs for my gaming tower that's run everything I'm interested in just fine. That includes recent Unreal 5.x games (Satisfactory, Fortnite, etc.) with high graphics settings on a 4k TV (60hz). There aren't a ton of games that require more than 8GB of GPU ram.


In 2021 I was considering getting a then-old K80 card because it had the highest VRAM on the market at 24GB and all newer cards had 8 or 12GB.

(Apparently this is because the K80 is two separate GPUs on one card, but I still think it counts if you only have one slot to put it in)


I'm still getting by mostly fine with a 4GB 1650 Super. (2020 vintage.)


> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.

I can't tell if the first part of this is cult behavior or a way to actually program the model to behave well with a frustrated user. Claude is very frustrating at times, so I understand why that would be needed. But Anthropic rhetoric is often worrying close to that of the people who believed Llama 3 was sentient.


If I'm doing the math right, that's like 17 t/s? I haven't played with Qwen 3.8 yet, but that seems really slow for a 27B on an M5 Max with sufficient RAM to hold it in memory.


So have you looked at what's happened in the US over the past 10 years?

The US has much further to fall, but it's falling very, very quickly and if there's ever another Democratic president they're going to have to rebuild a lot of the government from scratch.


I dunno. I'm just glad Congress can barely pass any legislation. What an Executive Order does, another Executive Order can just as easily undo.


The unelected bureaucracy was more like the chinese party system. The U.S. has a strong-president model by design: https://avalon.law.yale.edu/18th_century/fed70.asp. The check isn’t supposed to come from unelected bureaucrats, it’s that the strong president is elected every four years. It’s supposed to be a tight feedback loop. Engineers of all people should understand why that’s good.

When the next democrat president gets into office, he or she should do the same thing as Trump: put trusted deputies in charge of various departments and whip them to actually do what people elected the administration to do. That’s how our system is supposed to work. And democratic voters would I’m sure be much happier with the party if they sometimes actually got what they voted for.


Tesla has lost both house battery and car sales in my family -- we're talking hundreds of thousands of dollars -- simply because we don't trust him not to remotely shut off our power/cars for petty political reasons.

Also, if you want true privacy you should run AI models on local hardware. (Guess which country's models dominate SOTA/near SOTA open weights? Yes, it's China, and it's not even close. You can run full-fat DeepSeek locally for (just) under $10K USD.)


> You can run full-fat DeepSeek locally for (just) under $10K USD.)

Is that price not way off if you want actual decent performance, like at least 30-60 tokens per second and at least >256k context size?


Supposedly people are getting ~40 tps decode at Q8 on 2× DGX Spark (higher for Q4) which is what I assume they're suggesting is just under $10K USD. Prefill is just above 1.5k so TTFT is maybe 2 to 4 minutes? I don't have two DGX Sparks myself so I can't confirm and not 100% sure if that number is with or without speculative decoding already in use (if not probably around 60 to 80 if enabled?).


Funny, the Germans said the same thing a long time ago!


This was the basic theory behind Tesla eventually being able to stop killing Tesla drivers with FSD. It has not worked out.


Google also uses ML to train the Waymos, they just have better sensory capabilities. Lidar is much closer to how the human eye functions (inverted, I suppose) than cameras are anyway. But FSD is still really cool, you just have to make sure you don't take your hands off the wheel!


I think it's obvious that repeatedly noting that the absolute perfection isn't achieved yet is an absolutely useless way of tracking progress. No?

BTW, when was the last time Tesla driver died due to FSD fault?


It has actually, FSD is very good and the latest iteration is end to end now for many trips.


They don't control that. Pricing is controlled by the question of when RAM is once again made of semiconductors instead of unobtainium.


this is not how it's going to go if OpenAI and Anthropic get their way and the US outlaws use of open-weight models


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: