> the 24GB card that's been sitting in gaming PCs since 2022.
I'd wager most people have less. In 2022 a 3080 might have 12 GB if you were lucky, 10 if you weren't -- and you paid for the privilege. A current RTX 5080 is only 16GB.
The most recent Steam Hardware Survey[0] lists the most popular VRAM at 16GB in 25.9% of users. Second most popular build is 8GB in 25.3%. >=24GB is ~7%, which is quite a bit higher than I expected.
As RAM prices are so nuts, I picked up a used gaming laptop early this year with an RTX 3070 (i.e. 8GB) to match the specs for my gaming tower that's run everything I'm interested in just fine. That includes recent Unreal 5.x games (Satisfactory, Fortnite, etc.) with high graphics settings on a 4k TV (60hz). There aren't a ton of games that require more than 8GB of GPU ram.
> Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.
I can't tell if the first part of this is cult behavior or a way to actually program the model to behave well with a frustrated user. Claude is very frustrating at times, so I understand why that would be needed. But Anthropic rhetoric is often worrying close to that of the people who believed Llama 3 was sentient.
If I'm doing the math right, that's like 17 t/s? I haven't played with Qwen 3.8 yet, but that seems really slow for a 27B on an M5 Max with sufficient RAM to hold it in memory.
So have you looked at what's happened in the US over the past 10 years?
The US has much further to fall, but it's falling very, very quickly and if there's ever another Democratic president they're going to have to rebuild a lot of the government from scratch.
The unelected bureaucracy was more like the chinese party system. The U.S. has a strong-president model by design: https://avalon.law.yale.edu/18th_century/fed70.asp. The check isn’t supposed to come from unelected bureaucrats, it’s that the strong president is elected every four years. It’s supposed to be a tight feedback loop. Engineers of all people should understand why that’s good.
When the next democrat president gets into office, he or she should do the same thing as Trump: put trusted deputies in charge of various departments and whip them to actually do what people elected the administration to do. That’s how our system is supposed to work. And democratic voters would I’m sure be much happier with the party if they sometimes actually got what they voted for.
Tesla has lost both house battery and car sales in my family -- we're talking hundreds of thousands of dollars -- simply because we don't trust him not to remotely shut off our power/cars for petty political reasons.
Also, if you want true privacy you should run AI models on local hardware. (Guess which country's models dominate SOTA/near SOTA open weights? Yes, it's China, and it's not even close. You can run full-fat DeepSeek locally for (just) under $10K USD.)
Supposedly people are getting ~40 tps decode at Q8 on 2× DGX Spark (higher for Q4) which is what I assume they're suggesting is just under $10K USD. Prefill is just above 1.5k so TTFT is maybe 2 to 4 minutes? I don't have two DGX Sparks myself so I can't confirm and not 100% sure if that number is with or without speculative decoding already in use (if not probably around 60 to 80 if enabled?).
Google also uses ML to train the Waymos, they just have better sensory capabilities. Lidar is much closer to how the human eye functions (inverted, I suppose) than cameras are anyway. But FSD is still really cool, you just have to make sure you don't take your hands off the wheel!