Hacker Newsnew | past | comments | ask | show | jobs | submit | otabdeveloper4's commentslogin

Programming hasn't been solved by LLMs. AI chatbots give wrong answers to programming queries most of the time too.

Not most of the time

More often than it's worth most of the time

Not really

I do find them useful when querying like "how todo xyz in abc"


> without first having to solve the problem of effectively sandboxing Bash

"Sandboxing bash" is a problem that has been solved a zillion years ago already. Take your pick of any of the dozens of battle-proven solutions.


Which solution do you recommend?

Bonus points if it's available on both macOS and Linux and doesn't come from a random unmaintained GitHub repository with a note in the README that says "don't run this in production".


"Battle-proven" until an LLM decides it really needs to escape the sandbox you put it in and eventually succeeds.

For personal work, I run Codex in a VM that contains only what's necessary to do software development. Could it escape the VM? Sure, if there's a zero-day in VMWare Workstation.

Yeah, I'm using a pile driver when I really probably just need a hammer, but I've seen too many horror stories, and I don't trust guard rails. Even if there was an option to limit Bash calls to read-only operations, I would be 0% surprised to eventually run into "You're absolutely right! `rm -rf / --no-preserve-root` was a write operation! That's totally on me."


You're confusing the mechanical language of math with math itself.

The map is not the territory, etc. If math was just an elaborate linguistic Glass Bead Game then we wouldn't be funding it. The intuition is that the surface rules of math help uncover the underlying structure of reality.


AI irritates people. This is fine if you're making political cartoons or racist memes, but not so fine for ad copy.

A lot of posters or arts in general are to be seen once and never again. If you put pre-AI arts together in an exhibition like at the top of the article are you sure it wouldn't irritate people?

I walk past countless of such arts everyday, none of them registered after even 15 mins. Maybe you have better taste, but don't try to project your opinion as that of the majority.


Read what I said again. AI art falls into the uncanny valley range and solicits disgust, ire or suprise in most people.

Sometimes that's okay, but probably not in ad copy.


> Since they are invisibly inserted

They're not, all destructors are explicit. Seems like a skill issue on your end.


> Seems like a skill issue on your end.

You're talking to walter bright, the guy who wrote the digital mars C++ compiler


> appeal to authority

Okay.


Since AFAIK I'm still the only person to write a correct C++ (C++98) compiler from preprocessor to object file, I know all about destructors.

Here's a fun one for your amusement:

    foo(a, b, c);
The parameters are pass by value. a, b and c are objects that have destructors. Have a look at the code generated for that.

It is nice that the compiler does the dirty work for you, but the various paths with exceptions and recovery with invisible code may not be well tested.


AI slop isn't "new proofs of theorems", no more than Claude Code slop is "new software products".

> you're just not using the latest model, bro

Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.

P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.


> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.


The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price.

It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.


I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks.

By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.


No argument that DeepSeek is much cheaper, but you certainly get what you pay for.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.


It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...

Where did he claim otherwise?

Take your meds.


Impressive how feverishly you defend the steaming pile of shit that Gemini is.

> everything you believe is wrong, here's our god, you must follow and believe in him or you're going to hell

"Democracy" and "human rights" is basically that but on steroids. Surely you're not a cultural relativist that simply going to allow uncontacted people to be racist and misogynist willy-nilly?


Those things aren't at all equivalent, so I'm not sure I understand your point.

The very first thing Garage did after I installed it on a test three-node cluster is get corrupted and lose files.

No thanks.


> Fact: There was no world war.

Akshually there is, and WW3 has been going on since 2010. (Mostly in places that aren't Europe.)


Some people believe WW2 started many years earlier, but most historians don't put the start date at the conflicts going on before the war became intercontinental with sides working together on a global scale.

War has still been going on since 1945, much earlier if not always, so akshually WW2 never ended or is just how it's always been? (mostly in places that are not "first world" / The West)


There's a simple and very logical definition of what a "world war" is - it's when international law completely breaks down, and the only way to get back to a semblance of order in international relations is to have the victors of the ensuing shitstorm impose a new set of rules in some sort of grand finale pact.

1945 was clearly that. 1920 with the League of Nations attempted to be that, but failed.

Whatever happens after the current turbulence we can be sure that there won't be a UN and NATO and a World Bank at the end.

P.S. The first modern world war was the Thirty Year's War.


> Whatever happens after the current turbulence we can be sure

Have you put your money behind this on future gambling sites? How can you be so sure?


Because the UN, NATO, etc., are clearly not working now. They're not gonna be fixed to work as designed because that's quite impossible in today's world.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: