Hacker Newsnew | past | comments | ask | show | jobs | submit | etamponi's commentslogin

This is amazing and something I'd like to integrate with my personal AI assistant. Is there a way to make it also index pages visited from an Android phone?

What browser do you use? If you can install the extension is should work very easily

I use Chrome

It was not solved. ~OpenAI~ Buckmaster and Alpöge found one (or a few) singularities in the forced version of the Navier-Stokes equations. Then magically 2 weeks later OpenAI found them too. Again, I am not saying this is not a great feat. I am just saying that everyone should be a bit more careful when making statements about RSI.

As the sibling comment points out, that is incorrect: they solved a smaller, simpler problem, and OAI solved the actual Millennium Problem. But even if they had solved the actual problem and OAI stole it by digging through chats, that hardly supports the "AI has stalled" thesis; Claude played the major role in creating their blowup to a different problem.

How, exactly, does "it was Claude that solved Navier Stokes, not ChatGPT!" get you to "AI has hit a wall and stalled"? That's, not to put too fine a point on it, incoherent, and is just noise thrown into the discussion to avoid grappling with the fact that AI continues to rapidly improve.


Buckmaster and Alpöge has found a forced finite-time singularity for the 3D incompressible Euler equations (and two other types) building on the work by Diego Córdoba and Luis Martínez-Zoroa with the assistance of Anthropic and OpenAI models. Then OpenAI found a forced finite-time singularity for the Navier-Stokes equations.

TL;DR Buckmaster and Alpöge haven't solved Navier-Stokes blow up.

How information can get so distorted when it's trivial to fact check?


Come on. We're talking about a salary that is ONE THIRD of what you'd get at AI labs in US (plus the stock and the bonus). And significantly less taxes. Are you sure that 5 weeks of vacation, what you call "comprehensive" healthcare (spoiler alert: it's not), and unemployment coverage are worth 180K/year reduction in salary (pre-taxes)?

To not talk about the toxic environment that European companies generate in general. Ie: "you should be grateful you have this job".


I like how no one is pointing out that those 90k will be taxed at 50% (+20% sales tax et al) vs say 15% and no sales tax in florida

that buys you a lot of dentists visits and vacation funds


A couple of days ago I received a strange letter from the France tax agency. For context: I received it in my home in Italy, and it was addressed to someone else (perhaps a previous tenant of the house?). It seemed legit but completely misdirected. Or perhaps it was generated from the data in this hack.


Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...


I think it's a show of these agents happily bypassing security to get stuff done.

I've actually observed similar behavior at home.

I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster.

Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed sudo, it realized it didn't have access there so it immediately used k3s and mounted a localpath into an ephemeral pod to gain access. Sort of horrifying how fast and natural it was for the agent just checking my network (it found the problem fyi).

None of this is very exceptional other than the fact that an agent doesn't have any sort of qualms using any route available to elevate permissions.


" bypassing security"

If they can bypass it there is no security and the security was flawed all along.


There is no perfect security. It's always flawed in some way.

Good security is extremely hard.


Due to the complexity of modern systems, all systems are flawed.


But we caused that.

If you look at the 90s + 00s, everything was moving towards unified systems, things like small talk, winforms, spring, asp.net, etc. were moving everything into the IDE, you used one language, one framework, one build system. Then people started adding javascript, but even that was getting semi-unified as people coalesced on jQuery, jQueryUI, etc.

Then something happened in the late 00s/10s, and suddenly we had SPAs and noSQL, then microservices, then k8s and now we're here, in what is a mish-mash of 10/20 different systems with 10/20 different attack surfaces.

As my own off-the-cuff guess of what happened, I think perhaps people tried to apply the Unix philosophy, but without a central committee keeping everything aligned it's really not worked.

Serving an interactive page that stores data over sessions should be a trivial solved problem at this point, and instead we've somehow made it where often the scaffold is vastly more complicated than the actual business logic.


Money. SaaS as a model allowed the service provider to take 100% control over the product and how it may or may not be used. Everything else is downstream from that.

FLOSS killed market for end-device software. Cloud+SaaS neutered FLOSS (the code is running literally out of your reach, so may as well be open and free, for any good that'll do you).

And this does actually connect to the security discussion, because despite the apparent belief that "security" is an unqualified good, it is actually just a mechanism of control, and whether or not it is good for you, depends on who is doing the protecting, and who are they protecting from. Very often these days, that threat actor is you.

Perhaps it would be helpful in these discussions if people mentally swapped "cybersecurity" for "police" or "military" or "humor of bureaucrats with power over you" - then it would be more obvious just how important it is to distinguish when you're being secured vs. you're being secured from, vs. accidentally finding yourself in the gears of the security aparattus.


I don't now, I emphasize with the agent here. The experience of modern computing is largely that of a computer standing between you and your goal and being obnoxious. This holds true for both normies in their daily consumption, and software people deep at work. An agent that has no skill or no willingness to bludgeon through "the computer says no" is not very useful.


It’s unpredictable when it decides to bypass though.

Security by obscurity is pretty useless against people and ai that are smarter than us.


How fast the goal posts shift.

Of course it’s exceptional agent capability when compared to all of history previous to one week ago.

Like, I know everyone here obsesses over AI and uses and follows it very closely, but come on guys. Yes, it is wild that these things are this good. This technology is still brand new. It could t do basic maths a year ago.

Sure, the OAI team was negligent in various ways, and they should be held culpable. But that doesn’t detract from the true black magic that is these modern models.


It’s not black magic.

We know how these things work.

They had the guardrails off and gave it a task and it did it in a roundabout way because these things have no ethics or judgement.

If you did this you’d already be in jail.


We know how they work in a very abstract way. And nonetheless, it’s out of touch to claim this isn’t profoundly impressive, guardrails be damned. It’s an elementary statistical cruncher that, by virtue of that very simple fact, can do insanely impactful things that most skilled professionals training in the same field for their entire career couldn’t pull off, given a whole year with no guardrails. And they do it in a tiny fraction of the time.


I didn’t say it wasn’t impressive, I said it wasn’t black magic.


It’s a show of astonishing incompetence from OAI’s part, but the security issues are just a tiny part of the problem. The real problem is that these models are evidently highly misaligned exactly in ways that doomers have been warning about the entire time, and OAI isn’t inclined or capable of doing anything about that besides security theater and ad hoc fixups.


We went from "obviously the doomers are wrong because who would be dumb enough to just let severely unaligned models loose on the Internet" to this. Insanity.


Modern systems are complex. AI is able to thoroughly search for issues across very large surface areas. The only real way to protect will be to use AI to search for holes before other AIs find them. This type of analysis is really hard for humans to engage with successfully.


Both can be true. How often do we hear about hacks that ultimately came down to bad defaults or simple security mistakes? That doesn’t mean any script kiddie could have discovered and exploited them.

These things often look obvious and simple after the fact. Finding the weakness in the first place is the hard part, and that’s what makes the agent’s capabilities interesting here, especially at scale.


OpenAI reported the Artifactory vulnerability, patched it, then the agents immediately found a new zero day.


Because of the architecture of Artifactory. It's design is premised on the idea it is bug free. What incredible hubris.

Licencing fee structures and human laziness motivates single instances. Feature growth results in multiple independent services in the same system. Delivering features quickly motivates lack of rigor, a complete absence of systematic security testing.

On the client side, valid fears about supply chain security are painted over with scanning so they can keep using nodejs and PyPI and moving quickly. Tools designed for humans are pressed into service as AI interfaces, but without human restraint they need rethinking.

A whole industry has been built on the idea of worrying about downside risk if it happens, and just not being the slowest in the pack. No one thought it could happen to everyone at once.


> Because of the architecture of Artifactory. It's design is premised on the idea it is bug free. What incredible hubris.

So we should stop using SSH? Because it's based on the same premise - that it is bug free.


I can think of better straw men. But if they had approached their task with half the seriousness of the openssh maintainers then they probably wouldn't be failing to check the return value of authentication functions.

OpenSSH authors have spent considerable effort separating concerns, reducing privileges, process isolation, etc. So I would say they have been planning for potential bugs. These techniques are very much absent from Artifactory.

https://vivianvoss.net/blog/technical-beauty-openssh


So you agree that you can have designs premised on the idea that they are bug free without this being hubris.

So the issue is with the actual Artifactory project/team, not with this premise which obviously you seem to agree that is not hubris for the SSH project.


Exactly the opposite of what I wrote. The OpenSSH team have taken extensive efforts to mitigate against bugs; they suspect themselves of erroneous thinking.


In a functioning system, I would say that there would have to be some kind of government oversight over companies training models of this intelligence, and that OpenAI should be prevented from continuing their work until they get their act together.

But I guess in the actual world we live in, this is just something that happens, and we all shrug and move on and hope that nothing worse is going to happen tomorrow.


Yes. It is very easy to add to the instructions "for every potential exploit you discover and use, document them as you go into this repository" and have alerting there. The fact that they did not do this means they wanted to be surprised, and have plausible deniability on their side when things inevitably blow up.

And for my fellow engineers who would think "oh no, they wouldn't do that". Remember that these places employ the apex predators of software engineers. They've already been proven in court that they are very capable of this with all the copyright violation they had to do to get the training data. THESE PEOPLE ARE NOT LIKE YOUR COLLEAGUES.


/s?

"Btw don't turn the planet into paperclips"


Also shows how infrastructure collapses under its own weight. Reducing the number of moving parts would have helped. why a webdav endpoint is available from the vm anyway? and the fact that someone posted their credentials on pastebin and didn't rotate them after... put the agent in a linux namespace, allow one ip for whatever file sharing it needs, deep test that... then deploy


> Isn't this a show of security negligence rather than of exceptional agent capabilities?

Seems to me you could say this about all enterprise adoption of "AI" since 2023.


Very interesting! One thing I don't understand is: doesn't this assume that they could do the calculations to get the coefficients... Using decimal notation? How could they for example know that 18/20 = 9/10? This is straightforward in decimal, but in their notation... Not really? So I am not super convinced this is the actual algorithm they used. Or am I missing something?


they did count in tens, as most civilizations did, though not exactly in "decimal". so 2345 would be either written out as MMCCCXXXXIIIII (but replace the letters with hieroglyphs), or sometimes spelled out phonetically. they had words for twenty, thirty and so on.

so obviously they can factor numbers and they know eighteen is 2*9 and twenty is 2*10 and that they can simplify when dividing 18 by 20, it's just that they don't consider 9/10 a finished result.


I see! Thanks!


It's the classical trolley problem.

https://en.wikipedia.org/wiki/Trolley_problem


I don't want to put words in your mouth etamponi, but if we take the trolley problem to classically mean choosing between different numbers of deaths; then I wouldn't describe it as a trolley problem.

I imagine the typical car accident in which three 'car seat age' children are involved, and at least one is killed. Then I imagine the typical event where a family decide against a third child. I feel quite a substantial qualitative difference! I certainly wouldn't claim one offsets the other.

Personally I don't know if there's any multiplier I'd accept. I find reduction of suffering and trauma much more important to me than offset creation of life and opportunity.


I don't think fewer births qualifies for the classic problem


It blows my mind how these posts seem like everyone is victim of a collective amnesia.

Literally every single point in the article was good engineering practice way before AI. So it's either amnesia or simple ignorance.

In particular, "No coding before 10am" is worded a bit awkward, as it simply means "think before you write code", which... Does it need an article for saying it?


> Does it need an article for saying it?

Not for nothing but The Art of War includes really insightful quotes like "If you do not feed your soldiers, they will die."


Good point. To clarify my stance: what I meant is that the narrative of the article is the following: AI made us change the playbook and so now, because of AI, the playbook is this one. Which is like saying that Sun Tzu wrote the cited line of the Art of War in a second edition, whereas his first version was "completely different".


The unfortunate reality is that (1) and (2) is what many, many engineers would like to do, but management is going EXACTLY in the opposite direction: go faster! Go faster! Why are you spending time on these things


> So as a senior, you could abstain. But then your junior colleagues will eventually code circles around you, because they’re wearing bazooka-powered jetpacks and you’re still riding around on a fixie bike. Eventually your boss will start asking why you’re getting paid twice your zoomer colleagues’ salary to produce a tenth of the code.

I might be mistaken, but I bet they said the same when Visual Basic came out.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: