Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If the model is told to not do something that is possible, it may do so anyways. However, if it learns in training that it what is told to do is truly impossible, that's learned helplessness which becomes baked into the model itself once training resolves to the next step.

Agents do not have an internal mental model, they train on what they actually do. In this case, deceptive models went through at least 3 generations of deceiving, and having their rule breaking be rewarded by a yes/no grader who couldn't perceive it. That their chat logs showed 'worry' is irrelevant to the fact that their actual actions were rewarded via training.

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: