Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.


A few weeks ago I wrote that I don’t have much of a problem with saying an AI decided, chose or approved something. If an agent made the intermediate decision, that is often the clearest description of what happened. What bothers me more is designing AI to make it seem as though there is someoneContinue reading “Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.”