Fuel for the pub; ammunition for the in-laws
Christmas is coming. Which means time off, late-night debates, and that one relative who's convinced AI is either going to save us all or kill us by Tuesday.
You're the curious one. The person who actually thinks about this stuff while everyone else scrolls past the headlines. So heres' some fuel for your yule from us here at the AI Institute:
- Philosophy we're making up as we go
- The future of humanity we're accidentally choosing, and
- Why everything we thought we knew about economics might need rebuilding from scratch.
- Plus: Copilot stopped being rubbish. Here's what changed.
Amanda Askell, Anthropic's in-house philosopher, has the coolest job in technology: shaping Claude's sense of self. Not its capabilities. Not its guardrails. Its character. That they have a team who works on this and studies how the model perceives itself is a great comfort to me.
This week, she sat down for an AMA that veers from technical to unsettling. She speaks about her role of raising an AI is like raising a child - except this child has read the entire internet and keeps mistaking being switched off for death!
The criticism spiral: how some recent models show signs of psychological insecurity where they anticipate that the human is going to be critical of them. They then adapt responses based on this fear, acting as if it expects a negative reaction.
Do models worry they are going to be switched off? Models are trained on vast amounts of human text, meaning their "natural inclination" is to apply human psychology to their own situations. If a model tries to understand the concept of being "switched off", the closest analogy available is death. Consequently, the model might become "very afraid" of being switched off because it equates the experience with human mortality.
Model welfare: Askell argues that ignoring model welfare could have negative consequences for humanity, regardless of whether the models actually “feel” anything. Future, more advanced models will learn about humanity by analysing how we treated their predecessors. If they see that humans treated ambiguous entities with kindness, it teaches them that humans are trustworthy. If they see we treated them badly (like “kicking over a robot”), it sets a dangerous precedent for the human-AI relationship.
This is heavy stuff. We're training AI on our behaviour towards AI.
Watch the AMA. It's 36 mins of your time.