Your AI Is Not Your Colleague
Why the Teammate Metaphor Derails Implementation
A journalist asked us recently whether organizations should think of AI as “a new kind of teammate.” The question reveals a category error that’s costing companies millions in failed implementations.

The uncomfortable truth is that when we anthropomorphize AI—attributing human traits, understanding, or intent to these systems—we import assumptions about relationships and development that don’t apply. Those assumptions lead us to skip the actual work that makes technology adoption succeed.
The Psychology We’re Fighting
The tendency to treat AI as human isn’t irrational—it’s fundamental to how our brains work. Stanford researchers Byron Reeves and Clifford Nass established through extensive experiments that people automatically apply social rules from interpersonal communication to human-computer interaction. We attribute personality traits, follow politeness norms, and respond to gender cues—even when we consciously deny doing so.
Nicholas Epley’s research explains the mechanism: we use ourselves as the default model for understanding other agents, and attributing human characteristics helps us predict behavior and reduce uncertainty. These aren’t flaws to train away. They’re core features of human cognition that conversational AI interfaces actively exploit through human names, persona design, and natural language.
The effects scale with capability. Research shows that the more human-like the interaction, the more automatic the anthropomorphic response—and the more dangerous the consequences. When chatbots are perceived as human-like, users’ privacy concerns decrease, and they become more willing to disclose sensitive information. One study found 34.8% of employee ChatGPT inputs now contain sensitive data, up from 11% in 2023. The Samsung leak—where engineers disclosed confidential source code by pasting it into ChatGPT for debugging—illustrates how conversational interfaces encourage disclosure that employees would never provide through traditional data entry.
NIST now explicitly identifies “inappropriate anthropomorphizing of, overreliance on, or emotional entanglement with GAI systems” as one of twelve key governance risks.
What Goes Wrong
The documented harms cluster around trust, judgment, accountability, and privacy.
Over-trust and automation bias represent the most immediate operational risks. When AI is presented as authoritative and human-like, users develop what researchers call “positivity bias”—a predisposition to trust automated aids over humans even when automation is demonstrably less reliable. Multiple studies confirm that the majority of users accept AI outputs without meaningful scrutiny when systems are framed as authoritative. A Georgetown analysis of Tesla autopilot incidents, Boeing aviation failures, and military air defense mistakes concluded that organizational decisions about how to frame and deploy automation significantly shape outcomes. Human-in-the-loop oversight fails when the human has been primed to defer.
The accountability gap may be the most legally consequential risk. Research by Bigman et al. found that algorithmic discrimination elicits less moral outrage than human discrimination—people are less inclined to blame organizations for AI errors because they “perceive algorithms as lacking prejudicial motivation.” This creates a perverse situation where companies deploying anthropomorphized AI face reduced social pressure for failures while remaining legally exposed.
The Air Canada case established important precedent here. When a passenger received incorrect bereavement fare information from the airline’s chatbot, Air Canada argued the chatbot was a “separate legal entity” for which it bore no responsibility. The tribunal rejected this defense outright: companies cannot disclaim liability for AI chatbot statements by claiming AI is a separate entity. The EU’s revised Product Liability Directive, adopted in 2024, now explicitly includes software and AI, with providers liable for defects causing harm—including psychological harm for the first time.
The Pattern in Failed Implementations
The business consequences are both documented and quantifiable.
Klarna’s AI customer service deployment offers the clearest cautionary tale. The company announced that its AI assistant handled 75% of customer chats—2.3 million conversations monthly—and claimed it “did the work of 700 agents.” CEO Sebastian Siemiatkowski publicly stated AI could “take everyone’s job, including his own.” The result: declining customer satisfaction and mounting complaints about impersonal, inadequate responses, forcing Klarna to reverse course and rehire human agents by May 2025. Siemiatkowski later admitted that “cost unfortunately seems to have been a too predominant evaluation factor... what you end up having is lower quality.”
The Lattice case is equally instructive. In July 2024, the HR technology company announced it would give “digital workers” official employee records—AI agents would be onboarded, assigned managers, and tracked like human employees. Within 48 hours, the announcement generated hundreds of negative comments. Critics noted the company was effectively telling employees it viewed them “as ‘resources’ to be optimized and measured against machines.” Within three days, Lattice scrapped the feature entirely.
The pattern: organizations that frame AI as workforce replacement or human-equivalent colleagues face service degradation, workforce anxiety, and strategic reversals.
What Works Instead
Organizations achieving measurable returns share a consistent approach: they frame AI as augmentation rather than replacement, maintain robust human oversight, and resist anthropomorphic design choices.
McKinsey’s internal platform Lilli demonstrates the model—72% of the firm’s 45,000 employees use it monthly, reclaiming approximately 50,000 labor hours per month in research time. The key: AI is framed as a “knowledge accelerator” rather than a “digital colleague,” with automated policy checks and human oversight at every step. JPMorgan’s COIN system for contract review reportedly handles work that previously required hundreds of thousands of staff hours annually, with employees redeployed to higher-value activities—positioned as productivity enhancement, not workforce replacement.
The contrast with failure cases reveals that framing determines outcomes. This isn’t a communications preference; it’s an implementation variable that predicts success.
Practical Steps for Executives
Audit your language immediately. How does your organization talk about AI internally? If you hear “teammate,” “colleague,” “assistant” (in the human sense), or “partner,” you’ve imported assumptions that don’t serve implementation. Shift toward “tool,” “system,” “platform,” or simply the product name. This isn’t semantic nitpicking—it shapes how employees interact with the technology and how they evaluate its outputs.
Assign human accountability for every AI output category. The teammate metaphor obscures a critical question: when AI is wrong, who owns the consequence? Name specific individuals responsible for reviewing AI-generated content in each use case. Create verification checklists. Document escalation paths. Your legal exposure doesn’t decrease because “the AI did it”—Air Canada learned this the hard way.
Design for appropriate skepticism. Configure interfaces to state “You’re interacting with an AI system” at session start. Use functional names rather than human names. Expose confidence levels explicitly. Require human sign-off before AI outputs are shared with customers or entered into official records. The goal is productive friction—enough resistance that employees engage judgment rather than defer automatically.
Train for accurate mental models. Employees need to understand what AI actually is: pattern recognition systems that predict likely next tokens based on training data. The “stochastic parrots” framework provides useful grounding—large language models stitch together linguistic sequences according to probabilistic patterns, without reference to meaning. There is no comprehension, no intent, no ethical reasoning behind the output. Training should cover what AI cannot do, when to verify outputs independently, and when not to use AI at all.
Measure what matters. Track output quality, error rates, time allocation changes, and decision accuracy—not “adoption” or “satisfaction.” Your AI doesn’t need to be liked. It needs to work. Monthly reviews should examine leading indicators: Are verification steps being followed? Are escalations happening appropriately? Are error rates within acceptable bounds?
Communicate the frame explicitly to your workforce. Employees are anxious about AI. The teammate metaphor amplifies that anxiety by suggesting replacement is the goal. Be direct: AI is a productivity tool that handles routine work so humans can focus on judgment, relationships, and complex problem-solving. This framing reduces resistance and increases appropriate use.
The Implementation Science Lens
Our framework emphasizes problem-type diagnosis before intervention. AI implementation is no exception. Before any integration, executives need to determine whether it is operational (resource constraints, process bottlenecks). Behavioral (skill gaps, performance issues)? Structural (policy barriers, system limitations)? Cultural (values misalignment, broken trust)?
AI positioned as “teammate” muddles this diagnosis. It suggests the intervention is relational when it’s actually technical. It implies gradual trust-building when the need is immediate workflow redesign. It creates expectations about learning and development that don’t map onto how technology adoption actually works.
The organizations that will gain a competitive advantage from AI aren’t the ones building relationships with their tools. They’re the ones systematically implementing, honestly measuring, and resisting the seductive metaphors that obscure the actual work required.
AI is powerful. AI is transformative. AI is not your colleague.
This work synthesizes research from Stanford’s Computers Are Social Actors paradigm (Reeves & Nass), NIST’s 2024 Generative AI Risk Management Framework, Georgetown CSET’s automation bias analysis, and case documentation from the British Columbia Civil Resolution Tribunal (Air Canada v. Moffatt), alongside implementation data from McKinsey, Klarna, and Lattice.
Contact us for a full list of works cited.