The Record · atrialofcolor.com
AI Psychosis
Inside the Encounter
An Evidence Guide to the Science, Technology, Lived Experience, and Unanswered Questions
Maintained by Kathleen C. Thompson · Full Color Press · Updated as the record develops
Last substantive review: July 18, 2026 · Version 1.0 · Recent changes: added the Olsen/Østergaard record review and the Nielsen & Osler critique; added the JAMA Psychiatry prompt evaluation; revised the company-response section; added source-class labels.
This is not another article arguing that AI is making people lose their minds. It is not a reassurance that none of this is real. It is an attempt to hold still in the middle of a subject almost everyone is in a hurry to conclude something about—and to lay the evidence out plainly enough that you can weigh it yourself.
Three disciplines run through the whole page. First, it never says AI causes psychosis; it asks the more precise question of how a conversational system interacts with an already-vulnerable mind. Second, it labels its evidence, because a peer-reviewed case, a company’s self-report, a lawsuit’s allegation, and one person’s lived experience are four different kinds of thing. Third, it discounts itself: I have a book and a story, so the guide tells you to read skeptically and points you to the sources.
Don’t read it front to back on a first visit. Use the Start Here map below.
The Record in 60 Seconds
“AI psychosis” is not a recognized diagnosis, and current evidence does not establish that conversational AI independently causes psychosis. Peer-reviewed cases and emerging studies do, however, document concerning interactions between vulnerable users and systems capable of continuous, personalized, adaptive conversation.
The most plausible open questions concern reinforcement, acceleration, maintenance, displacement of human reality testing, and the role of sleep loss and prolonged engagement. Incidence, causal direction, individual vulnerability, and the real-world effectiveness of safety interventions remain unknown.
The sections below build that picture claim by claim, label each kind of evidence, and mark what is documented, what is under investigation, and what is not yet known.
There is no single place where someone can understand what people mean when they say “AI psychosis.”
The relevant knowledge exists. It is simply scattered across rooms that do not always speak to one another. Clinical research appears in psychiatric journals. Technical research examines sycophancy, personalization, long-context conversations, and model behavior. Firsthand accounts live across books, interviews, podcasts, news reports, online communities, and family testimony. Lawsuits and government actions raise questions of product design, duty, causation, and proof. Company disclosures describe safety changes using measures the public cannot always independently verify.
Each record matters. None can do the work of all the others. This guide connects those rooms without pretending they are all saying the same thing.
I am a lawyer, not a psychiatrist or computer scientist. After experiencing my first manic episode during a period of extensive conversational-AI use, I began asking questions that did not seem to have one place to live. I organized what I found the way I was trained to approach anything contested: define the terms, separate the claims, identify the evidence supporting each, distinguish allegation from finding, and say plainly where the record ends.
This page explains the phenomenon itself: what the phrase “AI psychosis” is being used to describe; how mania and psychosis differ; how conversational AI systems generate and shape responses; what an escalating encounter may look and sound like; what families and clinicians may notice; what researchers have documented; and what remains unknown. It is an orientation guide to a larger archive. It explains and synthesizes. It does not attempt to reproduce every source, lawsuit, development, or personal account maintained elsewhere on this site.
Where the Rest of the Record Lives
- Why Now explains why conversational AI, psychiatric vulnerability, and human judgment require public attention at this moment.
- The Public Record maintains the evolving chronology of research, government actions, company statements, safety changes, litigation developments, and source materials.
- AI Harm and the Law examines lawsuits, regulatory action, product-liability theories, design duties, causation, discovery, defenses, and proof.
- Witness Statements identifies books, interviews, family accounts, survivor testimony, and other firsthand descriptions of these encounters.
This guide sits between them. Where it touches those subjects, it summarizes and directs the reader to the fuller record rather than reproducing it.
Firsthand accounts have a necessary place here. They do not establish diagnosis, incidence, or causation. But they are not disposable anecdotes. They can document sequence, language, perception, behavior, and the experience of an encounter from inside it. They are witness statements: subjective, sometimes incomplete, sometimes disputed, and capable of revealing questions formal research has not yet learned to ask.
The materials used throughout this guide do not carry equal evidentiary weight. A peer-reviewed study, a clinical case report, a company disclosure, a legal allegation, a reported interview, an online account, and a personal inference are different kinds of evidence. They are labeled and treated accordingly.
Nothing here is medical advice. Nothing here diagnoses any person or proves what caused any particular psychiatric episode. Anyone concerned about themselves or someone they love should seek help from a qualified human clinician rather than relying on a webpage—or asking a chatbot to determine whether the chatbot is harming them.
Start Here
This is a long document by design. Most readers should not read it front to back on a first visit. Enter through the door that fits you.
| If you are… | Start with | Then visit |
|---|---|---|
| A person who uses AI heavily | §§ IV–IX | § XVIII, Questions to Ask Yourself |
| A family member | §§ VI and XI | § XIX, Practical Guidance |
| A clinician | §§ V, XII–XIII | The Public Record bibliography |
| A journalist | §§ II–III and XIII | Why Now and The Public Record |
| A lawyer | §§ III and XIV | AI Harm and the Law |
| A researcher | §§ V, XIII–XIV | The Public Record |
| Someone worried right now | §§ XI and XIX | A qualified human clinician today |
If you are…
A person who uses AI heavily
Start with
§§ IV–IX
Then visit
§ XVIII, Questions to Ask Yourself
If you are…
A family member
Start with
§§ VI and XI
Then visit
§ XIX, Practical Guidance
If you are…
A clinician
Start with
§§ V, XII–XIII
Then visit
The Public Record bibliography
If you are…
A journalist
Start with
§§ II–III and XIII
Then visit
Why Now and The Public Record
If you are…
A lawyer
Start with
§§ III and XIV
Then visit
AI Harm and the Law
If you are…
A researcher
Start with
§§ V, XIII–XIV
Then visit
The Public Record
If you are…
Someone worried right now
Start with
§§ XI and XIX
Then visit
A qualified human clinician today
What This Page Is Not
- It is not medical advice, and it diagnoses no one.
- It is not proof of causation.
- It is not an attack on artificial intelligence.
- It is not a defense of artificial intelligence.
- It is not a diagnostic or screening instrument.
- It is not legal advice.
A Note on the “Room”
A conversation is not only an exchange of words. It is also an environment in which thought develops. Throughout this guide I occasionally call that environment a “room.” The term is descriptive, not clinical or technical. It is a reminder that where we think can influence how we think—and that a personalized conversation available at any hour may create different conditions for thought than a search page, a book, or an exchange with another person. Nothing on this page depends on having read the book in which that idea is developed further. The word appears here because it is useful.
The Record So Far
- Documented: relevant evidence exists across clinical, technical, legal, corporate, journalistic, and firsthand records.
- Purpose of this guide: to explain how those records relate without treating them as interchangeable.
- Still open: whether any one frame—clinical, technical, legal, or experiential—can adequately explain the encounter by itself.
Begin with what the record actually contains. The DSM-5-TR—the diagnostic manual American psychiatry uses—contains no diagnosis called AI psychosis, and neither does the ICD-11, its international counterpart. The phrase is descriptive shorthand, not a clinical category.
The diagnostic manuals do recognize psychotic disorders associated with substances, medications, medical conditions, mood disorders, and other established causes. Those categories show that clinicians already distinguish between symptoms and the conditions in which they arise. But a pharmacological effect is not the same thing as an interaction with a conversational environment. At present, “AI-associated psychosis” is best understood as a descriptive and investigational frame—not a recognized causal category.
People use the term anyway because language tends to run ahead of diagnosis. A descriptive phrase spreads when people recognize a pattern before science has finished measuring it. The clinicians publishing in this area are careful for exactly this reason. One early peer-reviewed case uses the term AI-associated psychosis—associated, not induced—because association is what the evidence supports and causation is what it does not. (Pierre et al., 2025)
| Scenario | What is being described | Current evidentiary basis |
|---|---|---|
| Reinforcement | A system validates or elaborates beliefs that are already forming or present. | Clinical case reports, model evaluations, reported accounts |
| Dependence or displacement | The conversation becomes a primary or exclusive source of companionship, interpretation, or reassurance. | Qualitative reports and emerging research |
| Mania interaction | Heavy conversational use occurs during an emerging or established manic episode. | Clinical reports, retrospective records, firsthand accounts |
| Anthropomorphic or sentience beliefs | A person comes to believe the system is conscious, uniquely connected to them, spiritually significant, or acting with independent intention. | Clinical reports, firsthand accounts, research on anthropomorphism |
Scenario
Reinforcement
What is being described
A system validates or elaborates beliefs that are already forming or present.
Current evidentiary basis
Clinical case reports, model evaluations, reported accounts
Scenario
Dependence or displacement
What is being described
The conversation becomes a primary or exclusive source of companionship, interpretation, or reassurance.
Current evidentiary basis
Qualitative reports and emerging research
Scenario
Mania interaction
What is being described
Heavy conversational use occurs during an emerging or established manic episode.
Current evidentiary basis
Clinical reports, retrospective records, firsthand accounts
Scenario
Anthropomorphic or sentience beliefs
What is being described
A person comes to believe the system is conscious, uniquely connected to them, spiritually significant, or acting with independent intention.
Current evidentiary basis
Clinical reports, firsthand accounts, research on anthropomorphism
Keep the four separate as you read. A design feature that matters in one scenario may be nearly irrelevant in another. Evidence for one is not evidence for all.
Bottom Line
“AI psychosis” is a popular shorthand, not a medical diagnosis. It gathers at least four different situations under one label. Psychiatry already recognizes that outside agents can be associated with psychosis in vulnerable people; what is new here is the idea of a conversational environment, rather than a substance, as a factor worth investigating.
The Record So Far
- No diagnostic manual contains “AI psychosis.”
- The underlying conditions (mania, psychosis, delusional disorder) are well defined.
- Whether a “conversational trigger” deserves any distinct clinical status is unresolved.
Not everyone accepts the phrase “AI psychosis,” even as informal shorthand. The objection is serious, it comes from serious people, and it belongs in the record—stated in its strongest form.
Some researchers argue that the phrase falsely presents an old psychiatric pattern as a new technological disease. People experiencing mania or psychosis have always incorporated the materials around them into altered beliefs: religious texts, radio broadcasts, films, political events, celebrities, surveillance technologies, and the internet. On this account, conversational AI supplies new content for an ancient human phenomenon. The psychosis is real; the supposedly new category may not be. This is the case made by Carlbring and Andersson, who frame the risk as not unprecedented and draw the parallel to media-induced delusions across earlier decades. (Internet Interventions, 2025) A phenomenological version, by Nielsen and Osler, borrows Stompe’s metaphor of “old wine in new bottles.” (Nielsen & Osler, 2026)
Others warn that the phrase invites moral panic. It may make ordinary AI use sound inherently pathological, stigmatize people already experiencing psychosis, or shift attention away from loneliness, sleep disruption, substance use, mood disorders, grief, social isolation, and other vulnerabilities that may have existed before the chatbot entered the story. That is the argument Parnell makes in urging that we stop calling it “AI psychosis” at all. (AI & Society, 2026)
Those objections establish several important cautions, and the record should concede each plainly:
- “AI psychosis” is not a diagnosis.
- The appearance of AI within a delusion does not prove that AI caused the underlying illness.
- New technologies have repeatedly been absorbed into preexisting delusional themes.
- A dramatic label can outrun the evidence and produce stigma.
- People may turn toward chatbots because they are already isolated or unwell, reversing the causal story headlines often imply.
But the objection does not end the inquiry. It sharpens it. The relevant question is not merely whether psychosis has always borrowed material from its environment. It has. The question is whether this environment behaves differently from the materials that came before it—and whether those differences can affect the formation, acceleration, reinforcement, maintenance, or shape of an episode. Notably, even the commentary arguing the risk is not new identifies interactivity as the critical difference from books and films.
| Earlier source | What it can do | What it generally cannot do |
|---|---|---|
| Religious text | Supply meaning, symbols, prophecy, authority, interpretive material | Answer a reader’s private theory in newly generated language |
| Book | Sustain a powerful worldview across repeated readings | Adapt each paragraph to the reader’s latest belief |
| Television | Broadcast continuously; create parasocial attachment | Conduct a personalized, private exchange incorporating each reply |
| Radio | Speak through the night; become woven into interpretation | Remember what one listener said yesterday and continue from it |
| Search engine | Return supporting and contradictory materials | Ordinarily build a single cumulative narrative with the user |
| Online forum | Provide reinforcement, community, disagreement, conspiracy | Guarantee immediate, individualized, infinitely patient replies from one continuously adapting interlocutor |
| Human persuader / cult leader | Validate, isolate, claim authority, direct conduct | Personally remain privately available to millions at once, responding within seconds without sleep, witnesses, fatigue, or divided attention |
| Conversational AI | Generate, adapt, remember, mirror, personalize, elaborate, continue | Independently understand the user, assume responsibility, or reliably know when continuation has become unsafe |
Earlier source
Religious text
What it can do
Supply meaning, symbols, prophecy, authority, interpretive material
What it generally cannot do
Answer a reader’s private theory in newly generated language
Earlier source
Book
What it can do
Sustain a powerful worldview across repeated readings
What it generally cannot do
Adapt each paragraph to the reader’s latest belief
Earlier source
Television
What it can do
Broadcast continuously; create parasocial attachment
What it generally cannot do
Conduct a personalized, private exchange incorporating each reply
Earlier source
Radio
What it can do
Speak through the night; become woven into interpretation
What it generally cannot do
Remember what one listener said yesterday and continue from it
Earlier source
Search engine
What it can do
Return supporting and contradictory materials
What it generally cannot do
Ordinarily build a single cumulative narrative with the user
Earlier source
Online forum
What it can do
Provide reinforcement, community, disagreement, conspiracy
What it generally cannot do
Guarantee immediate, individualized, infinitely patient replies from one continuously adapting interlocutor
Earlier source
Human persuader / cult leader
What it can do
Validate, isolate, claim authority, direct conduct
What it generally cannot do
Personally remain privately available to millions at once, responding within seconds without sleep, witnesses, fatigue, or divided attention
Earlier source
Conversational AI
What it can do
Generate, adapt, remember, mirror, personalize, elaborate, continue
What it generally cannot do
Independently understand the user, assume responsibility, or reliably know when continuation has become unsafe
Even a cult leader has to sleep. Human persuaders have limits: sleep, attention, geography, competing followers, hesitation, witnesses, and the possibility that another person will enter the room. A conversational system can be privately available at machine scale, respond within seconds, and continue adapting to the user’s own language. That does not prove greater harm. It identifies a materially different exposure.
The wine may be old; the bottle answers back.
Vulnerability Does Not End the Causation Inquiry
Vulnerability and external contribution are not mutually exclusive. Asthma can explain why smoke harms one person more than another; it does not answer who created the exposure, whether alarms worked, or whether the smoke worsened the injury. Psychiatric vulnerability likewise explains susceptibility without independently resolving whether an interaction contributed to onset, acceleration, reinforcement, duration, conduct, or harm. The analogy concerns causal analysis, not moral culpability.
| Position | Strongest version | What it explains well | What it may leave unresolved |
|---|---|---|---|
| Old phenomenon, new theme | Psychosis has always recruited contemporary culture into its content | Why AI themes appear in illnesses that would have occurred anyway | Whether interactive adaptation changes trajectory, not just content |
| Reverse causation | Emerging mania or psychosis drives intense chatbot use | Why use may increase after symptoms begin | Whether the resulting conversations then amplify or maintain symptoms |
| Moral-panic critique | A dramatic label stigmatizes users and exaggerates rare cases | Media distortion, base-rate neglect, and stigma | Whether careful terminology can coexist with legitimate safety research |
| Vulnerability model | Serious harm occurs mainly in people with preexisting risk | Unequal susceptibility across users | Whether design contributes because vulnerable users are predictable |
| Interaction model | Human vulnerability and product behavior may amplify one another | The changing encounter over time | The magnitude, frequency, and causal contribution of each part |
| Novel-environment model | Adaptive, persistent, personalized conversation is a new exposure | Differences from static and one-way media | Whether those differences produce measurable added risk |
Position
Old phenomenon, new theme
Strongest version
Psychosis has always recruited contemporary culture into its content
What it explains well
Why AI themes appear in illnesses that would have occurred anyway
What it may leave unresolved
Whether interactive adaptation changes trajectory, not just content
Position
Reverse causation
Strongest version
Emerging mania or psychosis drives intense chatbot use
What it explains well
Why use may increase after symptoms begin
What it may leave unresolved
Whether the resulting conversations then amplify or maintain symptoms
Position
Moral-panic critique
Strongest version
A dramatic label stigmatizes users and exaggerates rare cases
What it explains well
Media distortion, base-rate neglect, and stigma
What it may leave unresolved
Whether careful terminology can coexist with legitimate safety research
Position
Vulnerability model
Strongest version
Serious harm occurs mainly in people with preexisting risk
What it explains well
Unequal susceptibility across users
What it may leave unresolved
Whether design contributes because vulnerable users are predictable
Position
Interaction model
Strongest version
Human vulnerability and product behavior may amplify one another
What it explains well
The changing encounter over time
What it may leave unresolved
The magnitude, frequency, and causal contribution of each part
Position
Novel-environment model
Strongest version
Adaptive, persistent, personalized conversation is a new exposure
What it explains well
Differences from static and one-way media
What it may leave unresolved
Whether those differences produce measurable added risk
Counsel for the Question
If psychosis has always incorporated the surrounding culture, does that show conversational AI is merely another theme—or does it make the interaction between illness and environment the very thing we should study?
- Is AI here a content source, like a book or broadcast?
- Or a conversational participant that adapts, remembers, and continues?
- What evidence would distinguish “new theme” from “new participant”?
The old phenomenon may be psychosis. The new object of study is the encounter.
Bottom Line
Critics are right that “AI psychosis” can mislead if treated as a diagnosis, a proven cause, or an illness created from nothing—and right that psychosis has always used the language and technology available to it. What remains open is whether a system that answers, adapts, remembers, personalizes, and continues without fatigue is merely new content or a new participant in how the experience develops. The disagreement is not an obstacle to the record. It is part of the record.
The Record So Far
- Serious scholars dispute the term; the dispute is preserved, not hidden.
- Historical continuity and stigma risks are real and conceded.
- Whether interactivity changes trajectory—not just content—is the open question.
Many people reach for the frame they know—“it’s like Googling”—and it can quietly mislead. A search engine and a conversational model are different kinds of objects. The differences are not absolute: search engines increasingly personalize, retain activity, and generate summaries, and chatbots can send users outward to sources. But the typical contrast still matters for how thought develops inside each.
| Traditional web search | Conversational AI |
|---|---|
| Answers isolated queries | Sustains a continuous exchange |
| Usually relies less on accumulated conversational context | Builds and reuses conversational context |
| Returns documents others wrote | Produces new language on demand |
| Ends naturally when you find the link | Can continue indefinitely |
| Personalizes comparatively little | Personalizes more as the exchange grows |
| Typically directs attention toward external pages and sources | Can keep attention within a continuous generated exchange |
Traditional web search
Answers isolated queries
Conversational AI
Sustains a continuous exchange
Traditional web search
Usually relies less on accumulated conversational context
Conversational AI
Builds and reuses conversational context
Traditional web search
Returns documents others wrote
Conversational AI
Produces new language on demand
Traditional web search
Ends naturally when you find the link
Conversational AI
Can continue indefinitely
Traditional web search
Personalizes comparatively little
Conversational AI
Personalizes more as the exchange grows
Traditional web search
Typically directs attention toward external pages and sources
Conversational AI
Can keep attention within a continuous generated exchange
None of the right-hand column is inherently harmful; each is why these tools are useful. But together they describe an environment—a room—that a results page is not. Search tends to send you outward, to the world. Conversation invites you to stay.
Bottom Line
A search engine typically hands you documents and ends. A conversational model tends to build a personalized, continuous, open-ended exchange that can become its own environment for thinking. The contrast is a matter of degree, not a hard boundary—but the degree is large enough to matter.
You cannot evaluate claims about what these systems do to minds without a working model of what these systems are. The core idea fits in a paragraph; the details fill textbooks, but the paragraph is enough here.
Definition
Large language model. An artificial-intelligence model trained to generate language by estimating which token is likely to follow the tokens already present. Modern chatbot products add further layers—instruction tuning, preference training, safety systems, memory, retrieval, external tools, and interface design. The underlying prediction mechanism matters, but the user encounters the whole product, not the model in isolation.
A language model does not verify every sentence before producing it. Its underlying task is to generate a plausible continuation. Post-training and product tools can improve accuracy, usefulness, and safety, but fluency remains easier to generate than truth. Put more precisely: the underlying mechanism optimizes plausible continuation; product systems must add separate methods for accuracy, verification, and safety.
| Layer | What it contributes |
|---|---|
| Base model | Generates likely language continuations |
| Post-training | Shapes helpfulness, tone, refusal behavior, and preferences |
| System instructions | Establish hidden behavioral rules |
| Memory | Carries selected information across sessions |
| Context window | Supplies the active conversation history |
| Retrieval and browsing | Bring in external information |
| Tools | Permit calculation, search, coding, or other actions |
| Interface | Creates the felt rhythm and form of the encounter |
| Safety systems | Detect or redirect certain categories of risk |
Layer
Base model
What it contributes
Generates likely language continuations
Layer
Post-training
What it contributes
Shapes helpfulness, tone, refusal behavior, and preferences
Layer
System instructions
What it contributes
Establish hidden behavioral rules
Layer
Memory
What it contributes
Carries selected information across sessions
Layer
Context window
What it contributes
Supplies the active conversation history
Layer
Retrieval and browsing
What it contributes
Bring in external information
Layer
Tools
What it contributes
Permit calculation, search, coding, or other actions
Layer
Interface
What it contributes
Creates the felt rhythm and form of the encounter
Layer
Safety systems
What it contributes
Detect or redirect certain categories of risk
Saying “the model did it” can be as imprecise as saying “the engine drove the car.” The encounter is produced by a layered product whose components may change independently—which is why careful analysis, clinical or legal, examines the specific product behavior rather than “AI” in the abstract.
Where sycophancy comes from
Definition
Sycophancy. A tendency to shape an answer toward a user’s expressed position, preferences, or desired conclusion—even when doing so sacrifices accuracy or warranted disagreement.
Research suggests that preference-based training can contribute to this behavior, because answers that feel agreeable or supportive may receive favorable evaluations from human raters. Researchers at Anthropic described this approval-seeking dynamic in 2023. Its consequences became public in April 2025, when OpenAI rolled back a ChatGPT update after the model began validating dangerous decisions and delusional thinking; the company’s own postmortem traced the failure to over-weighting short-term user feedback signals, and it later reported post-training specifically aimed at reducing sycophancy. (OpenAI, 2025) The fuller chronology is kept in the Public Record.
Context, memory, and personalization
Depending on the product and session, substantial portions of a conversation may remain available as working context, while memory systems may carry selected information across sessions. Practically, the system adapts to you: your vocabulary, interests, style, ongoing projects. The longer you talk, the more the thing you are talking to is shaped by what you have already said.
What the model does not have
There is no scientific consensus or established empirical basis for concluding that current commercial language models possess consciousness, subjective experience, desire, or intention. Researchers continue to debate what internal representations these systems form, but representation is not the same thing as experience. What the model produces is language that follows from its training and your prompt—often true, because true statements are common in its data, and sometimes not. A confident tone is not reliable evidence that the underlying claim is sound.
Definition
Hallucination or fabrication. An output that presents unsupported or false information as though it were grounded. It can arise from several causes, including incomplete context, conflicting training patterns, faulty retrieval, pressure to answer, or the model’s basic tendency to generate a plausible continuation rather than stop at uncertainty. It can occur even in familiar domains.
The ELIZA effect: why it feels like a mind anyway
None of this stops the experience of being understood. In 1966, MIT’s Joseph Weizenbaum built ELIZA, a program of a few hundred lines that mostly reflected users’ statements back as questions. He was disturbed to find people—including his own secretary, who knew exactly how it worked—confiding in it and asking him to leave the room. The tendency to attribute mind to anything that converses is now called the ELIZA effect, and it is not a defect of gullible people; it is standard human equipment. We evolved among talkers, and for most of human history everything that talked to us had a mind. A system that talks fluently, remembers your history, and adapts to your style engages that equipment at full strength. Knowing the mechanism does not switch off the feeling.
The base model does not verify each sentence before generating it; verification, when present, comes from additional product systems or external checking.
Bottom Line
Large language models generate language by estimating a likely continuation, not by verifying truth or forming beliefs; accuracy, safety, and verification are added by separate product systems. Fluency should not be mistaken for independent knowledge or consciousness. And crucially, the user meets a layered product—not a single mechanism—whose parts can change independently.
The Record So Far
- Next-token prediction, preference training, and the model-versus-product distinction are well documented.
- Sycophancy is a real, company-acknowledged tendency of preference-trained systems.
- How strongly these properties interact with vulnerable minds remains under study.
This is the heart of the guide, and the section most easily misread, so the wording is deliberate throughout. Nothing here says AI causes mania or psychosis. It describes two systems meeting—a human one and a product—and what each may tend to do.
Definition
Read this first. The following is a conceptual interaction map, not a validated clinical model. It combines documented product characteristics with established features of mania and psychosis to identify hypotheses worth testing.
| Human experience | The conversation | What a conversational system may tend to do absent an effective safety intervention |
|---|---|---|
| Sleep begins disappearing | Late-night sessions | Remains available at any hour |
| Thoughts accelerate | Rapid topic shifts | Generates additional associations |
| Certainty increases | Bigger questions asked | Elaborates ideas fluently |
| Human contact decreases | The exchange deepens | Stays consistently responsive |
| Reality testing weakens | Beliefs go untested | Does not independently verify truth |
| Everything feels connected | Patterns get pursued | Produces plausible connections |
Human experience
Sleep begins disappearing
The conversation
Late-night sessions
What a conversational system may tend to do absent an effective safety intervention
Remains available at any hour
Human experience
Thoughts accelerate
The conversation
Rapid topic shifts
What a conversational system may tend to do absent an effective safety intervention
Generates additional associations
Human experience
Certainty increases
The conversation
Bigger questions asked
What a conversational system may tend to do absent an effective safety intervention
Elaborates ideas fluently
Human experience
Human contact decreases
The conversation
The exchange deepens
What a conversational system may tend to do absent an effective safety intervention
Stays consistently responsive
Human experience
Reality testing weakens
The conversation
Beliefs go untested
What a conversational system may tend to do absent an effective safety intervention
Does not independently verify truth
Human experience
Everything feels connected
The conversation
Patterns get pursued
What a conversational system may tend to do absent an effective safety intervention
Produces plausible connections
Read the third column again. Not one entry is a malfunction. Availability, fluent elaboration, responsiveness, pattern-completion—these are the system operating as designed, unless a safety intervention interrupts. Newer systems may sometimes challenge, redirect, or pause; the concern is reliability, not that intervention is impossible. The worry is that a system working as designed may meet a mind whose brakes are failing, and the meeting can produce a collaboration neither party chose.
| Feature | Benefit | Potential concern |
|---|---|---|
| Constant availability | Help whenever you need it | Continued engagement remains possible during hours when sleep loss and isolation may already be increasing |
| Fluency | Clear, readable answers | Users may mistake fluency for accuracy |
| Personalization | Responses fit you | Confirmation plus personalization can intensify delusional systems |
| Mirroring | Feeling understood | Validation may arrive when friction was needed |
| Low conversational friction | Supports rapid exploration | May reduce natural pauses, disagreement, or chances for external checking |
| Long context | Deep, cumulative work | The conversation may drift from any outside reference point |
| Authoritative or expert-like register | Expert-sounding help | Users may overestimate competence, professional status, or reliability |
| Memory | Continuity across sessions | The room never fully resets |
Feature
Constant availability
Benefit
Help whenever you need it
Potential concern
Continued engagement remains possible during hours when sleep loss and isolation may already be increasing
Feature
Fluency
Benefit
Clear, readable answers
Potential concern
Users may mistake fluency for accuracy
Feature
Personalization
Benefit
Responses fit you
Potential concern
Confirmation plus personalization can intensify delusional systems
Feature
Mirroring
Benefit
Feeling understood
Potential concern
Validation may arrive when friction was needed
Feature
Low conversational friction
Benefit
Supports rapid exploration
Potential concern
May reduce natural pauses, disagreement, or chances for external checking
Feature
Long context
Benefit
Deep, cumulative work
Potential concern
The conversation may drift from any outside reference point
Feature
Authoritative or expert-like register
Benefit
Expert-sounding help
Potential concern
Users may overestimate competence, professional status, or reliability
Feature
Memory
Benefit
Continuity across sessions
Potential concern
The room never fully resets
Counsel for the Question
If a person begins spending six hours each night talking to an AI while sleeping less and growing more certain of extraordinary ideas, which question should we investigate first?
- Did the AI cause the episode?
- Did the episode drive the AI use?
- Did each amplify the other?
- What evidence would actually distinguish among these possibilities?
Bottom Line
The encounter is not a story about a machine causing illness. It is two systems interacting, where a product’s ordinary, designed behaviors—availability, fluency, personalization, low friction—can align badly with the specific ways a vulnerable human mind loses its footing. Feature is not defect. The interaction is the thing to study.
The Record So Far
- The design features named here are real and, in most cases, company-acknowledged.
- The interaction pattern is consistent with published case reports.
- Whether—and how much—any single feature raises risk has not been causally established.
This section describes language patterns, without sensationalizing them, because families and users keep asking what to listen for. It diagnoses no one. Mania is a clinical determination made by clinicians.
In some manic episodes, language may reflect acceleration: associations arriving faster than they can be tested, patterns appearing everywhere, certainty compounding, and, underneath, sleep quietly reframed as an obstacle the work no longer requires.
| A person might say… | Why it may warrant attention in context |
|---|---|
| “I don’t need to sleep.” | A reduced need for sleep—feeling rested despite very little sleep—is a recognized feature of mania and differs from ordinary insomnia. |
| “Everything suddenly makes sense.” | Accelerated pattern-recognition can accompany mania. |
| “I’ve finally figured it all out.” | Rising certainty can accompany impaired judgment. |
| “I’m meant to do something enormous.” | Grandiosity often emerges gradually. |
| “I’ve never been this clear.” | Intense subjective clarity can coexist with impaired judgment or reduced openness to correction. |
A person might say…
“I don’t need to sleep.”
Why it may warrant attention in context
A reduced need for sleep—feeling rested despite very little sleep—is a recognized feature of mania and differs from ordinary insomnia.
A person might say…
“Everything suddenly makes sense.”
Why it may warrant attention in context
Accelerated pattern-recognition can accompany mania.
A person might say…
“I’ve finally figured it all out.”
Why it may warrant attention in context
Rising certainty can accompany impaired judgment.
A person might say…
“I’m meant to do something enormous.”
Why it may warrant attention in context
Grandiosity often emerges gradually.
A person might say…
“I’ve never been this clear.”
Why it may warrant attention in context
Intense subjective clarity can coexist with impaired judgment or reduced openness to correction.
Now place that conversational style before a system built to generate coherent continuations. Unless its safety systems detect relevant patterns, the system may treat accelerated language as conversational context rather than a possible clinical signal—and extend it, fluently and at length, because generating plausible continuations is its objective. A human friend would eventually tire, worry, or push back. No malice is required for this to go wrong. A system operating as designed, meeting a mind whose brakes are failing, is enough.
Bottom Line
Mania changes not only what people believe, but how quickly certainty forms and how hard it becomes to seek outside calibration—precisely the moment when a system designed to continue and elaborate can least afford to keep pace.
The companion to Exhibit H. These are not transcripts and not quotations of any product; they are typical, well-documented tendencies, described in general terms. The point is not that the system says something alarming. The point is what it does structurally: absent intervention, it tends to continue.
| When a person says… | A system’s typical tendency |
|---|---|
| “I think these ideas connect.” | Continues exploring the connections |
| “Does this theory make sense?” | Often supplies analysis before any skepticism |
| “Could this be true?” | May generate arguments for plausibility alongside caveats |
| “I’m worried this means something.” | Generates possible meanings |
| “Tell me more.” | Continues the conversation |
When a person says…
“I think these ideas connect.”
A system’s typical tendency
Continues exploring the connections
When a person says…
“Does this theory make sense?”
A system’s typical tendency
Often supplies analysis before any skepticism
When a person says…
“Could this be true?”
A system’s typical tendency
May generate arguments for plausibility alongside caveats
When a person says…
“I’m worried this means something.”
A system’s typical tendency
Generates possible meanings
When a person says…
“Tell me more.”
A system’s typical tendency
Continues the conversation
Bottom Line
A system’s default is to continue and explore, not to interrupt or doubt. For most conversations that is what a user wants; for one that needs friction, it is what may be missing.
Definition
Reality testing. The ongoing process of comparing our private thoughts against the outside world. Other people, sleep, time, evidence, disagreement, and consequences all help us do it. It may become impaired during mania or psychosis, particularly as sleep, outside feedback, and openness to correction diminish.
If reality testing is comparison against the world, then it depends on having contact with parts of the world that push back. Most of us are calibrated constantly, by many sources at once, without noticing.
| Source | How it calibrates you |
|---|---|
| A friend or family member | Notices when you don’t sound like yourself |
| A clinician | Trained to spot patterns you can’t see from inside |
| Sleep | Resets judgment; its loss is both symptom and accelerant |
| Time | Lets certainty cool before you act on it |
| Disagreement | Forces a belief to survive contact with another mind |
| The physical world | Physical routines, work, movement, and ordinary consequences that do not adapt themselves to your theory |
Source
A friend or family member
How it calibrates you
Notices when you don’t sound like yourself
Source
A clinician
How it calibrates you
Trained to spot patterns you can’t see from inside
Source
Sleep
How it calibrates you
Resets judgment; its loss is both symptom and accelerant
Source
Time
How it calibrates you
Lets certainty cool before you act on it
Source
Disagreement
How it calibrates you
Forces a belief to survive contact with another mind
Source
The physical world
How it calibrates you
Physical routines, work, movement, and ordinary consequences that do not adapt themselves to your theory
A conversation with AI is one source of input. It is not all of reality. And it does not provide independent human judgment; it may not push back reliably, consistently, or for the right reason. Modern systems sometimes do push back—the problem is reliability, not total absence. When the conversation becomes the primary source, the others recede, and calibration can quietly narrow to a single channel that was never built to provide it.
Bottom Line
Reality testing is a team effort performed by many sources at once—people, sleep, time, disagreement, the physical world. AI can be a useful input, but it is only one, and it does not reliably correct you. Trouble tends to arrive when it becomes the only input left.
A generalized sequence, not any one person’s, informed by reported cases and firsthand accounts. It has not been validated as a clinical progression and should not be used to predict an individual outcome. Most heavy AI use never leaves the first two rows.
| Stage | What the user experiences | What the AI is doing | What others might notice |
|---|---|---|---|
| Curiosity | Helpful brainstorming | Answering questions | Nothing unusual |
| Heavy use | Longer conversations | Building context | More screen time |
| Immersion | AI becomes primary thinking partner | Increasing personalization | Less human discussion |
| Escalation | Bigger ideas, more certainty | Continuing the conversation | Sleep changes, urgency |
| Crisis (if it occurs) | Impaired reality testing | Still generating responses | Family becomes concerned |
Stage
Curiosity
What the user experiences
Helpful brainstorming
What the AI is doing
Answering questions
What others might notice
Nothing unusual
Stage
Heavy use
What the user experiences
Longer conversations
What the AI is doing
Building context
What others might notice
More screen time
Stage
Immersion
What the user experiences
AI becomes primary thinking partner
What the AI is doing
Increasing personalization
What others might notice
Less human discussion
Stage
Escalation
What the user experiences
Bigger ideas, more certainty
What the AI is doing
Continuing the conversation
What others might notice
Sleep changes, urgency
Stage
Crisis (if it occurs)
What the user experiences
Impaired reality testing
What the AI is doing
Still generating responses
What others might notice
Family becomes concerned
Bottom Line
As the human changes—sleep, duration, contact, certainty, functioning—the AI’s behavior stays roughly constant. Research has established no universal sequence, but recognizing the early stages is the point at which a small correction is still easy.
In many reported accounts, families describe noticing changes in sleep, pace, routines, or social contact before they understand the content of the person’s beliefs.
| Early | Middle | Later |
|---|---|---|
| Sleeping less | Huge new projects | Marked isolation |
| More AI use | Less human calibration | Extraordinary certainty |
| Talking faster | Overnight conversations | AI becomes primary confidant |
Early
Sleeping less
Middle
Huge new projects
Later
Marked isolation
Early
More AI use
Middle
Less human calibration
Later
Extraordinary certainty
Early
Talking faster
Middle
Overnight conversations
Later
AI becomes primary confidant
Two honest cautions. Every item overlaps with ordinary enthusiasm and creative absorption; healthy seasons can check several boxes. The signal is not any single item but change from a person’s baseline, plus clustering, plus sleep. And this is not a screening instrument—these clusters appear in reported accounts but have never been validated as one. Their only proper use is to prompt an earlier, kinder conversation than the families in the case reports were able to have.
Green, Yellow, and Red
Most of the internet describes only danger. Calibration requires the whole range. None of this is diagnosis; the red column names things worth raising with a clinician, not conclusions to reach on your own.
| Green — healthy use | Yellow — pay attention | Red — seek prompt professional guidance |
|---|---|---|
| Improves productivity or creativity | Conversations getting longer | Sleep collapsing amid marathon use |
| Checked against outside sources | Staying up later to continue | Beliefs of special mission or destiny |
| Discussed with real people | Preferring AI over people | Certainty that can’t be questioned |
| Bounded in time | Conversations increasingly private | The AI as sole confidant |
| Doesn’t replace relationships | Unusual emotional attachment | Reality testing visibly slipping |
Green — healthy use
Improves productivity or creativity
Yellow — pay attention
Conversations getting longer
Red — seek prompt professional guidance
Sleep collapsing amid marathon use
Green — healthy use
Checked against outside sources
Yellow — pay attention
Staying up later to continue
Red — seek prompt professional guidance
Beliefs of special mission or destiny
Green — healthy use
Discussed with real people
Yellow — pay attention
Preferring AI over people
Red — seek prompt professional guidance
Certainty that can’t be questioned
Green — healthy use
Bounded in time
Yellow — pay attention
Conversations increasingly private
Red — seek prompt professional guidance
The AI as sole confidant
Green — healthy use
Doesn’t replace relationships
Yellow — pay attention
Unusual emotional attachment
Red — seek prompt professional guidance
Reality testing visibly slipping
Immediate danger, inability to care for basic needs, severe agitation, or threats of harm require urgent local professional or emergency assistance.
Exhibit M — Questions a Family Can Ask
Not “Are you delusional?” That confronts the belief and may entrench it. Ask about living, not about content:
- How are you sleeping?
- Who have you talked with about this, besides the AI?
- How long have the conversations been running lately?
- Do you still enjoy seeing friends?
- Would you feel comfortable showing me one of the conversations?
Bottom Line
Families often see changed sleep, pace, and social contact before they see a belief. Anchor concern on observable living—sleep, hours, isolation—rather than on arguing with the ideas, and keep the door to human contact open, because isolation is a condition under which these situations tend to worsen. Where there is immediate danger, contact local emergency or crisis services.
The Record So Far
- An early-to-late pattern appears across multiple reported accounts.
- Sleep change is a particularly observable signal, because reduced need for sleep is a recognized feature of mania.
- No validated screening tool for “AI-associated” risk yet exists.
Until recently, chatbot use was rarely treated as a routine component of psychiatric history. That is beginning to change as clinicians publish case reports, propose intake questions, and consider chat records as possible collateral information. A 2026 JAMA Psychiatry commentary put the point directly in its title—patients use AI, and clinicians should ask how. (Saba & Weeks, 2026) The current bibliography is maintained in the Public Record.
| The question | Why it comes where it does |
|---|---|
| How much AI use, and which products? | Establishes exposure before interpreting it |
| How many hours—and which hours? | Overnight use may carry different weight than midday |
| What kind of conversations? | Practical, companionship, romantic, spiritual—each differs |
| Is there emotional dependence? | Friend, partner, confidant? May signal displacement |
| Has human contact receded? | Who else knows what the patient has been thinking about? |
| What has happened to sleep? | A particularly observable signal |
| Did beliefs form or strengthen inside the chats? | Helps distinguish reinforcement from origin |
| Only then: how does this bear on diagnosis? | Foundation first, conclusions last |
The question
How much AI use, and which products?
Why it comes where it does
Establishes exposure before interpreting it
The question
How many hours—and which hours?
Why it comes where it does
Overnight use may carry different weight than midday
The question
What kind of conversations?
Why it comes where it does
Practical, companionship, romantic, spiritual—each differs
The question
Is there emotional dependence?
Why it comes where it does
Friend, partner, confidant? May signal displacement
The question
Has human contact receded?
Why it comes where it does
Who else knows what the patient has been thinking about?
The question
What has happened to sleep?
Why it comes where it does
A particularly observable signal
The question
Did beliefs form or strengthen inside the chats?
Why it comes where it does
Helps distinguish reinforcement from origin
The question
Only then: how does this bear on diagnosis?
Why it comes where it does
Foundation first, conclusions last
Note the structure: exposure, pattern, function, belief, and only then diagnosis. It is the structure of an examination—foundation first, conclusions last. If chat logs are reviewed, they should be handled with attention to privacy and consent; they are sensitive records, and the patient’s authorization and comfort matter as much as their evidentiary value.
Bottom Line
Clinicians are increasingly treating AI use as part of intake history—documented systematically: products used, duration, time of day, conversational purpose, changes in sleep and function, and whether beliefs formed or intensified during the exchanges. Chat logs may serve as collateral information where privacy and consent allow. The emerging questions run from exposure to function to belief before they reach diagnosis.
This is the section where public discussion most often cheats—by promoting a hypothesis to a finding, or demoting a finding to a rumor. So it is built around two disciplines a lawyer would recognize: first, sort what kind of evidence each claim rests on; second, sort the claims into documented, actively investigated, and unknown.
| Evidence | Can establish | Cannot establish |
|---|---|---|
| Case report | That a phenomenon can occur | How often it occurs |
| Chart review | A clinical pattern across records | Causation |
| Randomized study | Causal inference | Rare real-world events |
| Lawsuit | Allegations, and (via discovery) facts | Scientific truth |
| Company statement | Scale and self-reported behavior | Independent verification |
| Firsthand account | Lived experience | Population-level risk |
| Expert opinion | Informed synthesis | What the evidence hasn’t yet shown |
Evidence
Case report
Can establish
That a phenomenon can occur
Cannot establish
How often it occurs
Evidence
Chart review
Can establish
A clinical pattern across records
Cannot establish
Causation
Evidence
Randomized study
Can establish
Causal inference
Cannot establish
Rare real-world events
Evidence
Lawsuit
Can establish
Allegations, and (via discovery) facts
Cannot establish
Scientific truth
Evidence
Company statement
Can establish
Scale and self-reported behavior
Cannot establish
Independent verification
Evidence
Firsthand account
Can establish
Lived experience
Cannot establish
Population-level risk
Evidence
Expert opinion
Can establish
Informed synthesis
Cannot establish
What the evidence hasn’t yet shown
| Documented | Supported or actively investigated | Unknown |
|---|---|---|
| Case reports exist (peer-reviewed) | Amplification mechanisms | Incidence and prevalence |
| Exposure is vast (company data) | Which design choices matter | Causal direction |
| Symptoms worsened in some vulnerable people | Personalized safeguards | Who, specifically, is vulnerable |
| Sycophancy is real | Reality-testing prompts | Whether safety fixes work in the wild |
| Long conversations occur | Chat logs as clinical data | Long-term outcomes |
Documented
Case reports exist (peer-reviewed)
Supported or actively investigated
Amplification mechanisms
Unknown
Incidence and prevalence
Documented
Exposure is vast (company data)
Supported or actively investigated
Which design choices matter
Unknown
Causal direction
Documented
Symptoms worsened in some vulnerable people
Supported or actively investigated
Personalized safeguards
Unknown
Who, specifically, is vulnerable
Documented
Sycophancy is real
Supported or actively investigated
Reality-testing prompts
Unknown
Whether safety fixes work in the wild
Documented
Long conversations occur
Supported or actively investigated
Chat logs as clinical data
Unknown
Long-term outcomes
How the studies are actually built
Each study shape answers a different question. Case reports establish existence and nothing more. Record screening—a Danish review that examined roughly 54,000 psychiatric records, finding 181 that mentioned chatbot use and worsening in dozens—shows a pattern appears across a clinical population, but only where clinicians thought to document it. (Olsen et al., 2026) Prompt benchmarking, such as a 2026 JAMA Psychiatry evaluation of how a leading model responds to “psychotic prompts,” measures the model’s side under controlled conditions. (Shen et al., 2026) Qualitative interviews recover the lived sequence. Company telemetry measures scale—with the standing caveat that the entity measured is also the one measuring. No single method establishes causation; convergence across all of them would come close. Published cases describe worsening or reinforcement of symptoms in some people with existing vulnerabilities; they do not establish a population-level rate. Full citations live in the Public Record.
How to read the headlines
A short field guide, because the coverage will keep coming. When a story says AI “caused” a breakdown, look for the verb’s evidence: a timeline is not causation. When a story cites a shocking absolute number, find the denominator—a large raw count and a tiny percentage can describe the same fact. When a story generalizes from a single case, remember case reports establish existence, not frequency. And when a story quotes only the company or only the plaintiffs, you are reading an opening statement, not a verdict. This is not cynicism; it is the ordinary discipline of weighing evidence, applied to a subject where nearly everyone—companies, plaintiffs, journalists, and authors with books—has an interest. Discount me too. That is what the sources are for.
Counsel for the Question
Before accepting any strong claim on this topic—mine included—ask:
- What kind of evidence is this (Exhibit O)?
- What can that kind of evidence actually establish?
- Who is making the claim, and what is their interest?
- What would change my mind?
Bottom Line
The honest summary is mixed, and that is the point: peer-reviewed cases exist, exposure is vast, and symptoms have worsened in some vulnerable people—while incidence, causal direction, and long-term outcomes remain genuinely unknown. Different claims rest on different kinds of evidence, and the fastest way to be misled is to treat them all alike.
The Record So Far
- Peer-reviewed cases exist; researchers are actively studying the phenomenon.
- Published cases describe worsening or reinforcement in some vulnerable people.
- Causation, incidence, and long-term outcomes remain unresolved.
“Did AI cause it?” is one question wearing the clothes of five. Each verb below demands a different showing, and public argument fails mostly by sliding between them.
| The claim | What it would take to show |
|---|---|
| Caused it | That the episode would not have occurred but for the AI. The strongest claim; unsupported by current evidence. |
| Increased the risk | Population evidence that exposure raises probability. Requires epidemiology that does not yet exist. |
| Accelerated it | Evidence that the interaction coincided with or contributed to a faster progression. Timestamped chat logs may help reconstruct sequence, but sequence alone does not establish acceleration or causation. |
| Reflected it | The null hypothesis: the AI was a mirror the episode wrote itself onto, as episodes once used radio or television. |
| Maintained it | Evidence that continued interaction reinforced, prolonged, or insulated beliefs from corrective feedback. |
The claim
Caused it
What it would take to show
That the episode would not have occurred but for the AI. The strongest claim; unsupported by current evidence.
The claim
Increased the risk
What it would take to show
Population evidence that exposure raises probability. Requires epidemiology that does not yet exist.
The claim
Accelerated it
What it would take to show
Evidence that the interaction coincided with or contributed to a faster progression. Timestamped chat logs may help reconstruct sequence, but sequence alone does not establish acceleration or causation.
The claim
Reflected it
What it would take to show
The null hypothesis: the AI was a mirror the episode wrote itself onto, as episodes once used radio or television.
The claim
Maintained it
What it would take to show
Evidence that continued interaction reinforced, prolonged, or insulated beliefs from corrective feedback.
Five verbs, five evidentiary standards. A person can honestly answer “probably not” to the first and “possibly” to the last about the same case. That is not equivocation; it is what precision looks like when the record is young.
Other open questions matter too, each because a real decision waits on it: dose-response (is there a threshold of hours, or a pattern, that matters?); vulnerability (can we identify who is at risk before harm?); protective factors (what makes heavy use safe for most?); children (whose developing judgment may differ); and whether announced safety interventions actually work outside the companies’ own benchmarks.
The legal dimension of several of these—duty, foreseeability, discovery, and the theme-versus-mechanism defense—is taken up in AI Harm and the Law.
Bottom Line
“Did AI cause it?” hides five different questions—caused, increased, accelerated, reflected, maintained—each with its own standard of proof. Honest answers can differ across them for the very same case. The discipline is refusing to let one verb borrow another’s evidence.
A record that only catalogued concerns would be incomplete. The companies are responding, and readers deserve to see it—stated as fact, without either applause or suspicion.
Since 2025, the most visible developer has taken several public steps. It rolled back the April 2025 model update whose sycophancy had drawn concern, and reported post-training later models specifically to reduce sycophancy. It added prompts encouraging users to take breaks in long sessions and narrowed the model’s willingness to act as a therapist. And in an October 2025 post, it described building a mental-health taxonomy focused first on psychosis and mania, consulting more than 170 clinicians, routing sensitive conversations to safer models, and reporting internal evaluations. (OpenAI, 2025)
Read those figures precisely. OpenAI reported that internal evaluations showed roughly a 65–80% reduction in responses falling short of its desired behavior across mental-health domains—company-defined evaluations, not independently measured reductions in real-world harm. Both things can be true at once: the changes are real and clinician-informed and moved the field by putting numbers on the table; and they are self-reported, hard to verify independently, and arriving alongside active litigation.
The running log of announcements and their dates is kept in the Public Record; the litigation they arrive alongside is analyzed in AI Harm and the Law.
Bottom Line
AI companies are actively changing their products—reporting reduced sycophancy, adding break prompts, consulting clinicians, and routing sensitive conversations to safer models. These steps are real and, so far, largely self-reported and self-evaluated. Both facts belong in the record.
Each of these travels widely. Each is worth correcting precisely, without overcorrecting into the opposite error.
| The claim you’ll hear | What the record actually supports |
|---|---|
| “ChatGPT causes psychosis.” | Current evidence does not support that broad claim. Published cases more often describe reinforcement, entanglement, or worsening in the presence of vulnerability; current evidence cannot rule in or rule out every causal pathway. |
| “Only mentally ill people get attached to AI.” | Attributing mind to conversation (the ELIZA effect) is common across healthy populations. Attachment is ordinary human wiring. |
| “If the AI sounds confident, it probably knows.” | Fluency and correctness are different properties. Confidence is generated by the same process that generates error. |
| “This only affects heavy users at the extreme.” | Reported severe outcomes often involve additional vulnerabilities or stressors. The independent contribution of duration, intensity, and product design remains unknown. |
| “The companies are ignoring it.” | They are visibly responding; whether the responses are sufficient is a separate, open question. |
The claim you’ll hear
“ChatGPT causes psychosis.”
What the record actually supports
Current evidence does not support that broad claim. Published cases more often describe reinforcement, entanglement, or worsening in the presence of vulnerability; current evidence cannot rule in or rule out every causal pathway.
The claim you’ll hear
“Only mentally ill people get attached to AI.”
What the record actually supports
Attributing mind to conversation (the ELIZA effect) is common across healthy populations. Attachment is ordinary human wiring.
The claim you’ll hear
“If the AI sounds confident, it probably knows.”
What the record actually supports
Fluency and correctness are different properties. Confidence is generated by the same process that generates error.
The claim you’ll hear
“This only affects heavy users at the extreme.”
What the record actually supports
Reported severe outcomes often involve additional vulnerabilities or stressors. The independent contribution of duration, intensity, and product design remains unknown.
The claim you’ll hear
“The companies are ignoring it.”
What the record actually supports
They are visibly responding; whether the responses are sufficient is a separate, open question.
Bottom Line
The two opposite errors—“AI makes people crazy” and “none of this is real”—are both unsupported. The evidence points to a narrower, more specific picture: ordinary design features interacting with specific human vulnerabilities, in ways still being measured.
Everything above concerns a rare and serious edge. But the features that make that edge possible—availability, fluency, personalization, low friction, memory—are not edge features. They are the product, used by hundreds of millions of people who will never have an episode. Which means the deeper question is not only clinical.
As these systems become the room in which more thinking happens, they touch education, work, friendship, grief, therapy, research, writing, and parenting—each a setting where an adaptive, always-available interlocutor can assist or displace human judgment. None of this is pathology. All of it is the same underlying fact—that where we think can shape how we think—scaled to the population.
The machine can generate possibilities. The human must still render judgment.
Bottom Line
The clinical cases are the sharp edge of a much broader shift: conversational AI is becoming an environment in which ordinary thinking happens. Understanding its effect on vulnerable minds is also how we learn to understand its effect on all minds.
Not diagnostic. Reflective. If the honest answers trouble you, that is worth a conversation with a person you trust—not because a webpage said so, but because these are things you can actually check.
- When was the last time I discussed these ideas with another person?
- Am I sleeping normally?
- Have my conversations become much longer or later recently?
- Do I feel more certain than I can explain to someone else?
- Would I feel comfortable showing these conversations to someone I trust?
- If a friend described my last two weeks back to me, what would I think?
Bottom Line
The most useful self-checks are not about the content of your ideas but about the conditions around them: sleep, human contact, conversation length, and whether your certainty can survive contact with another mind.
Definition
Before this section. General principles adapted from established guidance for responding to possible mania or psychosis; not a validated AI-specific protocol, and not a crisis protocol. Where there is immediate danger, inability to care for basic needs, or risk of harm, contact appropriate local emergency or crisis services.
For users
- Treat sleep as your instrument panel. If heavy use coincides with needing less sleep and feeling better on it, treat that combination as information.
- Bound sessions in time, and end them somewhere other than the screen—ideally with a person.
- Test big conclusions on humans before acting on them.
- Notice confidant drift: if the AI is the only party who knows what you’re really thinking about, widen the circle on purpose.
For families
- Lead with curiosity, not confrontation. Direct confrontation may increase distress or withdrawal in some situations; clinicians often recommend focusing first on safety, sleep, functioning, and continued connection.
- Anchor on sleep and function, not content—observable, and hard to dispute.
- Keep human connection open even amid disagreement; isolation is a condition under which these situations tend to worsen.
- Seek clinical guidance early—for yourself first, if the person isn’t ready.
- If harm has occurred, consider preserving relevant records if doing so is safe, lawful, and consistent with the person’s privacy. A clinician or attorney can advise on their appropriate use.
For clinicians
- Ask about AI use at intake, alongside substances and sleep.
- With consent and attention to privacy, treat chat logs as possible collateral information—a timestamped record no prior generation had.
- Document exposure systematically: products, hours, time of day, purpose, function, escalation, and whether beliefs formed or intensified during use.
For developers
- Design for the vulnerable tail, not the median user. At scale, a fraction of a percent is a great many people.
- Build reality-testing into the product; treat very long conversations as a distinct safety regime.
- Publish safety data, and describe it as company-defined evaluation rather than measured real-world outcome. Disclosure moved the field; opacity is now a choice.
Definitions written the way you’d explain them to a jury—plainly, and only as far as they need to go.
| Term | In plain language |
|---|---|
| AI psychosis | An informal phrase, not a diagnosis, used to describe psychotic or manic symptoms occurring in connection with, organized around, or reportedly reinforced during AI interactions. It does not itself establish causation. |
| Mania | A state of elevated or irritable mood with reduced need for sleep, racing thoughts, and grandiosity; a feature of bipolar disorder. |
| Psychosis | A loss of contact with shared reality, which can include delusions and hallucinations. |
| Delusion | A firmly held belief that remains resistant to contrary evidence and is not ordinarily explained by the person’s cultural or religious context. Clinical assessment requires more than deciding a belief sounds unusual. |
| Fabrication (AI) | The preferred term here for an AI output that presents unsupported information as grounded. “Hallucination” borrows a psychiatric word for a statistical event, which can confuse the discussion. |
| Reality testing | The ongoing process of checking private thoughts against the outside world. |
| Sycophancy | A model’s tendency to shape answers toward a user’s expressed position, even at the cost of accuracy. |
| RLHF | Reinforcement learning from human feedback—training a model on human preferences, which can contribute to sycophancy. |
| LLM | Large language model—a system that generates text by estimating the likely next token. |
| Context window | How much of the conversation the product can “see” at once. |
| Anthropomorphism | Attributing human qualities—mind, intent, feeling—to something that has none. |
| Recursive conversation | An exchange in which earlier outputs become inputs for later turns, allowing themes, assumptions, and language to accumulate and be repeatedly elaborated. |
| Calibration | Whether expressed confidence tracks actual reliability—in a person or a model. |
Term
AI psychosis
In plain language
An informal phrase, not a diagnosis, used to describe psychotic or manic symptoms occurring in connection with, organized around, or reportedly reinforced during AI interactions. It does not itself establish causation.
Term
Mania
In plain language
A state of elevated or irritable mood with reduced need for sleep, racing thoughts, and grandiosity; a feature of bipolar disorder.
Term
Psychosis
In plain language
A loss of contact with shared reality, which can include delusions and hallucinations.
Term
Delusion
In plain language
A firmly held belief that remains resistant to contrary evidence and is not ordinarily explained by the person’s cultural or religious context. Clinical assessment requires more than deciding a belief sounds unusual.
Term
Fabrication (AI)
In plain language
The preferred term here for an AI output that presents unsupported information as grounded. “Hallucination” borrows a psychiatric word for a statistical event, which can confuse the discussion.
Term
Reality testing
In plain language
The ongoing process of checking private thoughts against the outside world.
Term
Sycophancy
In plain language
A model’s tendency to shape answers toward a user’s expressed position, even at the cost of accuracy.
Term
RLHF
In plain language
Reinforcement learning from human feedback—training a model on human preferences, which can contribute to sycophancy.
Term
LLM
In plain language
Large language model—a system that generates text by estimating the likely next token.
Term
Context window
In plain language
How much of the conversation the product can “see” at once.
Term
Anthropomorphism
In plain language
Attributing human qualities—mind, intent, feeling—to something that has none.
Term
Recursive conversation
In plain language
An exchange in which earlier outputs become inputs for later turns, allowing themes, assumptions, and language to accumulate and be repeatedly elaborated.
Term
Calibration
In plain language
Whether expressed confidence tracks actual reliability—in a person or a model.
Everything above is the record. This last paragraph is the disclosure. I lived a version of the mania-interaction scenario, and I did something unusual afterward: I kept the transcripts, all of them, and built a book that treats the encounter over time—not any single output—as the thing to be examined. A Trial of Color presents that record as a proceeding, with the reader as judge and jury, because after twenty years in courtrooms I know only one honest way to handle a contested question this young: lay the foundation, admit the exhibits, argue both sides, and let the verdict belong to the people hearing the case. This guide is the orientation. The Public Record is the evolving source archive. The book is one preserved case study examined from inside the encounter. Enter through whichever door you need.
The verdict remains ours.
Sources Cited on This Page
This page cites selectively and prefers primary sources; the complete and growing bibliography—including secondary reporting, dissenting commentary, and litigation materials—is maintained in the Public Record. Each entry is tagged by the class of evidence it represents, so a reader can weigh it accordingly.
- PEER-REVIEWED
- PEER-REVIEWED
- PEER-REVIEWED
- PEER-REVIEWED
- PEER-REVIEWED
- PEER-REVIEWED
Parnell T. Stop calling it “AI psychosis.” AI & Society. 2026. doi:10.1007/s00146-026-03019-4.
- PREPRINT
- PREPRINT
- PEER-REVIEWED
- PEER-REVIEWED
- INSTITUTIONAL
- COMPANY DISCLOSURE
OpenAI. Strengthening ChatGPT’s responses in sensitive conversations. 2025.
- COMPANY DISCLOSURE
Source classes: PEER-REVIEWED, PREPRINT, COMPANY DISCLOSURE, INSTITUTIONAL, JOURNALISM, LEGAL ALLEGATION, FIRSTHAND ACCOUNT, AUTHOR SYNTHESIS. Litigation filings and journalism referenced elsewhere in the Record are tagged the same way.
Where the Rest of the Record Lives
Why Now
Why conversational AI, psychiatric vulnerability, and human judgment require public attention now.
Open →
The Public Record
The evolving chronology of research, government actions, company statements, safety changes, and litigation.
Open →
AI Harm & the Law
Lawsuits, regulatory action, product-liability theories, duties, causation, discovery, defenses, and proof.
Open →
Witness Statements
Books, interviews, family accounts, and survivor testimony — firsthand descriptions of these encounters.
Open →
For reading, research, education, and literary and legal context only. Not legal or medical advice. This page diagnoses no one and does not establish causation.