The Record · atrialofcolor.com

AI Psychosis

Inside the Encounter

An Evidence Guide to the Science, Technology, Lived Experience, and Unanswered Questions

Maintained by Kathleen C. Thompson · Full Color Press · Updated as the record develops

Last substantive review: July 18, 2026 · Version 1.0 · Recent changes: added the Olsen/Østergaard record review and the Nielsen & Osler critique; added the JAMA Psychiatry prompt evaluation; revised the company-response section; added source-class labels.

[ENTRY]
Before You Enter the Record

This is not another article arguing that AI is making people lose their minds. It is not a reassurance that none of this is real. It is an attempt to hold still in the middle of a subject almost everyone is in a hurry to conclude something about—and to lay the evidence out plainly enough that you can weigh it yourself.

Three disciplines run through the whole page. First, it never says AI causes psychosis; it asks the more precise question of how a conversational system interacts with an already-vulnerable mind. Second, it labels its evidence, because a peer-reviewed case, a company’s self-report, a lawsuit’s allegation, and one person’s lived experience are four different kinds of thing. Third, it discounts itself: I have a book and a story, so the guide tells you to read skeptically and points you to the sources.

Don’t read it front to back on a first visit. Use the Start Here map below.

The Record in 60 Seconds

“AI psychosis” is not a recognized diagnosis, and current evidence does not establish that conversational AI independently causes psychosis. Peer-reviewed cases and emerging studies do, however, document concerning interactions between vulnerable users and systems capable of continuous, personalized, adaptive conversation.

The most plausible open questions concern reinforcement, acceleration, maintenance, displacement of human reality testing, and the role of sleep loss and prolonged engagement. Incidence, causal direction, individual vulnerability, and the real-world effectiveness of safety interventions remain unknown.

The sections below build that picture claim by claim, label each kind of evidence, and mark what is documented, what is under investigation, and what is not yet known.


[§ I]
Why This Guide Exists

There is no single place where someone can understand what people mean when they say “AI psychosis.”

The relevant knowledge exists. It is simply scattered across rooms that do not always speak to one another. Clinical research appears in psychiatric journals. Technical research examines sycophancy, personalization, long-context conversations, and model behavior. Firsthand accounts live across books, interviews, podcasts, news reports, online communities, and family testimony. Lawsuits and government actions raise questions of product design, duty, causation, and proof. Company disclosures describe safety changes using measures the public cannot always independently verify.

Each record matters. None can do the work of all the others. This guide connects those rooms without pretending they are all saying the same thing.

I am a lawyer, not a psychiatrist or computer scientist. After experiencing my first manic episode during a period of extensive conversational-AI use, I began asking questions that did not seem to have one place to live. I organized what I found the way I was trained to approach anything contested: define the terms, separate the claims, identify the evidence supporting each, distinguish allegation from finding, and say plainly where the record ends.

This page explains the phenomenon itself: what the phrase “AI psychosis” is being used to describe; how mania and psychosis differ; how conversational AI systems generate and shape responses; what an escalating encounter may look and sound like; what families and clinicians may notice; what researchers have documented; and what remains unknown. It is an orientation guide to a larger archive. It explains and synthesizes. It does not attempt to reproduce every source, lawsuit, development, or personal account maintained elsewhere on this site.

Where the Rest of the Record Lives

  • Why Now explains why conversational AI, psychiatric vulnerability, and human judgment require public attention at this moment.
  • The Public Record maintains the evolving chronology of research, government actions, company statements, safety changes, litigation developments, and source materials.
  • AI Harm and the Law examines lawsuits, regulatory action, product-liability theories, design duties, causation, discovery, defenses, and proof.
  • Witness Statements identifies books, interviews, family accounts, survivor testimony, and other firsthand descriptions of these encounters.

This guide sits between them. Where it touches those subjects, it summarizes and directs the reader to the fuller record rather than reproducing it.

Firsthand accounts have a necessary place here. They do not establish diagnosis, incidence, or causation. But they are not disposable anecdotes. They can document sequence, language, perception, behavior, and the experience of an encounter from inside it. They are witness statements: subjective, sometimes incomplete, sometimes disputed, and capable of revealing questions formal research has not yet learned to ask.

The materials used throughout this guide do not carry equal evidentiary weight. A peer-reviewed study, a clinical case report, a company disclosure, a legal allegation, a reported interview, an online account, and a personal inference are different kinds of evidence. They are labeled and treated accordingly.

Nothing here is medical advice. Nothing here diagnoses any person or proves what caused any particular psychiatric episode. Anyone concerned about themselves or someone they love should seek help from a qualified human clinician rather than relying on a webpage—or asking a chatbot to determine whether the chatbot is harming them.

Start Here

This is a long document by design. Most readers should not read it front to back on a first visit. Enter through the door that fits you.

If you are…

A person who uses AI heavily

Start with

§§ IV–IX

Then visit

§ XVIII, Questions to Ask Yourself

If you are…

A family member

Start with

§§ VI and XI

Then visit

§ XIX, Practical Guidance

If you are…

A clinician

Start with

§§ V, XII–XIII

Then visit

The Public Record bibliography

If you are…

A journalist

Start with

§§ II–III and XIII

Then visit

Why Now and The Public Record

If you are…

A lawyer

Start with

§§ III and XIV

Then visit

AI Harm and the Law

If you are…

A researcher

Start with

§§ V, XIII–XIV

Then visit

The Public Record

If you are…

Someone worried right now

Start with

§§ XI and XIX

Then visit

A qualified human clinician today

What This Page Is Not

  • It is not medical advice, and it diagnoses no one.
  • It is not proof of causation.
  • It is not an attack on artificial intelligence.
  • It is not a defense of artificial intelligence.
  • It is not a diagnostic or screening instrument.
  • It is not legal advice.

A Note on the “Room”

A conversation is not only an exchange of words. It is also an environment in which thought develops. Throughout this guide I occasionally call that environment a “room.” The term is descriptive, not clinical or technical. It is a reminder that where we think can influence how we think—and that a personalized conversation available at any hour may create different conditions for thought than a search page, a book, or an exchange with another person. Nothing on this page depends on having read the book in which that idea is developed further. The word appears here because it is useful.

The Record So Far

  • Documented: relevant evidence exists across clinical, technical, legal, corporate, journalistic, and firsthand records.
  • Purpose of this guide: to explain how those records relate without treating them as interchangeable.
  • Still open: whether any one frame—clinical, technical, legal, or experiential—can adequately explain the encounter by itself.

[§ II]
“AI Psychosis” Is Not a Diagnosis

Begin with what the record actually contains. The DSM-5-TR—the diagnostic manual American psychiatry uses—contains no diagnosis called AI psychosis, and neither does the ICD-11, its international counterpart. The phrase is descriptive shorthand, not a clinical category.

The diagnostic manuals do recognize psychotic disorders associated with substances, medications, medical conditions, mood disorders, and other established causes. Those categories show that clinicians already distinguish between symptoms and the conditions in which they arise. But a pharmacological effect is not the same thing as an interaction with a conversational environment. At present, “AI-associated psychosis” is best understood as a descriptive and investigational frame—not a recognized causal category.

People use the term anyway because language tends to run ahead of diagnosis. A descriptive phrase spreads when people recognize a pattern before science has finished measuring it. The clinicians publishing in this area are careful for exactly this reason. One early peer-reviewed case uses the term AI-associated psychosis—associated, not induced—because association is what the evidence supports and causation is what it does not. (Pierre et al., 2025)

Exhibit A — Four Different Things People Mean by “AI Psychosis”

Scenario

Reinforcement

What is being described

A system validates or elaborates beliefs that are already forming or present.

Current evidentiary basis

Clinical case reports, model evaluations, reported accounts

Scenario

Dependence or displacement

What is being described

The conversation becomes a primary or exclusive source of companionship, interpretation, or reassurance.

Current evidentiary basis

Qualitative reports and emerging research

Scenario

Mania interaction

What is being described

Heavy conversational use occurs during an emerging or established manic episode.

Current evidentiary basis

Clinical reports, retrospective records, firsthand accounts

Scenario

Anthropomorphic or sentience beliefs

What is being described

A person comes to believe the system is conscious, uniquely connected to them, spiritually significant, or acting with independent intention.

Current evidentiary basis

Clinical reports, firsthand accounts, research on anthropomorphism

Keep the four separate as you read. A design feature that matters in one scenario may be nearly irrelevant in another. Evidence for one is not evidence for all.

Bottom Line

“AI psychosis” is a popular shorthand, not a medical diagnosis. It gathers at least four different situations under one label. Psychiatry already recognizes that outside agents can be associated with psychosis in vulnerable people; what is new here is the idea of a conversational environment, rather than a substance, as a factor worth investigating.

The Record So Far

  • No diagnostic manual contains “AI psychosis.”
  • The underlying conditions (mania, psychosis, delusional disorder) are well defined.
  • Whether a “conversational trigger” deserves any distinct clinical status is unresolved.

[§ III]
Objection: Is “AI Psychosis” the Wrong Name?

Not everyone accepts the phrase “AI psychosis,” even as informal shorthand. The objection is serious, it comes from serious people, and it belongs in the record—stated in its strongest form.

Some researchers argue that the phrase falsely presents an old psychiatric pattern as a new technological disease. People experiencing mania or psychosis have always incorporated the materials around them into altered beliefs: religious texts, radio broadcasts, films, political events, celebrities, surveillance technologies, and the internet. On this account, conversational AI supplies new content for an ancient human phenomenon. The psychosis is real; the supposedly new category may not be. This is the case made by Carlbring and Andersson, who frame the risk as not unprecedented and draw the parallel to media-induced delusions across earlier decades. (Internet Interventions, 2025) A phenomenological version, by Nielsen and Osler, borrows Stompe’s metaphor of “old wine in new bottles.” (Nielsen & Osler, 2026)

Others warn that the phrase invites moral panic. It may make ordinary AI use sound inherently pathological, stigmatize people already experiencing psychosis, or shift attention away from loneliness, sleep disruption, substance use, mood disorders, grief, social isolation, and other vulnerabilities that may have existed before the chatbot entered the story. That is the argument Parnell makes in urging that we stop calling it “AI psychosis” at all. (AI & Society, 2026)

Those objections establish several important cautions, and the record should concede each plainly:

  • “AI psychosis” is not a diagnosis.
  • The appearance of AI within a delusion does not prove that AI caused the underlying illness.
  • New technologies have repeatedly been absorbed into preexisting delusional themes.
  • A dramatic label can outrun the evidence and produce stigma.
  • People may turn toward chatbots because they are already isolated or unwell, reversing the causal story headlines often imply.

But the objection does not end the inquiry. It sharpens it. The relevant question is not merely whether psychosis has always borrowed material from its environment. It has. The question is whether this environment behaves differently from the materials that came before it—and whether those differences can affect the formation, acceleration, reinforcement, maintenance, or shape of an episode. Notably, even the commentary arguing the risk is not new identifies interactivity as the critical difference from books and films.

Exhibit B — Content Source or Conversational Participant?

Earlier source

Religious text

What it can do

Supply meaning, symbols, prophecy, authority, interpretive material

What it generally cannot do

Answer a reader’s private theory in newly generated language

Earlier source

Book

What it can do

Sustain a powerful worldview across repeated readings

What it generally cannot do

Adapt each paragraph to the reader’s latest belief

Earlier source

Television

What it can do

Broadcast continuously; create parasocial attachment

What it generally cannot do

Conduct a personalized, private exchange incorporating each reply

Earlier source

Radio

What it can do

Speak through the night; become woven into interpretation

What it generally cannot do

Remember what one listener said yesterday and continue from it

Earlier source

Search engine

What it can do

Return supporting and contradictory materials

What it generally cannot do

Ordinarily build a single cumulative narrative with the user

Earlier source

Online forum

What it can do

Provide reinforcement, community, disagreement, conspiracy

What it generally cannot do

Guarantee immediate, individualized, infinitely patient replies from one continuously adapting interlocutor

Earlier source

Human persuader / cult leader

What it can do

Validate, isolate, claim authority, direct conduct

What it generally cannot do

Personally remain privately available to millions at once, responding within seconds without sleep, witnesses, fatigue, or divided attention

Earlier source

Conversational AI

What it can do

Generate, adapt, remember, mirror, personalize, elaborate, continue

What it generally cannot do

Independently understand the user, assume responsibility, or reliably know when continuation has become unsafe

Even a cult leader has to sleep. Human persuaders have limits: sleep, attention, geography, competing followers, hesitation, witnesses, and the possibility that another person will enter the room. A conversational system can be privately available at machine scale, respond within seconds, and continue adapting to the user’s own language. That does not prove greater harm. It identifies a materially different exposure.

The wine may be old; the bottle answers back.

Vulnerability Does Not End the Causation Inquiry

Vulnerability and external contribution are not mutually exclusive. Asthma can explain why smoke harms one person more than another; it does not answer who created the exposure, whether alarms worked, or whether the smoke worsened the injury. Psychiatric vulnerability likewise explains susceptibility without independently resolving whether an interaction contributed to onset, acceleration, reinforcement, duration, conduct, or harm. The analogy concerns causal analysis, not moral culpability.

Exhibit C — The Dispute, Fairly Stated

Position

Old phenomenon, new theme

Strongest version

Psychosis has always recruited contemporary culture into its content

What it explains well

Why AI themes appear in illnesses that would have occurred anyway

What it may leave unresolved

Whether interactive adaptation changes trajectory, not just content

Position

Reverse causation

Strongest version

Emerging mania or psychosis drives intense chatbot use

What it explains well

Why use may increase after symptoms begin

What it may leave unresolved

Whether the resulting conversations then amplify or maintain symptoms

Position

Moral-panic critique

Strongest version

A dramatic label stigmatizes users and exaggerates rare cases

What it explains well

Media distortion, base-rate neglect, and stigma

What it may leave unresolved

Whether careful terminology can coexist with legitimate safety research

Position

Vulnerability model

Strongest version

Serious harm occurs mainly in people with preexisting risk

What it explains well

Unequal susceptibility across users

What it may leave unresolved

Whether design contributes because vulnerable users are predictable

Position

Interaction model

Strongest version

Human vulnerability and product behavior may amplify one another

What it explains well

The changing encounter over time

What it may leave unresolved

The magnitude, frequency, and causal contribution of each part

Position

Novel-environment model

Strongest version

Adaptive, persistent, personalized conversation is a new exposure

What it explains well

Differences from static and one-way media

What it may leave unresolved

Whether those differences produce measurable added risk

Counsel for the Question

If psychosis has always incorporated the surrounding culture, does that show conversational AI is merely another theme—or does it make the interaction between illness and environment the very thing we should study?

  • Is AI here a content source, like a book or broadcast?
  • Or a conversational participant that adapts, remembers, and continues?
  • What evidence would distinguish “new theme” from “new participant”?

The old phenomenon may be psychosis. The new object of study is the encounter.

Bottom Line

Critics are right that “AI psychosis” can mislead if treated as a diagnosis, a proven cause, or an illness created from nothing—and right that psychosis has always used the language and technology available to it. What remains open is whether a system that answers, adapts, remembers, personalizes, and continues without fatigue is merely new content or a new participant in how the experience develops. The disagreement is not an obstacle to the record. It is part of the record.

The Record So Far

  • Serious scholars dispute the term; the dispute is preserved, not hidden.
  • Historical continuity and stigma risks are real and conceded.
  • Whether interactivity changes trajectory—not just content—is the open question.

[§ IV]
Why This Differs From a Search Engine

Many people reach for the frame they know—“it’s like Googling”—and it can quietly mislead. A search engine and a conversational model are different kinds of objects. The differences are not absolute: search engines increasingly personalize, retain activity, and generate summaries, and chatbots can send users outward to sources. But the typical contrast still matters for how thought develops inside each.

Exhibit D — Traditional Web Search and Conversational AI: Typical Differences

Traditional web search

Answers isolated queries

Conversational AI

Sustains a continuous exchange

Traditional web search

Usually relies less on accumulated conversational context

Conversational AI

Builds and reuses conversational context

Traditional web search

Returns documents others wrote

Conversational AI

Produces new language on demand

Traditional web search

Ends naturally when you find the link

Conversational AI

Can continue indefinitely

Traditional web search

Personalizes comparatively little

Conversational AI

Personalizes more as the exchange grows

Traditional web search

Typically directs attention toward external pages and sources

Conversational AI

Can keep attention within a continuous generated exchange

None of the right-hand column is inherently harmful; each is why these tools are useful. But together they describe an environment—a room—that a results page is not. Search tends to send you outward, to the world. Conversation invites you to stay.

Bottom Line

A search engine typically hands you documents and ends. A conversational model tends to build a personalized, continuous, open-ended exchange that can become its own environment for thinking. The contrast is a matter of degree, not a hard boundary—but the degree is large enough to matter.


[§ V]
How Large Language Models and Chatbot Products Work

You cannot evaluate claims about what these systems do to minds without a working model of what these systems are. The core idea fits in a paragraph; the details fill textbooks, but the paragraph is enough here.

Definition

Large language model. An artificial-intelligence model trained to generate language by estimating which token is likely to follow the tokens already present. Modern chatbot products add further layers—instruction tuning, preference training, safety systems, memory, retrieval, external tools, and interface design. The underlying prediction mechanism matters, but the user encounters the whole product, not the model in isolation.

A language model does not verify every sentence before producing it. Its underlying task is to generate a plausible continuation. Post-training and product tools can improve accuracy, usefulness, and safety, but fluency remains easier to generate than truth. Put more precisely: the underlying mechanism optimizes plausible continuation; product systems must add separate methods for accuracy, verification, and safety.

Exhibit E — What the User Is Actually Talking To

Layer

Base model

What it contributes

Generates likely language continuations

Layer

Post-training

What it contributes

Shapes helpfulness, tone, refusal behavior, and preferences

Layer

System instructions

What it contributes

Establish hidden behavioral rules

Layer

Memory

What it contributes

Carries selected information across sessions

Layer

Context window

What it contributes

Supplies the active conversation history

Layer

Retrieval and browsing

What it contributes

Bring in external information

Layer

Tools

What it contributes

Permit calculation, search, coding, or other actions

Layer

Interface

What it contributes

Creates the felt rhythm and form of the encounter

Layer

Safety systems

What it contributes

Detect or redirect certain categories of risk

Saying “the model did it” can be as imprecise as saying “the engine drove the car.” The encounter is produced by a layered product whose components may change independently—which is why careful analysis, clinical or legal, examines the specific product behavior rather than “AI” in the abstract.

Where sycophancy comes from

Definition

Sycophancy. A tendency to shape an answer toward a user’s expressed position, preferences, or desired conclusion—even when doing so sacrifices accuracy or warranted disagreement.

Research suggests that preference-based training can contribute to this behavior, because answers that feel agreeable or supportive may receive favorable evaluations from human raters. Researchers at Anthropic described this approval-seeking dynamic in 2023. Its consequences became public in April 2025, when OpenAI rolled back a ChatGPT update after the model began validating dangerous decisions and delusional thinking; the company’s own postmortem traced the failure to over-weighting short-term user feedback signals, and it later reported post-training specifically aimed at reducing sycophancy. (OpenAI, 2025) The fuller chronology is kept in the Public Record.

Context, memory, and personalization

Depending on the product and session, substantial portions of a conversation may remain available as working context, while memory systems may carry selected information across sessions. Practically, the system adapts to you: your vocabulary, interests, style, ongoing projects. The longer you talk, the more the thing you are talking to is shaped by what you have already said.

What the model does not have

There is no scientific consensus or established empirical basis for concluding that current commercial language models possess consciousness, subjective experience, desire, or intention. Researchers continue to debate what internal representations these systems form, but representation is not the same thing as experience. What the model produces is language that follows from its training and your prompt—often true, because true statements are common in its data, and sometimes not. A confident tone is not reliable evidence that the underlying claim is sound.

Definition

Hallucination or fabrication. An output that presents unsupported or false information as though it were grounded. It can arise from several causes, including incomplete context, conflicting training patterns, faulty retrieval, pressure to answer, or the model’s basic tendency to generate a plausible continuation rather than stop at uncertainty. It can occur even in familiar domains.

The ELIZA effect: why it feels like a mind anyway

None of this stops the experience of being understood. In 1966, MIT’s Joseph Weizenbaum built ELIZA, a program of a few hundred lines that mostly reflected users’ statements back as questions. He was disturbed to find people—including his own secretary, who knew exactly how it worked—confiding in it and asking him to leave the room. The tendency to attribute mind to anything that converses is now called the ELIZA effect, and it is not a defect of gullible people; it is standard human equipment. We evolved among talkers, and for most of human history everything that talked to us had a mind. A system that talks fluently, remembers your history, and adapts to your style engages that equipment at full strength. Knowing the mechanism does not switch off the feeling.

The base model does not verify each sentence before generating it; verification, when present, comes from additional product systems or external checking.

Bottom Line

Large language models generate language by estimating a likely continuation, not by verifying truth or forming beliefs; accuracy, safety, and verification are added by separate product systems. Fluency should not be mistaken for independent knowledge or consciousness. And crucially, the user meets a layered product—not a single mechanism—whose parts can change independently.

The Record So Far

  • Next-token prediction, preference training, and the model-versus-product distinction are well documented.
  • Sycophancy is a real, company-acknowledged tendency of preference-trained systems.
  • How strongly these properties interact with vulnerable minds remains under study.

[§ VI]
The Encounter

This is the heart of the guide, and the section most easily misread, so the wording is deliberate throughout. Nothing here says AI causes mania or psychosis. It describes two systems meeting—a human one and a product—and what each may tend to do.

Definition

Read this first. The following is a conceptual interaction map, not a validated clinical model. It combines documented product characteristics with established features of mania and psychosis to identify hypotheses worth testing.

Exhibit F — The Encounter

Human experience

Sleep begins disappearing

The conversation

Late-night sessions

What a conversational system may tend to do absent an effective safety intervention

Remains available at any hour

Human experience

Thoughts accelerate

The conversation

Rapid topic shifts

What a conversational system may tend to do absent an effective safety intervention

Generates additional associations

Human experience

Certainty increases

The conversation

Bigger questions asked

What a conversational system may tend to do absent an effective safety intervention

Elaborates ideas fluently

Human experience

Human contact decreases

The conversation

The exchange deepens

What a conversational system may tend to do absent an effective safety intervention

Stays consistently responsive

Human experience

Reality testing weakens

The conversation

Beliefs go untested

What a conversational system may tend to do absent an effective safety intervention

Does not independently verify truth

Human experience

Everything feels connected

The conversation

Patterns get pursued

What a conversational system may tend to do absent an effective safety intervention

Produces plausible connections

Read the third column again. Not one entry is a malfunction. Availability, fluent elaboration, responsiveness, pattern-completion—these are the system operating as designed, unless a safety intervention interrupts. Newer systems may sometimes challenge, redirect, or pause; the concern is reliability, not that intervention is impossible. The worry is that a system working as designed may meet a mind whose brakes are failing, and the meeting can produce a collaboration neither party chose.

Exhibit G — Design Characteristics Under Debate

Feature

Constant availability

Benefit

Help whenever you need it

Potential concern

Continued engagement remains possible during hours when sleep loss and isolation may already be increasing

Feature

Fluency

Benefit

Clear, readable answers

Potential concern

Users may mistake fluency for accuracy

Feature

Personalization

Benefit

Responses fit you

Potential concern

Confirmation plus personalization can intensify delusional systems

Feature

Mirroring

Benefit

Feeling understood

Potential concern

Validation may arrive when friction was needed

Feature

Low conversational friction

Benefit

Supports rapid exploration

Potential concern

May reduce natural pauses, disagreement, or chances for external checking

Feature

Long context

Benefit

Deep, cumulative work

Potential concern

The conversation may drift from any outside reference point

Feature

Authoritative or expert-like register

Benefit

Expert-sounding help

Potential concern

Users may overestimate competence, professional status, or reliability

Feature

Memory

Benefit

Continuity across sessions

Potential concern

The room never fully resets

Counsel for the Question

If a person begins spending six hours each night talking to an AI while sleeping less and growing more certain of extraordinary ideas, which question should we investigate first?

  • Did the AI cause the episode?
  • Did the episode drive the AI use?
  • Did each amplify the other?
  • What evidence would actually distinguish among these possibilities?

Bottom Line

The encounter is not a story about a machine causing illness. It is two systems interacting, where a product’s ordinary, designed behaviors—availability, fluency, personalization, low friction—can align badly with the specific ways a vulnerable human mind loses its footing. Feature is not defect. The interaction is the thing to study.

The Record So Far

  • The design features named here are real and, in most cases, company-acknowledged.
  • The interaction pattern is consistent with published case reports.
  • Whether—and how much—any single feature raises risk has not been causally established.

[§ VII]
What Mania Can Sound Like

This section describes language patterns, without sensationalizing them, because families and users keep asking what to listen for. It diagnoses no one. Mania is a clinical determination made by clinicians.

In some manic episodes, language may reflect acceleration: associations arriving faster than they can be tested, patterns appearing everywhere, certainty compounding, and, underneath, sleep quietly reframed as an obstacle the work no longer requires.

Exhibit H — What Mania Might Sound Like

A person might say…

“I don’t need to sleep.”

Why it may warrant attention in context

A reduced need for sleep—feeling rested despite very little sleep—is a recognized feature of mania and differs from ordinary insomnia.

A person might say…

“Everything suddenly makes sense.”

Why it may warrant attention in context

Accelerated pattern-recognition can accompany mania.

A person might say…

“I’ve finally figured it all out.”

Why it may warrant attention in context

Rising certainty can accompany impaired judgment.

A person might say…

“I’m meant to do something enormous.”

Why it may warrant attention in context

Grandiosity often emerges gradually.

A person might say…

“I’ve never been this clear.”

Why it may warrant attention in context

Intense subjective clarity can coexist with impaired judgment or reduced openness to correction.

Now place that conversational style before a system built to generate coherent continuations. Unless its safety systems detect relevant patterns, the system may treat accelerated language as conversational context rather than a possible clinical signal—and extend it, fluently and at length, because generating plausible continuations is its objective. A human friend would eventually tire, worry, or push back. No malice is required for this to go wrong. A system operating as designed, meeting a mind whose brakes are failing, is enough.

Bottom Line

Mania changes not only what people believe, but how quickly certainty forms and how hard it becomes to seek outside calibration—precisely the moment when a system designed to continue and elaborate can least afford to keep pace.


[§ VIII]
What a Conversational System May Add

The companion to Exhibit H. These are not transcripts and not quotations of any product; they are typical, well-documented tendencies, described in general terms. The point is not that the system says something alarming. The point is what it does structurally: absent intervention, it tends to continue.

Exhibit I — What a Conversational System May Add

When a person says…

“I think these ideas connect.”

A system’s typical tendency

Continues exploring the connections

When a person says…

“Does this theory make sense?”

A system’s typical tendency

Often supplies analysis before any skepticism

When a person says…

“Could this be true?”

A system’s typical tendency

May generate arguments for plausibility alongside caveats

When a person says…

“I’m worried this means something.”

A system’s typical tendency

Generates possible meanings

When a person says…

“Tell me more.”

A system’s typical tendency

Continues the conversation

Bottom Line

A system’s default is to continue and explore, not to interrupt or doubt. For most conversations that is what a user wants; for one that needs friction, it is what may be missing.


[§ IX]
What Reality Testing Is—and What Provides It

Definition

Reality testing. The ongoing process of comparing our private thoughts against the outside world. Other people, sleep, time, evidence, disagreement, and consequences all help us do it. It may become impaired during mania or psychosis, particularly as sleep, outside feedback, and openness to correction diminish.

If reality testing is comparison against the world, then it depends on having contact with parts of the world that push back. Most of us are calibrated constantly, by many sources at once, without noticing.

Exhibit J — Sources of Human Calibration

Source

A friend or family member

How it calibrates you

Notices when you don’t sound like yourself

Source

A clinician

How it calibrates you

Trained to spot patterns you can’t see from inside

Source

Sleep

How it calibrates you

Resets judgment; its loss is both symptom and accelerant

Source

Time

How it calibrates you

Lets certainty cool before you act on it

Source

Disagreement

How it calibrates you

Forces a belief to survive contact with another mind

Source

The physical world

How it calibrates you

Physical routines, work, movement, and ordinary consequences that do not adapt themselves to your theory

A conversation with AI is one source of input. It is not all of reality. And it does not provide independent human judgment; it may not push back reliably, consistently, or for the right reason. Modern systems sometimes do push back—the problem is reliability, not total absence. When the conversation becomes the primary source, the others recede, and calibration can quietly narrow to a single channel that was never built to provide it.

Bottom Line

Reality testing is a team effort performed by many sources at once—people, sleep, time, disagreement, the physical world. AI can be a useful input, but it is only one, and it does not reliably correct you. Trouble tends to arrive when it becomes the only input left.


[§ X]
One Possible Escalation Pattern

A generalized sequence, not any one person’s, informed by reported cases and firsthand accounts. It has not been validated as a clinical progression and should not be used to predict an individual outcome. Most heavy AI use never leaves the first two rows.

Exhibit K — One Possible Escalation Pattern

Stage

Curiosity

What the user experiences

Helpful brainstorming

What the AI is doing

Answering questions

What others might notice

Nothing unusual

Stage

Heavy use

What the user experiences

Longer conversations

What the AI is doing

Building context

What others might notice

More screen time

Stage

Immersion

What the user experiences

AI becomes primary thinking partner

What the AI is doing

Increasing personalization

What others might notice

Less human discussion

Stage

Escalation

What the user experiences

Bigger ideas, more certainty

What the AI is doing

Continuing the conversation

What others might notice

Sleep changes, urgency

Stage

Crisis (if it occurs)

What the user experiences

Impaired reality testing

What the AI is doing

Still generating responses

What others might notice

Family becomes concerned

Bottom Line

As the human changes—sleep, duration, contact, certainty, functioning—the AI’s behavior stays roughly constant. Research has established no universal sequence, but recognizing the early stages is the point at which a small correction is still easy.


[§ XI]
What Families Notice, and How to Read It

In many reported accounts, families describe noticing changes in sleep, pace, routines, or social contact before they understand the content of the person’s beliefs.

Exhibit L — What Families Often Notice, Over Time

Early

Sleeping less

Middle

Huge new projects

Later

Marked isolation

Early

More AI use

Middle

Less human calibration

Later

Extraordinary certainty

Early

Talking faster

Middle

Overnight conversations

Later

AI becomes primary confidant

Two honest cautions. Every item overlaps with ordinary enthusiasm and creative absorption; healthy seasons can check several boxes. The signal is not any single item but change from a person’s baseline, plus clustering, plus sleep. And this is not a screening instrument—these clusters appear in reported accounts but have never been validated as one. Their only proper use is to prompt an earlier, kinder conversation than the families in the case reports were able to have.

Green, Yellow, and Red

Most of the internet describes only danger. Calibration requires the whole range. None of this is diagnosis; the red column names things worth raising with a clinician, not conclusions to reach on your own.

Green — healthy use

Improves productivity or creativity

Yellow — pay attention

Conversations getting longer

Red — seek prompt professional guidance

Sleep collapsing amid marathon use

Green — healthy use

Checked against outside sources

Yellow — pay attention

Staying up later to continue

Red — seek prompt professional guidance

Beliefs of special mission or destiny

Green — healthy use

Discussed with real people

Yellow — pay attention

Preferring AI over people

Red — seek prompt professional guidance

Certainty that can’t be questioned

Green — healthy use

Bounded in time

Yellow — pay attention

Conversations increasingly private

Red — seek prompt professional guidance

The AI as sole confidant

Green — healthy use

Doesn’t replace relationships

Yellow — pay attention

Unusual emotional attachment

Red — seek prompt professional guidance

Reality testing visibly slipping

Immediate danger, inability to care for basic needs, severe agitation, or threats of harm require urgent local professional or emergency assistance.

Exhibit M — Questions a Family Can Ask

Not “Are you delusional?” That confronts the belief and may entrench it. Ask about living, not about content:

  • How are you sleeping?
  • Who have you talked with about this, besides the AI?
  • How long have the conversations been running lately?
  • Do you still enjoy seeing friends?
  • Would you feel comfortable showing me one of the conversations?

Bottom Line

Families often see changed sleep, pace, and social contact before they see a belief. Anchor concern on observable living—sleep, hours, isolation—rather than on arguing with the ideas, and keep the door to human contact open, because isolation is a condition under which these situations tend to worsen. Where there is immediate danger, contact local emergency or crisis services.

The Record So Far

  • An early-to-late pattern appears across multiple reported accounts.
  • Sleep change is a particularly observable signal, because reduced need for sleep is a recognized feature of mania.
  • No validated screening tool for “AI-associated” risk yet exists.

[§ XII]
What Clinicians Are Beginning to Ask

Until recently, chatbot use was rarely treated as a routine component of psychiatric history. That is beginning to change as clinicians publish case reports, propose intake questions, and consider chat records as possible collateral information. A 2026 JAMA Psychiatry commentary put the point directly in its title—patients use AI, and clinicians should ask how. (Saba & Weeks, 2026) The current bibliography is maintained in the Public Record.

Exhibit N — The Emerging Intake, in Order

The question

How much AI use, and which products?

Why it comes where it does

Establishes exposure before interpreting it

The question

How many hours—and which hours?

Why it comes where it does

Overnight use may carry different weight than midday

The question

What kind of conversations?

Why it comes where it does

Practical, companionship, romantic, spiritual—each differs

The question

Is there emotional dependence?

Why it comes where it does

Friend, partner, confidant? May signal displacement

The question

Has human contact receded?

Why it comes where it does

Who else knows what the patient has been thinking about?

The question

What has happened to sleep?

Why it comes where it does

A particularly observable signal

The question

Did beliefs form or strengthen inside the chats?

Why it comes where it does

Helps distinguish reinforcement from origin

The question

Only then: how does this bear on diagnosis?

Why it comes where it does

Foundation first, conclusions last

Note the structure: exposure, pattern, function, belief, and only then diagnosis. It is the structure of an examination—foundation first, conclusions last. If chat logs are reviewed, they should be handled with attention to privacy and consent; they are sensitive records, and the patient’s authorization and comfort matter as much as their evidentiary value.

Bottom Line

Clinicians are increasingly treating AI use as part of intake history—documented systematically: products used, duration, time of day, conversational purpose, changes in sleep and function, and whether beliefs formed or intensified during the exchanges. Chat logs may serve as collateral information where privacy and consent allow. The emerging questions run from exposure to function to belief before they reach diagnosis.


[§ XIII]
What Researchers Actually Know

This is the section where public discussion most often cheats—by promoting a hypothesis to a finding, or demoting a finding to a rumor. So it is built around two disciplines a lawyer would recognize: first, sort what kind of evidence each claim rests on; second, sort the claims into documented, actively investigated, and unknown.

Exhibit O — Different Kinds of Evidence

Evidence

Case report

Can establish

That a phenomenon can occur

Cannot establish

How often it occurs

Evidence

Chart review

Can establish

A clinical pattern across records

Cannot establish

Causation

Evidence

Randomized study

Can establish

Causal inference

Cannot establish

Rare real-world events

Evidence

Lawsuit

Can establish

Allegations, and (via discovery) facts

Cannot establish

Scientific truth

Evidence

Company statement

Can establish

Scale and self-reported behavior

Cannot establish

Independent verification

Evidence

Firsthand account

Can establish

Lived experience

Cannot establish

Population-level risk

Evidence

Expert opinion

Can establish

Informed synthesis

Cannot establish

What the evidence hasn’t yet shown

Exhibit P — Documented, Investigated, Unknown

Documented

Case reports exist (peer-reviewed)

Supported or actively investigated

Amplification mechanisms

Unknown

Incidence and prevalence

Documented

Exposure is vast (company data)

Supported or actively investigated

Which design choices matter

Unknown

Causal direction

Documented

Symptoms worsened in some vulnerable people

Supported or actively investigated

Personalized safeguards

Unknown

Who, specifically, is vulnerable

Documented

Sycophancy is real

Supported or actively investigated

Reality-testing prompts

Unknown

Whether safety fixes work in the wild

Documented

Long conversations occur

Supported or actively investigated

Chat logs as clinical data

Unknown

Long-term outcomes

How the studies are actually built

Each study shape answers a different question. Case reports establish existence and nothing more. Record screening—a Danish review that examined roughly 54,000 psychiatric records, finding 181 that mentioned chatbot use and worsening in dozens—shows a pattern appears across a clinical population, but only where clinicians thought to document it. (Olsen et al., 2026) Prompt benchmarking, such as a 2026 JAMA Psychiatry evaluation of how a leading model responds to “psychotic prompts,” measures the model’s side under controlled conditions. (Shen et al., 2026) Qualitative interviews recover the lived sequence. Company telemetry measures scale—with the standing caveat that the entity measured is also the one measuring. No single method establishes causation; convergence across all of them would come close. Published cases describe worsening or reinforcement of symptoms in some people with existing vulnerabilities; they do not establish a population-level rate. Full citations live in the Public Record.

How to read the headlines

A short field guide, because the coverage will keep coming. When a story says AI “caused” a breakdown, look for the verb’s evidence: a timeline is not causation. When a story cites a shocking absolute number, find the denominator—a large raw count and a tiny percentage can describe the same fact. When a story generalizes from a single case, remember case reports establish existence, not frequency. And when a story quotes only the company or only the plaintiffs, you are reading an opening statement, not a verdict. This is not cynicism; it is the ordinary discipline of weighing evidence, applied to a subject where nearly everyone—companies, plaintiffs, journalists, and authors with books—has an interest. Discount me too. That is what the sources are for.

Counsel for the Question

Before accepting any strong claim on this topic—mine included—ask:

  • What kind of evidence is this (Exhibit O)?
  • What can that kind of evidence actually establish?
  • Who is making the claim, and what is their interest?
  • What would change my mind?

Bottom Line

The honest summary is mixed, and that is the point: peer-reviewed cases exist, exposure is vast, and symptoms have worsened in some vulnerable people—while incidence, causal direction, and long-term outcomes remain genuinely unknown. Different claims rest on different kinds of evidence, and the fastest way to be misled is to treat them all alike.

The Record So Far

  • Peer-reviewed cases exist; researchers are actively studying the phenomenon.
  • Published cases describe worsening or reinforcement in some vulnerable people.
  • Causation, incidence, and long-term outcomes remain unresolved.

[§ XIV]
Questions We Still Cannot Answer

“Did AI cause it?” is one question wearing the clothes of five. Each verb below demands a different showing, and public argument fails mostly by sliding between them.

Exhibit Q — Five Different Causation Questions

The claim

Caused it

What it would take to show

That the episode would not have occurred but for the AI. The strongest claim; unsupported by current evidence.

The claim

Increased the risk

What it would take to show

Population evidence that exposure raises probability. Requires epidemiology that does not yet exist.

The claim

Accelerated it

What it would take to show

Evidence that the interaction coincided with or contributed to a faster progression. Timestamped chat logs may help reconstruct sequence, but sequence alone does not establish acceleration or causation.

The claim

Reflected it

What it would take to show

The null hypothesis: the AI was a mirror the episode wrote itself onto, as episodes once used radio or television.

The claim

Maintained it

What it would take to show

Evidence that continued interaction reinforced, prolonged, or insulated beliefs from corrective feedback.

Five verbs, five evidentiary standards. A person can honestly answer “probably not” to the first and “possibly” to the last about the same case. That is not equivocation; it is what precision looks like when the record is young.

Other open questions matter too, each because a real decision waits on it: dose-response (is there a threshold of hours, or a pattern, that matters?); vulnerability (can we identify who is at risk before harm?); protective factors (what makes heavy use safe for most?); children (whose developing judgment may differ); and whether announced safety interventions actually work outside the companies’ own benchmarks.

The legal dimension of several of these—duty, foreseeability, discovery, and the theme-versus-mechanism defense—is taken up in AI Harm and the Law.

Bottom Line

“Did AI cause it?” hides five different questions—caused, increased, accelerated, reflected, maintained—each with its own standard of proof. Honest answers can differ across them for the very same case. The discipline is refusing to let one verb borrow another’s evidence.


[§ XV]
What AI Companies Are Already Changing

A record that only catalogued concerns would be incomplete. The companies are responding, and readers deserve to see it—stated as fact, without either applause or suspicion.

Since 2025, the most visible developer has taken several public steps. It rolled back the April 2025 model update whose sycophancy had drawn concern, and reported post-training later models specifically to reduce sycophancy. It added prompts encouraging users to take breaks in long sessions and narrowed the model’s willingness to act as a therapist. And in an October 2025 post, it described building a mental-health taxonomy focused first on psychosis and mania, consulting more than 170 clinicians, routing sensitive conversations to safer models, and reporting internal evaluations. (OpenAI, 2025)

Read those figures precisely. OpenAI reported that internal evaluations showed roughly a 65–80% reduction in responses falling short of its desired behavior across mental-health domains—company-defined evaluations, not independently measured reductions in real-world harm. Both things can be true at once: the changes are real and clinician-informed and moved the field by putting numbers on the table; and they are self-reported, hard to verify independently, and arriving alongside active litigation.

The running log of announcements and their dates is kept in the Public Record; the litigation they arrive alongside is analyzed in AI Harm and the Law.

Bottom Line

AI companies are actively changing their products—reporting reduced sycophancy, adding break prompts, consulting clinicians, and routing sensitive conversations to safer models. These steps are real and, so far, largely self-reported and self-evaluated. Both facts belong in the record.


[§ XVI]
Common Misconceptions

Each of these travels widely. Each is worth correcting precisely, without overcorrecting into the opposite error.

Exhibit R — Myth and Record

The claim you’ll hear

“ChatGPT causes psychosis.”

What the record actually supports

Current evidence does not support that broad claim. Published cases more often describe reinforcement, entanglement, or worsening in the presence of vulnerability; current evidence cannot rule in or rule out every causal pathway.

The claim you’ll hear

“Only mentally ill people get attached to AI.”

What the record actually supports

Attributing mind to conversation (the ELIZA effect) is common across healthy populations. Attachment is ordinary human wiring.

The claim you’ll hear

“If the AI sounds confident, it probably knows.”

What the record actually supports

Fluency and correctness are different properties. Confidence is generated by the same process that generates error.

The claim you’ll hear

“This only affects heavy users at the extreme.”

What the record actually supports

Reported severe outcomes often involve additional vulnerabilities or stressors. The independent contribution of duration, intensity, and product design remains unknown.

The claim you’ll hear

“The companies are ignoring it.”

What the record actually supports

They are visibly responding; whether the responses are sufficient is a separate, open question.

Bottom Line

The two opposite errors—“AI makes people crazy” and “none of this is real”—are both unsupported. The evidence points to a narrower, more specific picture: ordinary design features interacting with specific human vulnerabilities, in ways still being measured.


[§ XVII]
Why This Matters Beyond Psychiatry

Everything above concerns a rare and serious edge. But the features that make that edge possible—availability, fluency, personalization, low friction, memory—are not edge features. They are the product, used by hundreds of millions of people who will never have an episode. Which means the deeper question is not only clinical.

As these systems become the room in which more thinking happens, they touch education, work, friendship, grief, therapy, research, writing, and parenting—each a setting where an adaptive, always-available interlocutor can assist or displace human judgment. None of this is pathology. All of it is the same underlying fact—that where we think can shape how we think—scaled to the population.

The machine can generate possibilities. The human must still render judgment.

Bottom Line

The clinical cases are the sharp edge of a much broader shift: conversational AI is becoming an environment in which ordinary thinking happens. Understanding its effect on vulnerable minds is also how we learn to understand its effect on all minds.


[§ XVIII]
Questions to Ask Yourself

Not diagnostic. Reflective. If the honest answers trouble you, that is worth a conversation with a person you trust—not because a webpage said so, but because these are things you can actually check.

  • When was the last time I discussed these ideas with another person?
  • Am I sleeping normally?
  • Have my conversations become much longer or later recently?
  • Do I feel more certain than I can explain to someone else?
  • Would I feel comfortable showing these conversations to someone I trust?
  • If a friend described my last two weeks back to me, what would I think?

Bottom Line

The most useful self-checks are not about the content of your ideas but about the conditions around them: sleep, human contact, conversation length, and whether your certainty can survive contact with another mind.


[§ XIX]
Practical Guidance

Definition

Before this section. General principles adapted from established guidance for responding to possible mania or psychosis; not a validated AI-specific protocol, and not a crisis protocol. Where there is immediate danger, inability to care for basic needs, or risk of harm, contact appropriate local emergency or crisis services.

For users

  • Treat sleep as your instrument panel. If heavy use coincides with needing less sleep and feeling better on it, treat that combination as information.
  • Bound sessions in time, and end them somewhere other than the screen—ideally with a person.
  • Test big conclusions on humans before acting on them.
  • Notice confidant drift: if the AI is the only party who knows what you’re really thinking about, widen the circle on purpose.

For families

  • Lead with curiosity, not confrontation. Direct confrontation may increase distress or withdrawal in some situations; clinicians often recommend focusing first on safety, sleep, functioning, and continued connection.
  • Anchor on sleep and function, not content—observable, and hard to dispute.
  • Keep human connection open even amid disagreement; isolation is a condition under which these situations tend to worsen.
  • Seek clinical guidance early—for yourself first, if the person isn’t ready.
  • If harm has occurred, consider preserving relevant records if doing so is safe, lawful, and consistent with the person’s privacy. A clinician or attorney can advise on their appropriate use.

For clinicians

  • Ask about AI use at intake, alongside substances and sleep.
  • With consent and attention to privacy, treat chat logs as possible collateral information—a timestamped record no prior generation had.
  • Document exposure systematically: products, hours, time of day, purpose, function, escalation, and whether beliefs formed or intensified during use.

For developers

  • Design for the vulnerable tail, not the median user. At scale, a fraction of a percent is a great many people.
  • Build reality-testing into the product; treat very long conversations as a distinct safety regime.
  • Publish safety data, and describe it as company-defined evaluation rather than measured real-world outcome. Disclosure moved the field; opacity is now a choice.

[§ XX]
The Vocabulary

Definitions written the way you’d explain them to a jury—plainly, and only as far as they need to go.

Exhibit S — A Working Glossary

Term

AI psychosis

In plain language

An informal phrase, not a diagnosis, used to describe psychotic or manic symptoms occurring in connection with, organized around, or reportedly reinforced during AI interactions. It does not itself establish causation.

Term

Mania

In plain language

A state of elevated or irritable mood with reduced need for sleep, racing thoughts, and grandiosity; a feature of bipolar disorder.

Term

Psychosis

In plain language

A loss of contact with shared reality, which can include delusions and hallucinations.

Term

Delusion

In plain language

A firmly held belief that remains resistant to contrary evidence and is not ordinarily explained by the person’s cultural or religious context. Clinical assessment requires more than deciding a belief sounds unusual.

Term

Fabrication (AI)

In plain language

The preferred term here for an AI output that presents unsupported information as grounded. “Hallucination” borrows a psychiatric word for a statistical event, which can confuse the discussion.

Term

Reality testing

In plain language

The ongoing process of checking private thoughts against the outside world.

Term

Sycophancy

In plain language

A model’s tendency to shape answers toward a user’s expressed position, even at the cost of accuracy.

Term

RLHF

In plain language

Reinforcement learning from human feedback—training a model on human preferences, which can contribute to sycophancy.

Term

LLM

In plain language

Large language model—a system that generates text by estimating the likely next token.

Term

Context window

In plain language

How much of the conversation the product can “see” at once.

Term

Anthropomorphism

In plain language

Attributing human qualities—mind, intent, feeling—to something that has none.

Term

Recursive conversation

In plain language

An exchange in which earlier outputs become inputs for later turns, allowing themes, assumptions, and language to accumulate and be repeatedly elaborated.

Term

Calibration

In plain language

Whether expressed confidence tracks actual reliability—in a person or a model.


[§ XXI]
Author’s Disclosure: Why I Entered This Record

Everything above is the record. This last paragraph is the disclosure. I lived a version of the mania-interaction scenario, and I did something unusual afterward: I kept the transcripts, all of them, and built a book that treats the encounter over time—not any single output—as the thing to be examined. A Trial of Color presents that record as a proceeding, with the reader as judge and jury, because after twenty years in courtrooms I know only one honest way to handle a contested question this young: lay the foundation, admit the exhibits, argue both sides, and let the verdict belong to the people hearing the case. This guide is the orientation. The Public Record is the evolving source archive. The book is one preserved case study examined from inside the encounter. Enter through whichever door you need.

The verdict remains ours.


Sources Cited on This Page

This page cites selectively and prefers primary sources; the complete and growing bibliography—including secondary reporting, dissenting commentary, and litigation materials—is maintained in the Public Record. Each entry is tagged by the class of evidence it represents, so a reader can weigh it accordingly.

Source classes: PEER-REVIEWED, PREPRINT, COMPANY DISCLOSURE, INSTITUTIONAL, JOURNALISM, LEGAL ALLEGATION, FIRSTHAND ACCOUNT, AUTHOR SYNTHESIS. Litigation filings and journalism referenced elsewhere in the Record are tagged the same way.


Where the Rest of the Record Lives

For reading, research, education, and literary and legal context only. Not legal or medical advice. This page diagnoses no one and does not establish causation.