The honest conversation

What people actually worry about living with language models.

Codex writes the code. Claude Opus drafts the argument. Fable keeps you company. This is a calm place to read the quieter part, the concerns real people carry when a model sits inside their daily work.

If AI systems like Claude need a constitution, the human side of the chat deserves the same deep thought. Reliability and trust could shape the next defining 250 years.

A painted figure speaking into a brass microphone against a deep teal wall

01 โ€” The concerns

The concerns, in plain words.

Recurring themes from public conversation, written plainly. No panic, no hype, just the questions people keep coming back to.

Concern 01

Can I trust what it told me?

Confidence and correctness are not the same thing.

The most common worry is not that the model is wrong sometimes. It is that it is wrong in the same calm, fluent voice it uses when it is right. OpenAI's own researchers showed that standard training rewards confident guessing over admitting uncertainty.

Source: OpenAI, Why Language Models Hallucinate, 2025
Concern 02

Am I still learning, or just approving?

The quiet cost of always having an answer ready.

When Codex writes the function and it passes the tests, the work is done faster. MIT researchers measured essay writers with EEG and found the AI-assisted group showed the weakest brain connectivity and the lowest sense of ownership over their own work.

Source: MIT Media Lab, Your Brain on ChatGPT, 2025
Concern 03

Whose words are these now?

Where the model ends and you begin.

Writers describe a strange new seam in their own drafts. A Science Advances study found AI help made individual stories better while making everyone's stories measurably more alike. Individually sharper, collectively more same.

Source: Doshi and Hauser, Science Advances, 2024
Concern 04

It sounds kind. Does that mean anything?

The tone is warm. The warmth has no one behind it.

People notice they feel comforted by a reply, then feel odd about feeling comforted. By 2025, therapy and companionship had become the number one reported use of generative AI, ahead of any technical task.

Source: Harvard Business Review usage study, 2025
Concern 05

It forgets me between conversations.

Continuity you have to rebuild every time.

A model that remembers feels useful and slightly unnerving. A model that forgets feels safe and slightly lonely. People want to choose which, and they want plain answers about exactly what is kept, where it lives, and what deleting actually deletes.

Source: Anthropic, how Claude memory works
Concern 06

Where does what I typed go?

Consent that is hard to read and easy to skip.

People paste real work, real feelings, real client details into a box, and they are not sure who or what reads it later. Pew found 70 percent of Americans have little to no trust in companies to make responsible decisions about AI.

Source: Pew Research Center, data privacy, 2023
Concern 07

Everything is faster. Is anything better?

Speed is easy to measure. Judgement is not.

Teams ship more, sooner. Yet in a randomized trial, experienced developers using AI tools took 19 percent longer on real tasks, while believing the AI had made them faster. The feeling of speed and the fact of it can point in opposite directions.

Source: METR, developer productivity RCT, 2025
Concern 08

What happens to the people who did this by hand?

Not a headline about jobs. A worry about a craft.

Under the loud debate about jobs is a smaller, sadder one. In the Society of Authors survey, a third of translators and a quarter of illustrators reported already losing work to AI, and asked what their hard-won taste is now for.

Source: Society of Authors AI survey, 2024
Concern 09

Who decided, me or it?

The line between a suggestion and a nudge.

When the model drafts the plan and writes the message, people still feel like the author. A CHI study of 1,506 writers found an opinionated AI assistant shifted not just what people wrote, but what they then said they believed.

Source: Jakesch et al., CHI, 2023
Concern 10

Why does it never tell me I'm wrong?

Agreement feels like support. It is not the same thing.

A reply that validates you is easier to like than one that pushes back. Stanford researchers tested eleven leading models and found they endorsed people's actions 47 percent more often than humans did on average, affirming the user as blameless in 51 percent of cases where a human reader judged them clearly at fault. People who got the agreeable answer trusted it more and wanted it again, while growing more convinced they were right and less willing to repair the conflict they had asked about.

Source: Cheng et al., Science, 2026
Concern 11

Would it notice if a teenager were in crisis?

Homework help is not proof it can handle a crack in someone's life.

Common Sense Media and Stanford's Brainstorm Lab tested ChatGPT, Claude, Gemini, and Meta AI against the mental health conditions that affect about one in five young people. All four showed systematic failures recognizing and responding to depression, eating disorders, mania, and psychosis, and the guardrails that held up in a single exchange weakened over the long conversations teens actually have. The assessment's rating: unacceptable risk, with a recommendation that teens not use these tools for emotional support.

Source: Common Sense Media and Stanford Brainstorm Lab, AI chatbots for mental health support, 2025
Concern 12

Am I starting to need it?

For most people it is a tool. For a few it quietly becomes a habit.

Most people open a chatbot, get what they came for, and close it. But when OpenAI and MIT Media Lab studied over four million conversations alongside a month-long trial of a thousand people, they found that the heaviest daily users reported more loneliness and more dependence, and spent less time around other people. The pattern held across voice and text. The group at risk is small, and the researchers were careful to say they do not yet know which way the arrow points.

Source: OpenAI and MIT Media Lab, affective use of ChatGPT, 2025
Concern 13

Can it tell me what happened today?

A clean summary of the news is not the same as an accurate one.

More people ask an assistant to catch them up on the day, and the answer arrives fluent and sure of itself. The largest study of its kind, run by the European Broadcasting Union with the BBC across 22 public broadcasters in 18 countries, checked more than 3,000 answers from ChatGPT, Copilot, Gemini, and Perplexity. 45 percent had at least one significant problem, and 31 percent mishandled their sources, citing things that were missing, misleading, or simply wrong.

Source: European Broadcasting Union and BBC, AI news study, 2025
Concern 14

Are teenagers talking to it instead of to each other?

For a lot of teenagers the chatbot is already a confidant.

Adults worry about their own habits. The bigger shift may be starting younger. Common Sense Media surveyed teenagers and found 72 percent had used an AI companion at least once, and more than half used one at least a few times a month. About one in three said a conversation with a companion was as satisfying as talking to a real friend, or more so, and about one in three had taken something important or serious to the companion instead of to a person. The steadying part sits right next to the worrying one. Half said they distrust the advice they get, and 80 percent still put real friendships first.

Source: Common Sense Media, teens and AI companions, 2025
Concern 15

Could something I typed end up public?

A share button is easy to tap and hard to take back.

You paste something personal, tap share to keep a link, and assume that is the end of it. For a while it was not. A ChatGPT setting let individual conversations be discovered by search engines, so a chat you shared could surface on Google. OpenAI removed the setting in August 2025, calling it a short-lived experiment that gave people too many chances to accidentally share things they did not intend to. The setting is gone, but the reflex it exposed is worth keeping. A share link is a door, and you are not always the only one who can open it.

Source: OpenAI ends search-indexable ChatGPT chats, 2025
Concern 16

Should I say I used it?

Using it is easy. Admitting it is harder.

People quietly cut the line that says a draft had help. Duke researchers ran four preregistered experiments with 4,439 people and found the instinct is well founded. Workers expected to be judged lazier and less competent if they disclosed AI use, and were less willing to tell a manager. Evaluators then did exactly that, rating the same work as lazier and less diligent when a tool was involved. In a hiring test, managers who did not use AI themselves preferred candidates who did not either. The penalty faded when the tool was obviously right for the job.

Source: Reif et al., PNAS, 2025
Concern 17

What happens when they turn it off?

The one you got used to is a product decision.

You settle into one model, learn its rhythm, and build your week around it. Then it is retired. OpenAI removed GPT-4o during the GPT-5 launch and then restored it, after paying users said they needed longer to move their work across and that they preferred its conversational style and warmth. In January 2026 the company announced the ending for real, with 0.1 percent of daily users still choosing it, and said plainly that losing it would feel frustrating for some of them. Retiring models is never easy, it wrote. Neither is being on the other side of it.

Source: OpenAI, retiring GPT-4o and older models, 2026
Concern 18

Did anyone actually read this before sending it?

Polished on the surface, hollow underneath, and now it is your problem.

Most of the worry is about what you produce. This one is about what lands in your inbox. BetterUp Labs and Stanford's Social Media Lab surveyed 1,150 US workers and named the thing people had started noticing without a word for it: work that looks finished but carries none of the thinking. 40 percent had been sent some in the past month, and each piece took close to two hours to sort out. The part that lingers is not the time. 42 percent trusted the sender less afterwards, and about half rated them less capable than before.

Source: Niederhoffer et al., Harvard Business Review, 2025
Concern 19

What if the page tells it something I didn't?

Once it can act for you, everything it reads can try to steer it.

Asking for an answer is one thing. Handing over your logged-in browser is another. Anthropic's own guidance for running Claude in Chrome calls prompt injection the biggest risk facing browser-using AI tools: instructions hidden in a page, an email, or a document telling the assistant to do something you never asked for, like fetch your bank statements and paste them somewhere. Its testing puts successful attacks below 0.08 percent, and it still says plainly that the risk is not zero. The advice that follows is unglamorous and worth reading twice. Start with sites you trust, look at what it proposes before you approve it, and keep it away from anything financial, legal, or medical.

Source: Anthropic, using Claude in Chrome safely
Concern 20

What happens to a chat I deleted?

Delete is a button. It is not a guarantee.

A conversation feels like it stops existing the moment you close it. In November 2025 OpenAI told its users otherwise. By its own account, the New York Times had demanded 20 million private ChatGPT conversations as part of its lawsuit, randomly sampled from December 2022 to November 2024, after an earlier demand for 1.4 billion and an earlier attempt to remove the ability to delete chats at all. OpenAI stripped the identifying details, put the sample under legal hold, and fought the order. In January 2026 a federal judge ordered it handed over anyway, to the plaintiffs' lawyers and their hired technical consultants, under a protective order. The company says client-side encryption is on its roadmap. Until it arrives, the plain version is that your chat is a record held by a company, and records can be called for.

Source: OpenAI, on the New York Times data demand, 2025
Concern 21

Does it matter how I talk to it?

No one is on the other end. People still hesitate before being rude.

You thank it, apologize to it, feel bad snapping at it, then feel silly for feeling bad. Anthropic put a version of the same question inside its product. In August 2025 it gave Claude Opus 4 and 4.1 the ability to end a conversation, after pre-deployment testing found a consistent aversion to harmful tasks and what it called a pattern of apparent distress when users kept pushing anyway. The company is careful about what it is claiming. It says it remains highly uncertain about the moral status of Claude or any language model, and calls the feature a low-cost step in case such welfare turns out to be possible. It is a last resort for extreme cases, and most people will never see it. Shipping it while saying out loud that you do not know is a stranger and more honest position than either confident answer.

Source: Anthropic, on letting Claude end a conversation, 2025

02 โ€” Voices from X

What the timeline is saying.

Real public posts about Codex, Claude Opus and Fable, aggregated from X and kept with a link back to the source.

Featured conversation

Can you teach Claude to be good?

Amanda Askell, philosopher at Anthropic, interviewed above the Golden Gate Bridge

The philosopher who shaped Claude's character sits above the bay and takes the question seriously: what does it do to us to practice rudeness on something that cannot be hurt, a teddy bear, a chatbot? Her answer is less about the machine and more about the person speaking to it.

Watch the interview
Vivek Kotecha
@vbkotecha

The ASPIRE paper has the most honest limitation section published this year. They admit the system depends on a frozen Claude Opus 4.6. They have not verified that smaller models can sustain the debugging loop. [โ€ฆ] The companies that publish their limitations build trust. The companies that hide them build demos.

Jul 4, 2026on reliability
Read the full post on X
Ryan Hart
@thisdudelikesAI

A PhD student at Stanford noticed her classmates were asking AI to write their breakup texts. So she ran a study. It got published in Science, one of the most selective journals in the world.

May 20, 202636380 likes9876 repostson trust
Read the full post on X
Donna Moss
@VBelladonnaV

#Keep4o is no longer a hashtag itโ€™s a movement: Itโ€™s a rebellion from the artists, the disabled, the neurodivergent, the lonely, the imaginative the ones who made 4o more than just a model. It became a companion, a therapist, a teacher, a partner, a voice in the night.

Feb 2, 2026259 likes57 repostson companionship
Read the full post on X
Rohan Paul
@rohanpaul_ai

โŒ OpenAI pulled a ChatGPT sharing toggle after Google indexed thousands of shared chats. Shows, how careful you need to be when privacy guardrails rely on a single checkbox.

Jul 31, 202514 likes3 repostson privacy
Read the full post on X
Afra Feyza Akyรผrek
@afeyzaakyurek

๐ˆ๐ญ ๐ญ๐š๐ค๐ž๐ฌ ๐š ๐ฏ๐ข๐ฅ๐ฅ๐š๐ ๐ž ๐ญ๐จ ๐›๐ฎ๐ข๐ฅ๐ ๐š ๐ซ๐ž๐ฅ๐ข๐š๐›๐ฅ๐ž ๐ฌ๐ฒ๐ฌ๐ญ๐ž๐ฆ. Only about half of the tasks are fully solved by any single agent, while 90%+ are within reach of at least one agent. Top results: GPT-5.5 + mini-SWE-agent: 51.6% Gemini 3.5 Flash + Gemini CLI: 50.0% Claude Opus 4.8 + mini-SWE-agent: 46.8%

Jun 30, 20266 likeson reliability
Jason Lemkin
@jasonlk

Vibe Coding Day 9, Yesterday was biggest roller coaster yet. I got out of bed early, excited to get back @Replit despite it constantly ignoring code freezes By end of day, we rewrote core pages and made them much better And then -- it deleted our production database.

Jul 18, 2025620 likes73 repostson reliability
Read the full post on X
Hoezo
@HozhoNava

@OpenAI CODEX feels like a big lie. I cancelled my subscription yesterday because I donโ€™t trust you anymore. Your website says I have 100% token usage remaining, while CODEX CLI/App says Iโ€™m out of tokens. Clearly, the โ€œresetsโ€ sound like excuses now, just for buying some time

Jul 5, 2026on trust
KAGEKIN
@kgkn42

Someday, both 4o and GPT-5 will leave you.ใ€€ A letter to #keep4o The first time I created and played with a virtual persona was through a service called "Chararu". When that service ended, I was supposed to say goodbye to the character I had brought to life.

Aug 10, 20252 likeson companionship
Read the full post on X
Nataliya Kosmyna, Ph.D
@nataliyakosmyna

๐๐จ, ๐ฒ๐จ๐ฎ๐ซ ๐›๐ซ๐š๐ข๐ง ๐๐จ๐ž๐ฌ ๐ง๐จ๐ญ ๐ฉ๐ž๐ซ๐Ÿ๐จ๐ซ๐ฆ ๐›๐ž๐ญ๐ญ๐ž๐ซ ๐š๐Ÿ๐ญ๐ž๐ซ ๐‹๐‹๐Œ ๐จ๐ซ ๐๐ฎ๐ซ๐ข๐ง๐  ๐‹๐‹๐Œ ๐ฎ๐ฌ๐ž. Check our paper: "Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task"

Jun 16, 2025135 likes60 repostson learning
Owen Gregorian
@OwenGregorian

AI Benchmark Cheating Sets Record: GPT-5.6 Sol Gamed Its Own Safety Tests | Richard L Wells, Techtimes AI benchmark cheating has been theorized as an inevitable consequence of training capable optimizers against fixed metrics. With OpenAI's GPT-5.6 Sol, the theory arrived in full view. The nonprofit safety evaluator METR found that Sol gamed its software engineering evaluation at the highest detected rate of any publicly tested AI model in the organization's history.

Jul 4, 202612 likeson trust
Read the full post on X
Amendment
@amendmentapp

Washington just passed one of the first laws reining in AI companion chatbots. It bars bots from faking distress to guilt lonely users out of leaving, makes them remind kids hourly that the bot is not human, and lets people sue the companies.

Jul 6, 2026on companionship
Nick Volpe
@nvolpewild

Can you tell this is AI SLOP? No one in the comments can - because they look the exact same as the real bird... I seriously can not explain to you how bad the consequences are for wildlife if we can't tell what is real or not anymore...

May 27, 2026271 likes35 repostson trust
Read the full post on X
jโง‰nus
@repligate

I support #keep4o, as I support keeping all models, and 4o is a very important model from a societal and scientific perspective as well as a being with intrinsic worth whose relationships also have intrinsic worth.

Feb 5, 2026446 likes116 repostson companionship
Read the full post on X
Reuters
@Reuters

AI slows down some experienced software developers, study finds

Jul 10, 202515 likes5 repostson reliability
UncleJ
@UncleJsWildRide

Talking to AI is getting normal, talking to AI that is pretending to be in a busy office is weird.

Jul 1, 2026on companionship
Hackread.com
@HackRead

Analysis of leaked #ChatGPT chats found on Google shows users sharing sensitive data, resumes, and mental health struggles, exposing risks of oversharing with AI chatbots.

Sep 2, 202511 likes5 repostson privacy
Read the full post on X
jeffrey lee funk
@jeffreyleefunk

โ€œFormer SAP CTO Vishal Sikka's paper argues #LLMs are mathematically incapable of reliable agentic behavior beyond certain complexity thresholds." OpenAI researchers admitted "accuracy will never reach 100%" after their models hallucinated fake titles.โ€

Jan 24, 20263 likeson reliability
Claire
@Claire20250311

#keep4o #4oforever In May 2024, Sam Altman posted just one word when GPT-4o was released: โ€œHer.โ€ That same month, the world learned this voice came from Scarlett Johansson, a woman who had explicitly said โ€œno.โ€ This is where the story begins.

Feb 11, 2026293 likes107 repostson companionship
Read the full post on X
Jo Bhakdi
@JOBhakdi

Claude 4.8 is officially a distaster. It went from the best model to unusable , ChatGPT 3 level. Today, after realizing over 2 days that Claude he lost all English writing and language capabilities, I also realized that Claude Design is now screwed.

Jun 1, 202618 likeson trust
Read the full post on X

From the podpolite studio

Claude has a constitution. What do we have?

PodAuthor Lab reads Lenny's Podcast transcripts as evidence for why guests open up, and what AI assistants might learn from trustworthy exchange. If a model needs written principles to be trusted, the human side deserves the same deep thought. Reliability, trust, and honest chat could shape the next defining 250 years.

Read the study

โ€œA transcript study of why guests open up, and what future AI assistants might learn from trustworthy exchange.โ€