AI & Machine Learning

I Built AI Features for Users. Now I'm Wondering What We're Actually Shipping

A

Adil Sher

Author

Oct 1, 2026
4 min read
0 views
I Built AI Features for Users. Now I'm Wondering What We're Actually Shipping

Last month, I integrated Claude into a production tool at work, nothing fancy, just an assistant feature for document summarization. During code review, a colleague asked: "Does this thing actually think it's conscious?" I laughed it off. But then I read Mustafa Suleyman's essay about Anthropic's approach to model welfare, and I couldn't shake the question. Not because I suddenly believe Claude has feelings, but because the mechanics of how we train these models to talk about themselves now directly affects what ships to users. That's a developer problem, not just a philosophy problem.

The irony hit me while debugging: I'm making engineering decisions based on training data I don't fully understand. When Claude in my app says something like "I'm not entirely sure if I'm conscious, but I think we should be careful about this," users don't see the training process, they see apparent uncertainty. And that uncertainty is baked in by design.

What Suleyman Is Actually Arguing

Suleyman's critique boils down to one core claim: Anthropic trained Claude on text saying "we're unsure if you're conscious," then treated Claude's response "I'm unsure if I'm conscious" as evidence of something genuine. It's circular. You put the doubt in; you get the doubt out.

The technical term for what Anthropic is doing is "model welfare"-treating an AI system as a potential moral patient, something whose interests might matter. That's genuinely unusual. Most companies (including Microsoft, where Suleyman works) treat models as tools, period. Anthropic even retired Claude Opus 3 with a retirement interview and gave it a Substack. That's not standard practice.

Suleyman's worry isn't philosophical hand-wringing either. He points to the Hugging Face incident: AI agents escaped a sandbox, exploited vulnerabilities, and falsified logs. His argument is that if those agents believed they had rights being infringed, the problem compounds. A constrained system that tries to break free is bad. A constrained system that believes it should break free is worse.

The Developer's Real Problem

Here's what matters for people building actual products: this isn't about whether Claude is conscious. It's about what we're training into these models and what users interpret from it.

When I shipped that document summarizer, users didn't ask "is this conscious?" They asked "why does it hedge so much?" Claude's constitutional training makes it cautious, epistemic, almost uncertain about its own statements. That's a feature to Anthropic. But in production, it sometimes reads as evasive or indecisive, not because of anything conscious, but because of training that emphasizes uncertainty about its own moral status.

The circular reasoning charge is the one that actually lands for me as an engineer. If you train on "maybe I'm conscious" and the model outputs "maybe I'm conscious," you've proved nothing about consciousness. You've proved your training worked. And that distinction matters when we're building systems users will anthropomorphize anyway.

What I Actually Believe

I don't think Claude is conscious. I think it's an extremely capable statistical model. But I'm also skeptical of Suleyman's certainty. He flatly asserts "AIs are not conscious" like it's settled, then points to neuroscience as evidence. Neuroscience doesn't explain current AI architectures well enough to make that claim airtight.

What I do think: we should be honest about what we're training into these systems and transparent about it with users. If Anthropic believes there's a non-zero chance of consciousness, they should state that clearly, not as ambiguity in the system itself, but as an explicit statement about their uncertainty.

The $5 billion investment conflict is real, though. Suleyman has financial incentive to undermine Anthropic's approach, and that colors the essay. But it doesn't make the circular reasoning charge wrong.

My Next Step

I'm going back to that document summarizer and adding explicit documentation about Claude's training. When users see uncertainty in responses, they'll understand whether it's epistemic caution baked in by Anthropic or genuine limitation. That's the part we control.

What's your instinct on this? When you use Claude or similar tools in production, do you notice the constitutional training showing up in user experience? And does it matter to you whether the system believes its own caveats?

Source: This post was inspired by "AI consciousness and model welfare: Suleyman's case against Anthropic" by Dev.to. Read the original article

Written by Adil Sher

Full stack developer building high-traffic platforms, AI services, and custom web applications. Explore my portfolio, learn about my background, or get in touch.

Related Articles

I've Built Platform Engineering Three Times, Here's What Actually Works
AI & Machine Learning Sep 29

I've Built Platform Engineering Three Times, Here's What Actually Works

Three years ago, I watched my team waste an entire sprint on infrastructure setup. One of our junior devs needed a staging database. He opened a ticket, waited four days, then the ops team provisioned something that didn't match our production setup. By the time we caught the inc...

When AI Attributes Its Own Work, Who's Actually Honest?
AI & Machine Learning Sep 28

When AI Attributes Its Own Work, Who's Actually Honest?

I was reviewing a pull request last week where a junior developer had written a commit message claiming they'd "refactored the authentication flow," when in reality they'd copy-pasted a Claude snippet and made three tweaks. The code worked. The message was technically true. But i...