I asked ChatGPT about when AI is helpful in writing, and Chatgpt said, โ€œYour mother never stood a chance. ๐Ÿ˜‚โ€

What AI slop, productive struggle, and a typo taught me about how we should design AI for work and education

There is some irony in using AI to write an essay about when we should use AI to write.

So Iโ€™m going to lean into it.

I developed this piece through a long conversation with ChatGPT. Iโ€™m sharing that conversation alongside the finished piece because I increasingly think the interaction between a human and AI may tell us as much as the artifact they eventually produce.

And that may have profound implications for how we use AI at work and in education.

I started with the wrong framework

My original hypothesis was fairly simple:

When the stakes are high, AI should be an assistant. Use it to research, brainstorm, challenge ideas, and edit, but keep the human firmly in control. Words matter. Accuracy matters. Clarity of thought matters.

When the stakes are low, delegate much more aggressively. If AI writes a routine email for me, I don't need to prove that I could have written it myself.

That still seems directionally useful. But during my conversation with AI, another distinction emerged that I think is much more important:

How much of the value of the task comes from the artifact we produce, and how much comes from the thinking required to produce it?

Those aren't the same thing.

If I need to turn five bullet points into a routine email, the email is the value. Spending fifteen minutes crafting the prose myself probably doesn't make me a meaningfully better thinker.

But imagine I'm trying to decide whether my company's strategy is wrong.

Writing isn't merely a mechanism for communicating the answer. Writing may be part of how I discover the answer.

The awkward sentence matters. The paragraph I can't make work matters. The moment when I realize, "Wait, I don't actually know why I believe this," really matters.

Generative AI can eliminate those moments remarkably efficiently.

And sometimes that's a problem.

There is already evidence for the upside. In a randomized experiment involving 453 college-educated professionals, Shakked Noy and Whitney Zhang found that ChatGPT reduced the time required for professional writing tasks by 40% while increasing independently rated output quality by 18%.

AI can absolutely make us more productive.

The harder question is what happens to the thinking underneath that productivity.

AI has made it possible for our prose to become better than our thinking

We all recognize "AI slop" by now.

It's fluent. Organized. Often grammatically immaculate.

And somehow there is nothing there.

I think there's a deeper problem underneath our irritation with overly enthusiastic headings and predictable sentence construction.

Before generative AI, low cognitive investment was often visible in the output. Half-hearted thinking tended to produce half-hearted work.

Generative AI partially decouples those things.

Give an LLM a half-formed thought, and it can turn it into an articulate argument in seconds.

The result can sound much more developed than the thinking underneath it actually is.

AI can make an idea look finished before we've finished thinking.

And emerging research suggests this concern isn't entirely theoretical.

A 2025 study from Microsoft Research surveyed 319 knowledge workers about 936 real-world examples of using generative AI. Greater confidence in AI was associated with less critical-thinking effort, while greater confidence in one's own ability was associated with more. Interestingly, AI didn't simply eliminate critical thinking. It shifted it toward verification, integration, and oversight.

That distinction matters.

AI may not eliminate thinking.

It may relocate it.

Which means we need to get much better at recognizing which thinking should move and which thinking we need to protect.

So what determines whether AI makes our work better?

My own experience suggests at least three things matter enormously:

1. How much I know about the subject.

Expertise helps me recognize when something is wrong, incomplete, superficial, or technically accurate but missing the point.

2. How well I interact with the AI.

The quality difference between "give me an answer" and a sustained conversation involving questions, disagreement, evidence, counterarguments, and revision can be enormous.

3. How much cognitive energy I invest in scrutinizing the result.

This one may be the easiest to overlook.

I can understand a subject well and know how to use AI effectively, but still accept a mediocre answer because I'm tired, rushed, or simply don't care enough about that particular task.

AI makes that incredibly easy because the mediocre answer can look excellent.

This led me to something I think is more important than "prompt engineering."

Judgment may be the defining AI skill

Prompt engineering is useful, but prompts and models will change.

Judgment transfers.

Can I recognize when an answer doesn't quite make sense?

Can I distinguish evidence from inference?

Can I recognize confidence without justification?

Can I identify the important thing that wasn't mentioned?

Can I tell when an argument is technically correct but intellectually shallow?

Can I recognize when the AI is reinforcing my assumptions rather than challenging them?

Ultimately, perhaps AI literacy comes down to three deceptively difficult questions:

What's missing?

What should I ask next?

What deserves skepticism?

This isn't entirely a new conception of AI literacy. HCI researchers Duri Long and Brian Magerko argued as early as 2020 that AI literacy needs to include the competencies required to critically evaluate AI, not merely operate it.

Research on AI-assisted decision making adds another wrinkle. Experiments with "cognitive forcing" interventions have found that requiring people to engage with a problem before accepting AI recommendations can reduce overreliance on incorrect AI advice. But there's a catch: people often like those systems less. The designs that made people think harder weren't necessarily the designs they preferred using.

That sounds like an HCI problem worth taking seriously.

And then, hilariously, the AI gave me a perfect demonstration.

AI confidently misunderstood my typo

During our conversation, I said this could be an interesting thesis for the advanced degree in "CHI" I want to pursue.

I meant HCI: Human-Computer Interaction.

ChatGPT confidently interpreted CHI differently and produced several paragraphs based on that interpretation.

The response was excellent.

Except it was wrong.

I immediately recognized the mismatch because I knew what I meant and understood the context.

The AI had generated a plausible interpretation and continued fluently.

The human exercised judgment.

That tiny mistake captured the entire problem beautifully.

Fluency is not understanding.

Plausibility is not correctness.

And the better AI becomes at producing plausible language, the more important our ability to recognize the difference becomes.

What if we stopped asking students not to use AI?

This is where the conversation became much more interesting to me.

Our current debate about AI in education often revolves around a relatively narrow question:

Did the student use AI to write the paper?

What if that's the wrong question?

Imagine an assignment where students are explicitly told:

Write this paper with AI.

But instead of grading only the finished paper, the student's AI conversation becomes part of the artifact being evaluated.

Now the professor can potentially observe something that traditional homework often obscures: portions of the student's thinking process.

Did the student accept the first answer?

Did they notice contradictions?

Did they ask for evidence?

Did they challenge a weak claim?

Did their questions become more sophisticated?

Did they ask for a counterargument?

Did they realize their original position was wrong?

Did they distinguish what they knew from what they merely thought they knew?

The transcript could potentially provide evidence of domain knowledge, inquiry, verification, skepticism, synthesis, and metacognition.

Instead of:

Assignment โ†’ black box โ†’ paper โ†’ grade

we could begin observing:

Initial understanding โ†’ inquiry โ†’ AI response โ†’ evaluation โ†’ challenge โ†’ verification โ†’ revision โ†’ synthesis

The AI conversation stops being merely evidence of whether someone "cheated."

It potentially becomes evidence of learning.

I want to emphasize potentially here. Whether interaction traces can reliably measure understanding, rather than simply measuring a student's ability to perform the behaviors that look like critical thinking, is an empirical question. Recent researchers are beginning to analyze student-AI dialogue through a metacognitive lens, but this is far from a solved measurement problem.

And that distinction feels incredibly important.

The problem is productive struggle

Education already has language for something we kept rediscovering in our conversation: productive struggle and the closely related research on "desirable difficulties."

Learning science has repeatedly demonstrated that conditions which make learning feel harder in the moment can sometimes improve later retention and transfer. Performance during practice is not the same thing as learning.

Generative AI creates an interesting collision between that insight and traditional software design.

Software generally tries to reduce friction.

Fewer clicks. Less effort. Faster completion.

But if the objective is learning, minimizing effort can be exactly the wrong optimization.

The question becomes:

Which work is unnecessary friction, and which work is productive struggle?

If I'm studying history, having AI correct my spelling probably removes irrelevant friction.

Having AI construct the historical argument for me might remove the exact cognitive work I'm supposed to be developing.

But context matters.

If I'm already a historian and I'm trying to turn an argument I've spent months developing into accessible prose, AI drafting may be extraordinarily useful.

Same technology.

Same behavior.

Different objective.

Different effect.

We now have remarkably direct evidence of this problem

A 2025 PNAS field experiment involving nearly 1,000 high-school math students tested two different versions of an AI tutor.

One behaved more like a standard ChatGPT interface.

The other had safeguards designed to support learning rather than simply provide solutions.

Both dramatically improved students' performance while they had access to AI.

Then researchers took the AI away.

Students who had used the relatively unrestricted GPT interface performed 17% worse than students who had never had access to it. The negative learning effect was largely mitigated in the safeguarded tutoring condition. Researchers found that students with unrestricted access frequently used AI as a crutch, asking for and copying solutions. ๎ˆ€cite๎ˆ‚turn1search0๎ˆ‚turn1search2๎ˆ

Think about that distinction.

AI made the students better at the task in front of them.

It did not necessarily make them better at the underlying capability.

Performance and learning had diverged.

That might be one of the most important distinctions we need to understand about AI.

Khan Academy is approaching the same question from the product side

Khan Academy has spent the past several years developing Khanmigo, its AI tutor.

I find one of its current product metrics particularly fascinating: next-item correctness.

Khan Academy isn't only measuring whether Khanmigo helps a student solve the current problem. It measures whether the student correctly solves the next problem on the same skill without Khanmigo's help.

In other words:

Did the AI produce successful performance?

is different from:

Did the human learn?

From October 2025 through April 2026, Khan Academy ran roughly 20 substantive product tests across more than 15 million tutoring threads. Giving Khanmigo structured information about students' recent performance and prerequisite skill gaps improved next-item correctness by a combined 6.1%. They also measure whether interactions involve active cognitive engagement rather than passive receipt of information.

This is particularly interesting because it moves the question out of philosophy and into product design.

What should the AI actually do differently?

Maybe AI needs to understand what cognitive work we're trying to preserve

Today's AI systems are extraordinarily good at understanding tasks.

"Help me analyze this."

"Write this."

"Summarize this."

"Find the problem."

But they don't necessarily know why I am doing the task.

Compare:

"Help me calculate this because I need the answer."

with:

"Help me calculate this because I'm learning how to calculate it."

The surface task is identical.

The optimal AI behavior is completely different.

The first user may benefit from aggressive automation.

For the second user, aggressive automation may remove the learning objective itself.

This suggests a fascinating design challenge:

Could an AI determine which cognitive work is valuable for the human and adjust its level of assistance accordingly?

Sometimes answer.

Sometimes explain.

Sometimes ask a question.

Sometimes provide a hint.

Sometimes challenge an assumption.

Sometimes require the human to commit to an answer first.

And perhaps occasionally say, in effect:

"I can do this for you, but doing this part yourself may be important to what you're trying to learn."

Not because struggle is inherently virtuous.

Because the objective isn't minimizing human effort.

Potential minus interference

This is where another idea that has influenced how I think about people and systems becomes useful.

In The Inner Game of Tennis, W. Timothy Gallwey introduced the idea commonly expressed as:

Performance = Potential โˆ’ Interference.

I encountered Gallwey's work through Brenรฉ Brown, and I've found the framework useful far beyond tennis.

I wonder whether it gives us a useful lens for AI, too.

AI can dramatically expand human potential.

It can give us personalized tutoring, synthesis, research assistance, brainstorming, critique, translation, expertise, and creative collaboration at a scale we've never had before.

We shouldn't artificially constrain that potential because we're nostalgic for doing things the hard way.

Instead, we should ask:

What interference can AI remove?

And equally:

What valuable human capability might it accidentally remove with it?

The goal of good human-AI interaction may not be to minimize human effort.

It may be to minimize low-value human effort while preserving, or even increasing, high-value human thought.

And I want to be clear about the epistemic status of that idea.

It's a hypothesis.

There is research supporting important pieces of it: AI can increase productivity; AI can change critical-thinking behavior; overreliance is real; productive difficulty can improve learning; unrestricted generative AI can improve immediate performance while harming subsequent independent performance; and differently designed AI systems can produce different learning outcomes. ๎ˆ€

But those pieces do not yet prove the broader framework.

For example:

Does expertise reliably make human-AI collaboration better?

Can we distinguish genuine judgment from someone merely performing skeptical-looking behaviors?

Can an AI reliably recognize when a user's fluency exceeds their understanding?

Can interaction traces validly measure learning?

Can AI determine when struggle is productive versus simply frustrating?

Does AI-supported judgment transfer to situations where AI is absent?

And perhaps the question I find most interesting:

Can we design an AI that exercises good enough judgment about the human to help the human develop better judgment about the AI?

I don't know.

That's why I want to keep pulling on the thread.

Which brings me back to this essay

AI helped write it.

A lot.

I brought the initial hypothesis and my experiences using AI. During the conversation, the model introduced research, challenged my stakes-only framework, helped distinguish performance from capability, and helped connect my ideas to concepts from learning science and HCI.

I challenged it.

I extended ideas.

I caught it being wrong.

It helped me see connections I hadn't seen.

Then, before publishing, I asked it to research the claims we had generated together and look for evidence rather than treating our increasingly compelling conversation as evidence itself.

And now I'm deciding which of those ideas I actually believe.

I'm sharing the transcript because I think you should be able to judge that process for yourself.

Maybe authorship in an AI world becomes less about proving that every sentence originated inside one person's head.

Maybe the more meaningful question becomes:

Where did the judgment happen?

Because the future probably doesn't belong to humans who can write without AI.

And it probably doesn't belong to humans who can get AI to write for them.

It may belong to people who know when to trust it, when to challenge it, what to ask next, what's missing, and what deserves skepticism.

If you're still reading...

I'm currently looking for my next product leadership opportunity. I'm particularly interested in complex products at the intersection of data, AI, research, decision support, and human behavior, and in teams thinking seriously about what AI should enable humans to do rather than simply what AI can automate.

I'm also increasingly curious about this as a potential future HCI research direction.

If you're an educator, HCI or learning-sciences researcher, institution, AI lab, or organization working on human-AI collaboration, metacognition, AI literacy, productive struggle, educational AI, or the development of human judgment in AI-mediated environments, I'd love to connect.

I have considerably more questions than answers.

At this stage, I think that's exactly where I should be.