AI/ML
August 15, 2026
In April, I published a set of commitments about how I intended to work with AI. One of them read: "I'll become proficient at prompting. This is the new literacy."
I was half right.
Everything since has convinced me that prompting is the easy half. It's the half that improves on its own, almost as a side effect of volume. Across a team, across enough projects, prompts get sharper without anyone running a training session. People carry more context, ask better, get more back. That progress is real and it's visible.
What doesn't improve on its own is the ability to tell when the answer is wrong. I've watched that gap widen in proportion to how much we produce, in my own work and across the studio.
That gap has a name, and putting a name to it changed how I think about our work. It's called output literacy, and a piece Nielsen Norman Group published last month makes the case that it isn't a personal habit at all. It's the core skill of our discipline now.
What Is Output Literacy?
Output literacy is the ability to evaluate what AI gives back to you: to notice gaps and misreadings, to recognize a plausible fabrication, to know which details are worth verifying before anything ships. It's the counterpart to prompt fluency: the skill of getting a good answer out, rather than getting a good question in.
The two don't grow together, which is the part I'd missed. NN/g's Maria Rosala, studying how people actually work with these tools, found the combination that should worry anyone running a team: fluent, confident, prolific, writes excellent prompts, accepts the results at face value. She calls it the naive power user.
Which describes, in most organizations, the person everyone points to as the strongest AI adopter.
Enthusiasm, it turns out, is not evidence of judgment.
What This Looks Like in a Studio
Here's where it gets specific to us. Our version of output literacy has a name we've been using for decades: critique.
That's the argument Adam Elman made in NN/g's June piece on design critique in the AI era, and he names the structural reason. Traditional software is deterministic: same input, same output. We write a spec, engineering implements the behaviors we specified, QA validates against it. Generative systems are probabilistic. The same input doesn't guarantee the same output. Which means the spec, as a mechanism for controlling quality, no longer holds.
Elman's example is the plainest possible one. Ask a model about the weather and it might say too much, or too little, or report that rain is unlikely at a 30% chance, which is technically true and also exactly the sort of thing most people would want to know about. The model is making design decisions. Nobody designed them.
So the job changes. Instead of specifying exact behaviors, we define what good looks like — and, just as importantly, what it doesn't. Elman's team runs a judge-evaluate-iterate loop: write explicit criteria for an acceptable output, evaluate real outputs against them, feed the failures back into the implementation, and add new criteria as new failure patterns surface.
The part I'd underline for anyone leading a creative team is his warning about the criteria themselves. Vague ones don't work, because "does this feel too verbose?" just relocates the subjectivity into whoever happens to be reviewing. But arbitrary ones fail differently. Take response time, and set the bar at ten seconds. That's far too slow if someone is asking the system to turn off a light, and far too strict if they've asked something that genuinely requires reasoning. One number, wrong in both directions, and it'll look perfectly reasonable in a document. The criteria have to be objective enough that most reviewers land in the same place, and grounded enough in how people actually use the thing that they mean something. That's a hard authorship problem, and it belongs to us.
The same obligation runs into what we ship. Verification guidance exists in nearly every AI product and users route around it. One of Rosala's participants worked past a full-screen welcome message in Gemini for roughly ten minutes before dismissing it, with no sign she'd read a word. So make verification an action rather than a warning, and flag the fields that tend to be wrong instead of disclaiming everything uniformly, which only teaches people to ignore all of it equally. That's information hierarchy and affordance design. We've been solving this class of problem for thirty years. We're just being asked to solve it for a new kind of output, on a much shorter clock.
Elman lands on a sentence I wish I'd written in April: designers are the arbiters of good, through considered judgment and rigorous critique rather than through taste.
Craft Was Always the Answer. Now It Has a Name.
In April I argued that craft, the accumulated instinct for what's right, the empathy for the person on the other end of the screen, the judgment that gives AI its direction and its purpose, is the thing we bring to these tools that they cannot bring themselves.
I believed that then and I believe it more now. What these two pieces of research do is give me the mechanism.
Because look at what output literacy actually consists of in our work. Noticing that something is subtly off. Knowing which details tend to be wrong. Recognizing fluency that isn't saying anything. Sensing when an answer is too clean for a problem that messy. None of that is a new competency. It's what we've done in front of a wall of printouts for as long as the discipline has existed. The muscle exists. What's changed is the volume it has to handle, and the fact that it now has to be written down.
That last part is the real shift. Critique used to live in a room. It was tacit, conversational, carried in the judgment of whoever was senior enough to say "no, not that." Elman's loop asks us to make it explicit, to convert taste into criteria that survive being applied a thousand times by someone, or something, that wasn't in the room. That's uncomfortable. It's also the only version of craft that scales to the volume of output we're now producing.
It also clarified something I'd underweighted: where in the workflow craft earns its keep. I'd have said upstream framing the problem, directing the tool, setting the brief. That's real. But a great deal of it lives downstream, in the two seconds after the output appears. When the persona reads a half-degree false. When the copy is fluent and hollow. When the flow is technically coherent and would still confuse a real physician on a real Tuesday afternoon. You cannot prompt your way to that reaction. It comes from having built the thing badly once and remembering exactly how that felt.
The naive power user isn't missing a technique. They're missing twenty years.
Which reframes the fear I wrote about in April, and reframes it in our favor. The worry was that AI would make craft less valuable. The opposite happened. When generation costs almost nothing, the judgment applied to it is nearly the entire value of the work. Craft didn't get displaced by this technology. It got concentrated by it.
It Almost Worked Too Well
A recent example of my own.
We built an internal audit tool that evaluates a site across eight dimensions of its experience. We specified it carefully, documented the structure, and ran it. It was the kind of setup that makes you feel good about whatever comes out the other end. And what came out was clean. Well organized, thorough, and confident in a way that made it easy to accept.
Some of it was wrong. It flagged a color pairing as failing WCAG contrast when the ratio was comfortably within spec. It reported an element as unreadable, which turned out to be a screenshot captured before the page had finished loading. Nothing in the output looked wrong. Every finding was formatted identically and stated with identical confidence, the accurate ones and the artifacts alike. The only way to tell them apart was to go through them one at a time and check.
Then I found the more interesting problem. The grading rubric wasn't identifying the wrong things; directionally, it flagged real areas of concern. But the scale didn't match the numerical values being assigned across the eight dimensions. The math didn't reconcile. Anyone who sat down and checked the arithmetic would have found it, and nobody skimming the report ever would have.
So the failure wasn't a wrong answer. It was an answer that couldn't survive being checked, which is a harder thing to notice and a worse thing to hand a client. I rebuilt the scoring into a more granular structure so the numbers meant something at the dimension level.
All of this surfaced in internal testing. The work did inform what we eventually brought to the client, but only after we'd fixed the criteria and corrected the output.
The tool is better now, and so are the criteria behind it. That's the whole arc in miniature: the output was never the deliverable. The judgment applied to it was.
Which raises the question I want to take up next. Everything above is about work that gets reviewed before anyone sees it. But we build for a different situation, a clinician reading something at the point of a decision, with no review step and no time for one. When there's no one downstream to catch it, the standard has to change. That's Part 2.
References
Rosala, Maria. "How AI Literacy Shapes GenAI Use." Nielsen Norman Group, February 6, 2026. https://www.nngroup.com/articles/ai-literacy/
Elman, Adam. "The Core Skill of Design in the AI Era: Critique." Nielsen Norman Group, June 12, 2026. https://www.nngroup.com/articles/ai-era-critique/
Harris, Buddy. "Craft in the Age of Intelligence." Relevate Health Blog, April 2, 2026. https://www.relevatehealth.com/blogs/craft-in-the-age-of-intelligence
