Hey there,

In the early days of desktop publishing, organizations faced what seemed like a distribution problem. A few people in every office figured out PageMaker or QuarkXPress, and they produced documents that looked dramatically better than what everyone else was turning out. The obvious response was training. Teach more people the software. The less obvious discovery, which took years to surface, was that the people producing good documents weren't just using the tools better. They had visual judgment. They understood typography and white space and hierarchy in ways they'd absorbed over years and couldn't fully articulate. You could teach someone the mechanics of leading and kerning in an afternoon, but you couldn't teach them taste or instinct - the ability to look at a finished page and feel that something was off.

Organizations spent a decade learning this lesson, and most of them learned it in a roundabout way: by distributing templates that produced technically correct documents that looked cheap. The templates solved the consistency problem (same fonts, same margins, same logo placement) while completely failing to solve the judgment problem (does this communicate well?). The output looked professional from a distance. Up close, however, it didn't work.

We are about to learn this lesson again with AI, and we are going to learn it the expensive way again, because the failure mode is the same: the output looks complete.

I work with an architecture firm that's been integrating AI into its practice for several months. Last month, a partner told me about an RFP response they'd submitted that day. She and a senior practitioner had collaborated on the proposal, and the process had surfaced something she wanted to talk through.

The practitioner had done something smart. He'd loaded the client's brief into Claude alongside his draft and asked it to identify gaps: places where the response didn't address something the client had asked for. Claude found several. So, he went back and filled each one.

When the partner got his revised draft, the content was all there, but she felt that the writing was flat. She told me it read like a checklist. "I answered this, I answered this, I answered this. It was kind of boring." Every gap had been addressed, but none of it built toward a case for why this firm should win the project.

She took the draft, saw all the content they were trying to include, and rewrote it with a voice and a narrative. Then she ran her version through Claude for suggestions on clarity and structure, which she found immediately helpful for reorganizing the flow. But some of the content changes were, as she put it, "really off the mark." Suggestions that sounded plausible but misrepresented what the firm was actually trying to say. She caught those because she knew the intent behind each paragraph. Someone reviewing the same suggestions without that context would have accepted them.

"It could have been bad had I not seen it," she told me. "It would have been not so great, and I'm not sure that would have been made evident." The practitioner wouldn't have flagged the draft as a problem, because by his read, it wasn't a problem. The proposal addressed every requirement. It would have gone out complete, competent, and unpersuasive, and nobody would have known why it lost.

Two professionals used the same tool on the same document. Both of them used it correctly. And yet the results were not interchangeable.

The instinct in most organizations right now is to find the people who have figured out how to use AI well and extract their methods. Document the prompts, capture the workflows, and build a shared playbook. This is the instinct that produced those desktop publishing templates twenty-five years ago, and it comes from the same reasonable place: someone has solved the problem, so let's scale the solution.

The trouble is that this treats all work as if it were the same kind of work, and it isn't.

The economist Friedrich Hayek wrote about the problem of knowledge in organizations and economies. Some knowledge can be made explicit and widely shared: specifications, formal procedures, aggregate data. But other knowledge is local, situational, and bound up in what he called "the particular circumstances of time and place." Hayek was talking about economic planning, but the distinction maps remarkably well onto what happens when organizations try to scale individual AI practices.

At the architecture firm, the practitioner's gap analysis is centralizable knowledge. Load the brief, load the draft, ask the tool what's missing. You could write that into a standard procedure and hand it to every person on the team. It would work because it's a consistency operation. The quality of the output doesn't depend heavily on who runs it. This is the kind of step where organizational infrastructure genuinely raises the floor. Build it, share it, require it.

The partner's restructuring, however, is local knowledge. She looked at a complete document and recognized that completeness is not the same thing as persuasion. That recognition cannot be proceduralized because it depends on standards that are acquired through years of doing the work, standards the practitioner doesn't have yet. He didn't submit a flat draft knowing it was flat. He submitted it believing it was done. The gap between the practitioner and the partner is not a gap in tool usage. It is a gap in professional judgment about what "done" means, and a template doesn't solve that.

This distinction has organizational consequences that most AI strategies are not designed to handle.

When you systematize a consistency step, you get the expected benefit. Quality becomes more uniform, the floor rises, the worst outputs get better without the best outputs getting worse. This is real value, and organizations should pursue it aggressively.

When you systematize a judgment step, something different happens. The output still looks complete. In many cases it looks more polished than before, because the tool is good at surface-level coherence. But the quality drops in ways that are invisible to anyone who lacks the judgment the infrastructure was supposed to replace. A proposal addresses every requirement and wins no contracts. A strategy document hits every section heading and misses the actual strategic insight. A client communication sounds professional and says the wrong thing with great confidence.

The partner on that call put it simply: "You sort of want to believe that it's trustworthy. It looks really good. And then you read it and you're like, wait a second, what?"

"When you systematize a judgment step, the output still looks complete. The quality drops in ways that are invisible to anyone who lacks the judgment the infrastructure was supposed to replace."

The desktop publishing parallel is exact. The templates produced documents that looked fine to almost everyone. They were consistently formatted, professionally structured, and perfectly adequate. The only people who could see they were cheap were the people with enough design judgment to recognize what was missing. That's why the templates survived for a decade. The AI version works the same way. A technically complete proposal that lacks persuasive force looks, on its surface, like a good proposal. The failure is legible only to people with the experience to read for what's missing, which is exactly the population you cannot replace with infrastructure.

Most organizations are bad at looking at each step in a workflow and making a diagnostic call about what kind of step it is. Is this a consistency problem, where the value is in uniformity? Or is this a judgment problem, where the quality of the output depends on whoever is doing it? Without that diagnosis, they tend to treat every step the same way, either systematizing everything or systematizing nothing.

Consistency steps should get infrastructure; judgment steps should get protection from automation. Both involve AI. The work is knowing which is which. If you fail to systematize a consistency step, the cost is visible: inconsistent output, obvious gaps, missed requirements. Conversely, if you systematize a judgment step, the cost is invisible: output that looks finished and performs worse, with no obvious signal telling you why.

The architecture firm figured this out for one document, through the lived experience of two people who happened to be working on the same draft. Most organizations won't be so lucky. They will build their AI infrastructure, distribute it, and watch their output become more consistent and less good, and they will not connect those two observations because the first one looks like progress.

Break a Pencil,

P.S. If your team has built shared AI tools around a step that turned out to need a person, or left something to individual variation that should have been standardized long ago, I'd like to hear about it. Reply and tell me.

Reply

Avatar

or to participate