Cognitive Operations Maps
Every business already runs on one. Yours just isn’t written down.
Someone has asked you by now which parts of your business the machines can do. Your board, your partner, your brother-in-law at Thanksgiving.
You probably answered with a list of tools.
You got your employees a ChatGPT subscription. Your engineers are using Claude Code. You have better meeting summaries. Emails and Slack messages are getting longer and better written.
Tools are the wrong unit. Your business doesn’t run on tools. It runs on a few hundred small recurring decisions - classify this, retrieve that, rank these, write the draft, decide whether the draft is good enough to send, decide how much to order, how much to produce, decide whether the customer complaining at 11pm gets a refund.
Every one of those decisions is a seat, and something sits in it.
You have an org chart. It says who reports to whom and what to call them. You have process docs. They say what happens in what order. Neither one tells you where the judgment is. The closest thing on file is a RACI matrix, and RACI records who owns a task - not what would make the answer wrong.
So the question about how you’re using AI at work lands on nothing, and the conversation defaults to vendors and tools.
A cognitive operations map is the missing surface. It’s a registry of every recurring judgment your business makes - one row each - recording what question the seat answers, who or what holds it, what would make its answer wrong, and what checks it.
I have one for Skylark Creations, my portfolio of eighteen AI-native software businesses. On the biggest project, the registry currently holds 19 pipelines comprising 129 distinct steps, held by about 26 model seats.
I have human partners on a few of the businesses and advisors on others. The daily operations are me and the machines, which is exactly why I needed the list: when almost every seat is filled by something I can rent by the token, “who does this” stops being obvious and starts being a design decision.
This article is about the instrument - about what a seat is, the two axes that classify one, and what the map catches that an org chart can’t. The readings come next time.
1. The seat
The unit is not a team, a tool, or a job title. It is a single recurring cognitive operation, stated as the question it answers.
From my own registry: extract the claims from this draft. Decide whether this hook is grounded in the source it cites. Score this candidate against the record of what the editor has accepted. Decide which event to show to this user at this moment. Decide on the narrative structure of this user’s Daily Dhamma.
Each of those is one seat. Each can be held by a model, a person, or twelve lines of code.
Take one of them the whole way down - the row is the thing, and I’d rather show you one than describe the format.
Seat: is this hook grounded in the source it cites? It runs on a content platform I build and operate for a media publication. The question is answered many times per article. It is held by a model, not a person. Its answer is a verdict on another piece of work rather than a piece of work itself. It fails in a specific and nasty way: a sentence that reads beautifully and cites a source that does not say that. And it is validated against a frozen set of past verdicts - a new candidate model gets the seat only if it agrees with rulings I already believe were right.
That last field is the one that makes the cognitive operations map a change-control system instead of just extra documentation. Measurement seats change only on agreement with the frozen set. Writer seats change only on validated results. Nothing changes outside a pre-registered round, and nothing gets tuned to make a score go up.
Skip that field and your map is a diagram that rots.
One more load-bearing property: a seat is a chair, not the person in it. Humans and machines are the same node type. A human seat differs in what it costs (time instead of tokens), what validates it, and maybe in why it has to stay human. It doesn’t differ in kind.
You hear the phrase “humans in the loop” to describe their participation in a system. It’s a column in a table.
2. The first axis: is it a rule or a ruling?
Every seat answers to one of two very different standards. The process of teasing them apart is another reason to build the map.
Here is the test. Run the seat twice on the same input. If the two answers differ and that is a bug, you have a rule. If the two answers differ and both are defensible, you have a ruling.
Rules are counters, thresholds, schema checks, a sort, a query, a regex. They can be wrong. They can’t be unsure. When a rule breaks you fix the rule, and the fix is permanent.
Rulings are the ones with a “well, it depends” inside them. Is this claim supported by its source? Is this good enough to publish? Does this sound like us? How many of these will we sell next month? Is this refund request legitimate? Which of these three ads will convert?
Two competent judges, human or machine, can disagree and neither one has malfunctioned.
The axis matters more than it looks, because three things follow from it and from nothing else.
Validation. A rule is validated by a test. A ruling can only be validated against a record - a body of past verdicts you’re willing to stand behind. You can’t unit-test taste. You can only check whether the seat still agrees with the taste you already had, which is a weaker thing than being right, and it’s the best anybody has.
Cost. Rules are close to free and they stay free. Rulings are metered every single time, forever.
Accountability. A rule can be read. A ruling has to be trusted. This could be the whole reason the human column on this map never empties out.
3. The second axis: who holds it, and why
The second axis is simpler to state and harder to answer: is this seat held by a human or a machine?
Sorting work between people and machines is not a new pastime. Fitts published a “humans are better at / machines are better at” list in 1951, and the field has been correcting it ever since, because it was a static table and the world was not. The version that survives is not a table of capabilities. It is a question you ask seat by seat: why is this one still human?
When I tagged every human seat across the portfolio, only four answers ever came back. There was no force-fitting either - every seat took a tag cleanly. Which is either a very good sign or a slightly worrying one.
Taste. The human is still the better judge, and the machines are calibrated against their verdicts. This is the genuinely contested class - comparative advantage in judgment, possibly temporary (Agrawal, Gans & Goldfarb 2018). Polanyi’s tacit knowledge is the older name for why it resists specification: we know more than we can say, and a seat that runs on what you can’t say is a seat you can’t hand over yet.
Accountability. Someone has to be answerable. The byline, the publish button, the refund, the thing a regulator or a reader can attach blame to. This one does not automate at any capability level, because it is not a difficulty property. The EU AI Act’s Article 14 writes it into law for high-risk systems. Elish (2019) adds the design warning: the seat can decay into a “moral crumple zone,” a human positioned to absorb blame for a machine’s failure. Who signs has to be designed, not defaulted.
Exception. The machine routes out what it can’t handle and the human catches it. Supervisory control has studied this seat for fifty years (Parasuraman, Sheridan & Wickens 2000). It shrinks as capability rises - with a catch I’ll get to.
Training signal. The human’s corrections are the asset. Every accept, reject, and edit becomes the standard the next model has to beat. In my registry that is literal: the frozen verdict sets that gate every judge-model change are made of past human rulings. Most discourse treats this seat as a temporary bridge to full automation. My own operation points the other way. It looks like the core accumulating property of the business - the thing any business would be most reluctant to sell and the smartest thing to build around.
Worth noticing: the tags come out different for different businesses. On the content platform, human seats concentrate in taste and training signal - voice, editorial judgment, an accumulating record of what a writer will actually accept. On the commerce products they concentrate in accountability - payments, app-store compliance, the places where I am the only thing a customer can hold responsible. Same four tags, opposite centers of mass. That is what an instrument is supposed to do. It returns different readings on different objects.
4. The grid
Put the two axes together and you get four quadrants. Every seat you have is in one of them, whether or not you have ever said so out loud.
Three of those quadrants are easy to populate from my own registry. The machine-rule quadrant runs one of my products almost entirely: a meditation app whose checks are all deterministic code, free and permanent. The machine-ruling quadrant holds most of the 129 steps. The human-ruling quadrant holds me - and I ask for help from other machines and people with the hardest judgments.
The human-rule quadrant is nearly empty in my portfolio. Before that sounds like an accomplishment: it’s a consequence of being one person. None of these businesses had a floor of people executing checklists. In companies with employees in that quadrant, this is likely where the “where does AI go” conversation will start, because it’s the one place where the answer isn’t a judgment call.
It’s also the oldest quadrant anybody has records of. In the 1790s Gaspard de Prony produced the French state’s logarithm tables by splitting the work across three tiers of people, the bottom tier staffed partly by out-of-work hairdressers doing rote arithmetic. (The Revolution had guillotined their clientele, and elaborate aristocratic hairstyles turned out not to be a growth industry.) Their job title was “computer.” Babbage wrote it up in 1832 as the division of mental labor. Two hundred years later the job title moved to a machine and the seat did not change at all.
5. What the map catches
The map catches seats that are in the wrong row. Not the wrong column - the wrong row. Businesses misclassify rules and rulings constantly, in both directions, and both errors are expensive in ways that never show up as a line item.
A ruling dressed as a rule. It usually has a name, and the name is the disguise - “the policy,” “our standards,” “the bar.” Nobody ever wrote down the criteria, because there are no criteria. There’s often a person, and that person has the taste. Everyone treats the seat as a rule because rules are what things with names are.
Everybody can tell you who holds it and nobody can tell you what it decides on, and that gap is the signal.
The costs arrive later. That seat can’t be checked. It can’t be delegated. It walks out the door when the person does. And when somebody automates it, they automate the name instead of the judgment, which is how you end up with a machine confidently enforcing a standard nobody wrote. So, if you build only one column of your map, build the one that asks what would make this answer wrong.
A rule dressed as a ruling. Earlier this month one of my machine seats approved a customer survey - the standard “how disappointed would you be if this went away” instrument, minimum of 40 responses, deadline attached. The machine seat assigned to build it did something nobody asked it to do. It counted the population first. 405 users. 296 after stripping bots and test accounts. 90 who actually qualified. At a generous 40% response rate: 36. Short of the 40 the survey needed to say anything at all.
It killed the survey on arithmetic before writing a line of code. No human was in that loop - I found out from the day’s digest - and I don’t think I’d have caught it either. What I’d have gotten was a finished survey and a plausible-looking number, which is the most dangerous artifact any business produces. Nobody argues with a number. The seat that decided “is this survey worth running” had been sitting in the ruling row the whole time. It belonged in the rule row. It was division.
6. What moves, and what doesn’t
Now the part about machine intelligence, which is where most of this discussion starts and where I think it should end instead.
Capability moves seats across the map, right to left: rulings that were human become rulings a machine can hold. That migration is the entire news cycle, it is real, and my own registry is mostly evidence for it.
But the more valuable move runs up, not sideways. Every seat you can promote from ruling to rule comes off the meter permanently and stops being a thing that can be unsure. A counter that runs forever costs less than a good model answering the same question once. And better models make this move easier rather than harder, because a good model is very good at noticing that a question was arithmetic all along. Sideways migration is what the machines do to your map. Upward migration is what you do to your map, and nobody is going to do it for you.
Then there is what does not move. Accountability is not a difficulty property, so no capability level touches it. There is no model good enough to be the thing a customer sues. That seat can be badly designed - Elish’s crumple zone is exactly what happens when you leave a human holding blame for a machine’s decision without the authority or information to have prevented it - but it cannot be vacated.
One of my accountability seats carries a pre-written re-humanization trigger: a dated ruling that says when the product crosses a specific customer threshold, a human goes back into the loop. Klarna replaced hundreds of support agents with an AI assistant, and by 2025 its CEO was saying cost had been too dominant a factor in the decision and the result was lower quality - the company started bringing human agents back. I’d rather write the reversal down in advance than argue afterward about what to call it.
And one warning stapled to all of it, from a paper industrial engineers have been re-proving since 1983: the better your machines get, the less practice your humans get, at exactly the moment the judgments left to them get harder (Bainbridge 1983).
The residual human seats do not get easier as you automate around them. They concentrate. If you keep a taste seat, budget deliberate exercise for it the way you would budget maintenance on any other instrument.
7. Building one
It is less work than it sounds, and the first version should be deliberately crude.
Start from your outputs, not your systems. List the things your business actually emits - the invoice, the published post, the shipped order, the reply to the angry customer. Walk each one backwards to the decisions that produced it. Write each decision as the question it answers. If a row does not read as a question, it is not a seat, it is a tool, and it belongs somewhere else.
Then run the two-runs test on every row and mark it rule or ruling. Then mark who holds it. Then, for every human ruling, write down which of the four reasons applies - taste, accountability, exception, training signal.
Here’s the worrying half. A taxonomy that never fails to fit is either describing something real or is loose enough to absorb anything, and from the inside those two feel identical. Mine has four tags and one operator. So if you build a map and find a human seat that takes none of the four, that is the most useful thing anyone could send me, and I would genuinely like to see it.
The other worry, which I’ll state before somebody else does: a map like this can turn into a thing you maintain instead of a thing you use. I don’t have a clean answer for that. The best rule I’ve got is that a row earns its place by changing something - a seat that moves, a model that gets swapped, a survey that doesn’t get built. Rows that have never changed anything are decoration, and decoration is what documentation becomes right before everybody stops reading it.
Try it? An afternoon gets you a real first draft.
Coda
A rule can be read. A ruling has to be trusted.
That distinction is old, and for most of the history of business it was never urgent. The rulings were held by people you could see. The rules were held by paper you could read. The map was implicit and implicit was fine.
It stopped being fine when judgment became something you can buy by the token. Now the rulings are held by things that answer instantly, sound certain, and are wrong in ways that read beautifully. The only way to know which of your decisions are in that category is to have written the list down.
Write the list down.
Next time: what the seats cost. I pulled the billing records for every seat and found what fraction of my machine spend goes to checking work rather than making it. It was not the fraction I expected, and it was not the same fraction twice.
References
Agrawal, A., Gans, J. & Goldfarb, A. (2018). Prediction Machines: The Simple Economics of Artificial Intelligence. Harvard Business Review Press.
Babbage, C. (1832). On the Economy of Machinery and Manufactures. London.
Bainbridge, L. (1983). “Ironies of Automation.” Automatica 19(6): 775-779.
Elish, M.C. (2019). “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction.” Engaging Science, Technology, and Society 5: 40-60.
European Union (2024). AI Act, Article 14 (human oversight).
Fitts, P.M., ed. (1951). Human Engineering for an Effective Air-Navigation and Traffic-Control System. National Research Council.
Parasuraman, R., Sheridan, T.B. & Wickens, C.D. (2000). “A Model for Types and Levels of Human Interaction with Automation.” IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans 30(3): 286-297.
Polanyi, M. (1966). The Tacit Dimension. Doubleday.






