icm
The AI Workflow You Can Open Like a Folder: Interpretable Context Methodology, Explained
Most AI workflows are a black box. Interpretable Context Methodology (ICM) runs them as numbered folders of plain files, so you can read and edit every step. Here is how it works, where it fits and what the evidence says.
· 7 min read
icmai-automationworkflowssmall-businessMost AI workflows work like a black box. You hand over a big request, wait, and get a big answer back. If something is wrong, you find out at the end, and you cannot see where it went wrong.
There is another way to build them. It is called Interpretable Context Methodology, or ICM. The idea is simple enough to explain with a folder on your desktop.
What Interpretable Context Methodology is, in plain words
ICM comes from a paper by Jake Van Clief and David McDermott, published on arXiv in March 2026. The method is open source under the MIT license, and the code is on GitHub.
The authors' central idea is simple. Suppose the instructions for each step already live as files in an organized folder. Then you do not need a complicated framework to run them. One AI agent can read the right files at the right moment.
In practice, ICM comes down to four ideas:
- One stage, one job. Each step has its own numbered folder. A step that gathers data does not also write the final report.
- Plain text instructions. Each stage has a short file that says what to read, what to do and what to produce. Anyone with a text editor can read it.
- Every output is a file you can edit. When a stage finishes, it saves its work in a folder. You can open it, fix it and save it. The next stage starts from your edited version.
- Scripts do the mechanical work. Fetching data, moving files and sending emails do not need AI, so ordinary scripts handle them.
What it looks like
Here is an imagined weekly client report. This is an illustration, not a real client.
weekly-report/
CONTEXT.md where things are
01_gather/
CONTEXT.md what this step reads, does and writes
output/ the gathered numbers, as a file
02_draft/
CONTEXT.md
output/ the draft report, as a file
03_polish/
CONTEXT.md
output/ the final report, ready to send
_config/
voice.md how our reports should sound
The numbers in the folder names set the order. The output folders are the handoff points. The _config folder holds rules that stay the same every week, such as tone and format.
Between any two steps, a person can stop and look:
| Stage | What it does | Where you step in |
|---|---|---|
| 01 Gather | Collects this week's numbers | Check that nothing is missing |
| 02 Draft | Writes the first version of the report | Fix the angle before it is polished |
| 03 Polish | Applies your voice and format | Approve it before it goes out |
Why this matters for a business
The paper compares ICM with software frameworks that run AI agents through code. For workflows that run step by step with a person reviewing each one, the day-to-day differences are practical:
| If you want to... | In a code framework | In ICM |
|---|---|---|
| Change a prompt | Edit configuration in code | Edit a text file |
| Add or remove a step | Write new code | Add or delete a folder |
| See what happened | Build logging or a dashboard | Open the folder and read |
| Hand it to someone else | Document the setup | Copy the folder |
| Who can make changes | A developer | Anyone with a text editor |
For a business owner, the biggest benefit is that nothing is hidden. Every intermediate result is a readable file, so there is nothing to explain after the fact. You catch problems at the cheapest moment: after the first step, not after the last.
Less clutter for the AI, too
Each stage only loads the files it needs. The authors point to research showing that AI models do worse when the information that matters is buried in a long pile of text. Loading only what a step needs avoids the pile in the first place.
In one example workspace, the authors report that each stage gets roughly 2,000 to 8,000 tokens of focused context. They estimate that a single all-in-one prompt for the same job would run to 30,000 to 50,000. A token is a small piece of text, roughly three-quarters of a word.
Treat those numbers as an illustration. The authors say they have not run a controlled comparison between the two approaches.
What the evidence says
The paper is open about its limits, and so should we be.
- The evidence is informal. The observations come from conversations with an invite-only community of 52 practitioners, not from a formal study.
- Edits cluster at the start and the end. Of 33 practitioners, 30 said they edit the most at the first stage, where they set direction. They also edit heavily at the last stage, to check the output matches earlier decisions. They edit the least in the middle. This is self-reported and has not been verified.
- Non-developers got it working. Three practitioners with no coding experience used the workspace builder to create and run workspaces that produced animated videos. The authors call this a single data point from a small group.
- One model family. All the workspaces in the paper were built and run with Claude Code and Claude Opus 4.6. Results with other models are untested.
In short: the idea is sound and well argued, and early reports are encouraging. It is not proven. Test it on your own work before you rely on it.
Where ICM does not fit
The authors say plainly that ICM is not for everything. It is a poor fit when:
- Agents must talk to each other in real time. File handoffs are too slow.
- Many people use the same workflow at once. It was designed to run locally, for one person or team at a time.
- The AI must choose its own path mid-way. A person can decide between stages, but automatic branching pushes ICM toward becoming a framework.
It works best for work that is step by step, worth reviewing and repeated, such as weekly reports, client onboarding packets or content production.
How to try it this week
- Pick one repeated task. Choose something you do weekly or daily, with a clear start and finish.
- Split it into three steps. For example: gather, draft, polish.
- Write each step's instructions in a plain text file. What does it read? What should it produce?
- Decide where you review. Stop after each step at first. Skip checks only where you have come to trust the result.
- Keep the rules in one place. Put tone, format and style in a shared file, so they stay the same every run.
Where Cognitivv AI fits
We use ICM as one approach for sequential work that people should review at each step. [Describe one real workflow where we use staged review, or delete this paragraph.]
In the Cognitivv workspace, that fits alongside a private workspace per client, approval on anything customer-facing and a monthly report of what ran. If you want help finding the first workflow worth structuring this way, take a look at our workflow audit.
Source: Jake Van Clief and David McDermott, "Interpretable Context Methodology: Folder Structure as Agentic Architecture," arXiv:2603.16021v2, March 2026.
Comments