← All articles

icm

The AI Workflow You Can Open Like a Folder: Interpretable Context Methodology, Explained

Most AI workflows are a black box. Interpretable Context Methodology (ICM) runs them as numbered folders of plain files, so you can read and edit every step. Here is how it works, where it fits and what the evidence says.

· 7 min read

icmai-automationworkflowssmall-business

Share

The AI Workflow You Can Open Like a Folder: Interpretable Context Methodology, Explained: Most AI workflows are a black box. Interpretable Context Methodology (ICM) runs them as numbered folders of pl

Most AI workflows work like a black box. You hand over a big request, wait, and get a big answer back. If something is wrong, you find out at the end, and you cannot see where it went wrong.

There is another way to build them. It is called Interpretable Context Methodology, or ICM. The idea is simple enough to explain with a folder on your desktop.

What Interpretable Context Methodology is, in plain words

ICM comes from a paper by Jake Van Clief and David McDermott, published on arXiv in March 2026. The method is open source under the MIT license, and the code is on GitHub.

The authors' central idea is simple. Suppose the instructions for each step already live as files in an organized folder. Then you do not need a complicated framework to run them. One AI agent can read the right files at the right moment.

In practice, ICM comes down to four ideas:

  1. One stage, one job. Each step has its own numbered folder. A step that gathers data does not also write the final report.
  2. Plain text instructions. Each stage has a short file that says what to read, what to do and what to produce. Anyone with a text editor can read it.
  3. Every output is a file you can edit. When a stage finishes, it saves its work in a folder. You can open it, fix it and save it. The next stage starts from your edited version.
  4. Scripts do the mechanical work. Fetching data, moving files and sending emails do not need AI, so ordinary scripts handle them.

What it looks like

Here is an imagined weekly client report. This is an illustration, not a real client.

weekly-report/
  CONTEXT.md             where things are
  01_gather/
    CONTEXT.md           what this step reads, does and writes
    output/              the gathered numbers, as a file
  02_draft/
    CONTEXT.md
    output/              the draft report, as a file
  03_polish/
    CONTEXT.md
    output/              the final report, ready to send
  _config/
    voice.md             how our reports should sound

The numbers in the folder names set the order. The output folders are the handoff points. The _config folder holds rules that stay the same every week, such as tone and format.

Between any two steps, a person can stop and look:

StageWhat it doesWhere you step in
01 GatherCollects this week's numbersCheck that nothing is missing
02 DraftWrites the first version of the reportFix the angle before it is polished
03 PolishApplies your voice and formatApprove it before it goes out

Why this matters for a business

The paper compares ICM with software frameworks that run AI agents through code. For workflows that run step by step with a person reviewing each one, the day-to-day differences are practical:

If you want to...In a code frameworkIn ICM
Change a promptEdit configuration in codeEdit a text file
Add or remove a stepWrite new codeAdd or delete a folder
See what happenedBuild logging or a dashboardOpen the folder and read
Hand it to someone elseDocument the setupCopy the folder
Who can make changesA developerAnyone with a text editor

For a business owner, the biggest benefit is that nothing is hidden. Every intermediate result is a readable file, so there is nothing to explain after the fact. You catch problems at the cheapest moment: after the first step, not after the last.

Less clutter for the AI, too

Each stage only loads the files it needs. The authors point to research showing that AI models do worse when the information that matters is buried in a long pile of text. Loading only what a step needs avoids the pile in the first place.

In one example workspace, the authors report that each stage gets roughly 2,000 to 8,000 tokens of focused context. They estimate that a single all-in-one prompt for the same job would run to 30,000 to 50,000. A token is a small piece of text, roughly three-quarters of a word.

Treat those numbers as an illustration. The authors say they have not run a controlled comparison between the two approaches.

What the evidence says

The paper is open about its limits, and so should we be.

  • The evidence is informal. The observations come from conversations with an invite-only community of 52 practitioners, not from a formal study.
  • Edits cluster at the start and the end. Of 33 practitioners, 30 said they edit the most at the first stage, where they set direction. They also edit heavily at the last stage, to check the output matches earlier decisions. They edit the least in the middle. This is self-reported and has not been verified.
  • Non-developers got it working. Three practitioners with no coding experience used the workspace builder to create and run workspaces that produced animated videos. The authors call this a single data point from a small group.
  • One model family. All the workspaces in the paper were built and run with Claude Code and Claude Opus 4.6. Results with other models are untested.

In short: the idea is sound and well argued, and early reports are encouraging. It is not proven. Test it on your own work before you rely on it.

Where ICM does not fit

The authors say plainly that ICM is not for everything. It is a poor fit when:

  • Agents must talk to each other in real time. File handoffs are too slow.
  • Many people use the same workflow at once. It was designed to run locally, for one person or team at a time.
  • The AI must choose its own path mid-way. A person can decide between stages, but automatic branching pushes ICM toward becoming a framework.

It works best for work that is step by step, worth reviewing and repeated, such as weekly reports, client onboarding packets or content production.

How to try it this week

  1. Pick one repeated task. Choose something you do weekly or daily, with a clear start and finish.
  2. Split it into three steps. For example: gather, draft, polish.
  3. Write each step's instructions in a plain text file. What does it read? What should it produce?
  4. Decide where you review. Stop after each step at first. Skip checks only where you have come to trust the result.
  5. Keep the rules in one place. Put tone, format and style in a shared file, so they stay the same every run.

Where Cognitivv AI fits

We use ICM as one approach for sequential work that people should review at each step. [Describe one real workflow where we use staged review, or delete this paragraph.]

In the Cognitivv workspace, that fits alongside a private workspace per client, approval on anything customer-facing and a monthly report of what ran. If you want help finding the first workflow worth structuring this way, take a look at our workflow audit.

Source: Jake Van Clief and David McDermott, "Interpretable Context Methodology: Folder Structure as Agentic Architecture," arXiv:2603.16021v2, March 2026.

Share

Comments

Only the site owner can see it. It is never shown publicly.

0/2000 characters. Comments are plain text and are reviewed before they appear.

Keep reading

Want help putting this into practice?

Join the waitlist and we will reach out to map your first workflow.

hello@cognitivvai.com