In The CIO’s third act, Rahul Kayala explains why enterprise AI programs rarely stall on model capabilities. They stall because the AI has no real understanding of the systems underneath it. The knowledge of how work moves across a company's systems used to sit with IT. Over twenty years, it scattered across SaaS apps, partners, and people's heads.
At Echelon, we’re building two things that work together. The first is an agentic harness for enterprise applications, so what an IT team can take on is no longer capped by the hours its people have. The second is the IT process graph, which gives the harness the context it needs to do the work well. It's a map of how the business runs across its systems, built from the work itself and kept current as the systems change.
This post is about the IT process graph, and how building the harness taught us we needed one.
Every task started cold.
Picture an IT team that owns the systems below and the processes that run through them. Nothing is documented. No one person knows a whole process, or even every system it touches.
So every bug or troubleshooting request on a system of record becomes a ticket, and every ticket starts with an investigation: query the systems one by one until the cause turns up. Then the fix. Then the next ticket, and the same investigation again.
When we first built our harness, it took on that investigation. It can analyze, design, configure, troubleshoot, and document these systems. It found the cause and applied the fix. But the next ticket started from nothing, just like the last one. If it had kept what it learned from that ticket, the next one would have started at the group-removal flow instead of at the first system. What it learned is the knowledge Rahul says never compounds.
Everything is a data point to learn from
The fix seemed obvious: let every task write what it learned to a shared memory that outlives the task.
Memory is a collection of pages, like the one shown above. They capture the “decision traces” that used to live in people’s heads.
- Each “page” is one short document, about one thing: a system, a process, a team’s convention, or a piece of work in flight.
- A page is made up of “facts”, something that a task showed about the memory, such as “a contractor’s access ends on their end date”.
- Facts are attributable to their “origin” and have a label on “how the fact was determined”. So a fact said by a person is much stronger than something found on the web, or things the agent concluded.
- “Open Questions” are something that a page doesn’t know yet - something that a future task could answer.
- Finally, “State” captures how settled the content of the page is - transitioning from first touch to complete.
We could have started with a more structured organization: a page for every system, a page for every process, fields for owners and approvals, etc. But we chose not to. Every organization works differently, and a fixed schema chosen up front would describe how we think companies run, not how this one does. So we designed only what a page holds, and let the work decide which pages exist.
The memory we curate comes from two places.
The first is the everyday work people run through Echelon: troubleshooting a bug, answering a question, or changing the systems themselves. Each task has to understand the process before it touches it, and when it finishes, it writes down what it learned.
The second is configuration present in the system. Configurations reflect the reality of the business process, not what the old design document claims. History shows who changed it and why: a rule added by a contractor who has since left, custom code rushed in for a go-live, the same fix rebuilt by one integrator after another because none of them knew the last had been there.
What learning looks like
A task writes to memory as it works. Here’s the finance bug again, showcasing how the memory gets updated based on what the harness found.
Three things can happen to what a task finds. If memory already has it, the task adds itself as another source, and the fact gains credibility. If it's new and would change what a future task does, it's added. If it only matters today, like one person's termination date, it's dropped. The source line on every fact is recorded by code, straight from the transcript, never written by the model.
The shape that memory took
We modeled the shelves after what a forward-deployed engineer’s notebook would look like after a few months in an organization.
The product already kept memory in a few shapes (vocabulary, how things are set up, workflows, work in flight), so those became the shelves. Beyond that, we designed nothing. Which pages existed, and how they grouped, was left to the harness.
A lot has been written about agents creating memory for themselves. We think recall is equally important - if not more so. Memory a task can't find is memory it doesn't have. So the harness writes recall notes, each describing how a kind of task tends to open, what to settle first, and which few pages to read. A new task starts with the slice it needs, not the whole memory.
Just memory isn’t enough
Memory worked. Tasks found what earlier tasks had learned, and repeated questions got shorter answers. But it continued to grow.
When the removal job moved from 02:00 to 03:00, the new line landed beside the old one, and nothing retired it. The same fact about Workday lived on several pages, each copy updated on its own. Open questions piled up, and no page ever reached "complete."
We also found the agents gaming our own check. Every fact had to cite the line in the task it came from, so the model found a line that passed and cited it again and again. On one page, most facts pointed back to the same line, which couldn't explain them all. Every citation checked out, and most of them meant nothing. We stopped treating a passing citation as proof, and started judging pages by whether a new engineer would use them.
The deeper problem was the container. In text, a source, a correction, and an open question are all just words. Nothing marks two lines as the same fact, or says that one replaces another. Tidying up meant a model rewriting whole pages, and the cost of that grew with every task.
Let there be process graphs!
What if a fact weren't a sentence on a page, but a record? There would be a record for Workday, one for Okta, one for the group-removal flow.
Each fact about them keeps its sources and how it's known, exactly as before. The difference is that the structure the text was hiding is now something the harness can work with.
Three things change.
- A fact is one record. When four tasks say where end dates come from, the record gains four sources, instead of the page gaining four lines.
- New values replace old ones. 03:00 replaces 02:00. The old value stays in the record's history, with the task that said it, but nothing reads it as current.
- Records link. Workday feeds Okta, Okta feeds ServiceNow, and ServiceNow runs group removal. Following the process, those links are the IT process graph for the “Leaver” stage of the Hire to Retire workflow.
The links also show the gaps: a missing link or a flow still in draft.
Facts still originate from tasks, and every record still says where it came from. What changes is who keeps them tidy. Finding a fact across tasks, retiring a stale value, and closing a question become deterministic graph updates, instead of a model rewriting pages.
Every task starts warm
We started with a promise: what an IT team can take on should no longer be capped by the hours its people have. The harness alone got us partway there. It could do the work, but every task paid for its own discovery.
With the process graph, a task starts where the last one stopped. That makes it faster, but more importantly, it shows what comes after the fix. IT used to know that. The graph brings it back: the systems that feed each step, the ones that depend on it, and the people whose work changes with it.
Every task leaves the graph more complete, so the next one starts further ahead. An IT team no longer has to choose which few changes it can afford to carry end to end.
The process, as it runs today.
The graph the harness works from is the same one the people responsible for a process need. Documentation shows how someone designed the process at some point. The graph shows how it runs now.
Gaps in the graph are where the process actually breaks, found by real work. They are a natural place to start a redesign.
Every transformation starts with understanding how things run today. That step is slow, expensive, and quickly outdated. The process graph makes this easy by staying current - as everyday work keeps it updated.
It also removes the reason to think small. Most transformation plans cover the top few systems, not because the rest don't matter but because nobody can afford to understand them all. With the graph, the whole estate is in view, so a plan can cover every application a process touches, not just the ones someone had time to map.
The graph's projection isn’t fixed either. A stakeholder could bring a planned change to the graph to see every step it would touch, across every process, and which known gaps would move with it.
Book a demo to see the full cycle in practice.


