Spreadsheet Audit for AI Agents
What part of the model did the AI agent change and where is this data from?
Every finance org has an answer for that question when a human built the model. Almost none have one when an AI agent did - or it’s hard to dig that information from each session.
A board deck goes out with a revenue forecast that’s off by half a million dollars. Someone asks the obvious question: who touched this number, and why did they change it?
Now replace that analyst with Copilot, Claude, or a purpose-built agent like Endex generating the model directly inside the workbook. Ask the same question. For most finance teams today, there’s no answer. Not because nobody cares - because the tooling to give one barely exists yet.
We went looking for independent audit tools. There aren’t many.
We spent this past week mapping every tool that claims to give you visibility into spreadsheet changes — native platform features, third-party diffing tools, and the audit claims baked into the AI agents themselves. The finding: outside of Microsoft and Google’s own ecosystems, there’s almost nothing built specifically to independently audit what an AI agent did to your numbers.
Most of what exists falls into one of two buckets. Either it’s a general-purpose file history feature that was never designed with an autonomous agent in mind - it tells you a file changed, not why, or on whose authority. Or it’s the agent auditing itself - the same model that wrote the formula also generating the explanation for why the formula is correct.
Would you accept a preparer signing off on their own work with no second reviewer? That’s the control you’re accepting every time an agent’s self-reported audit trail is the only record you have.
Every finance function is built around segregation of duties for exactly this reason — the person who prepares isn’t the person who reviews. That principle doesn’t disappear because the preparer is now a model instead of an analyst. If anything, it matters more, because a model can generate a plausible-looking rationale for a wrong number just as fluently as it generates one for a right number.
How the tools actually stack up
If Claude, Endex, or Copilot are writing to your workbook right now, here’s what actually catches those edits - ranked by how well each one fits an AI-agent workflow, not a human one.
Native tools were built for a world where people made the edits
Microsoft 365 and Google Sheets version history will restore a file. Excel’s Show Changes will tell you a cell moved. Purview will tell you who had access to open the file in the first place. None of that is wrong - it’s just answering a question from a decade ago: did someone unauthorized get into the file?
The question finance teams need answered now is different: did an authorized agent make an unauthorized change to the logic? That’s a much harder problem, and it’s the one native tooling wasn’t built to solve.
It’s the same thing as your manager walking by your desk
Before any of this was automated, review looked like this: you’re deep in a model, and your manager stops by and asks you to walk them through it. Not because they suspect something is wrong — because that’s how the control works. Someone independent checks the reasoning while it’s still fresh, before the number goes anywhere important.
That review doesn’t go away just because the person building the model is now an agent. It just has to happen differently, because there’s no desk to walk by and no analyst to ask. You need an independence layer that shows what the agent did - in the moment it did it, and again months later when internal audit or a regulator asks for the trail.
That’s the layer Rockhopper is built to be. Not a replacement for the agent doing the work - the independent reviewer standing next to it, in real time, with a record that holds up when someone asks the question your model risk framework already requires you to be able to answer.