At a recent product launch for an AI finance platform, the most interesting moment was not the demo. It was the chat.
While the agents ran, the audience typed questions. Does the AI calculate the numbers itself, or does it call predetermined code? Do I need to set up my own Anthropic account? Does the data leave our servers? Claude already does this, so what is the practical difference? Can it cope with messy data and rules that change every month without turning into another spreadsheet to maintain by hand?
Nobody prompted those questions. They are what finance people now ask by reflex, because they have all seen a confident AI answer that turned out to be wrong, and they have all inherited a tool that became one more thing to maintain.
They are the right questions. Here they are, organised, with what a good answer sounds like and what should make you pause. They apply to every category: close automation, FP&A platforms, AI agents, modelling tools, Layerz included.
1. Does the AI do the math, or does it call code?
This is the question to ask first, because the answer shapes every other one.
A language model is very good at deciding what to calculate: which lines, which periods, which method. It is not the right place to perform the calculation, and serious tools do not ask it to. Agents delegate arithmetic to tools routinely and get the right answer; the claim that "AI can't count" is easy to disprove and misses the point. The point is that a number produced inside the model's reasoning cannot be inspected, cannot be rerun identically, and may differ on the next run of the same prompt.
A good answer: "The AI writes the logic, an engine computes it. Here is where you see which is which." The vendor can show you a number and the formula or query that produced it.
Pause if: the answer is vague, changes register halfway through, or amounts to "the model is very accurate now".
2. Where does each number come from?
Revenue lives in the CRM, the billing system, the ledger and probably a spreadsheet. They rarely agree to the cent. Something has to decide which one is the reference for which use.
An agent wired to five systems with no rule about precedence will pick one, and may pick a different one tomorrow. That is not a model failure, it is a missing specification. The best tools make the precedence explicit and keep it outside the agent: a mapping, a configuration, a written convention.
A good answer: "Each metric has a declared source. Here is where that rule is written, and here is who can change it."
Pause if: the answer is "the AI figures out the best source".
3. What happens when an assumption changes?
The most common failure of AI-built models is not a wrong number. It is a correct number written as a value where a relationship should be. It is right today, and it silently stops responding the day an input moves. I wrote about why this failure passes every visual check separately.
You can test this in two minutes during a demo. Ask the vendor to change one driver by 10% and show you the outputs.
A good answer: everything downstream moves, immediately, and the tool can show you the chain from the driver to each output.
Pause if: the change requires a new run of the agent, or some lines do not move and nobody can say why.
4. How is accuracy measured?
"Accurate" is an adjective. Ask for the measurement.
The most convincing test available to a finance tool is a replay: run it on periods that are already closed, where you know what happened, and compare. For a close or reconciliation tool, that means reproducing last quarter's results. For a forecasting or modelling tool, it means rerunning the forecast from an earlier vantage point and showing the gap with actuals. Coverage (was all the needed data there?) and completeness (did every step run?) are useful companions.
A good answer: a described method, a number, and the offer to run it on your own history before you commit.
Pause if: the only evidence is a demo on the vendor's sample data.
5. What does it remember between sessions?
A finance workflow is a month, not a conversation. The tool has to know next month what it learned this month: your chart of accounts, your conventions, the exception you explained in March.
There are two kinds of memory worth distinguishing. Memory of rules: instructions, conventions, context files. Memory of the model: the actual structure of your numbers, the drivers, the links between lines, versioned so you can see what changed. Many tools now do the first well. Fewer do the second, and it is the one that makes a forecast reusable rather than regenerated. More on this in what a financial model must keep between sessions.
A good answer: the vendor shows you where each kind of memory lives, who owns it, and its history.
Pause if: memory means "the chat history".
6. When I correct it, does the correction stick?
Every tool will get something wrong in the first weeks. What matters is what happens after you fix it.
The good pattern: the tool asks when it meets a case it does not know, records your answer, and proposes to turn it into a rule you can read and edit. The correction becomes part of the system rather than something you repeat every month.
A good answer: "Your correction becomes a written rule. Here is where it is."
Pause if: you have to re-explain the same thing each period, or corrections disappear into a place you cannot inspect.
7. Which AI, and whose account?
There are two legitimate models.
AI included. The vendor runs the model, the tokens are in the price. Simple to buy, predictable to budget, nothing to configure.
Bring your own AI. You connect the assistant you already use (Claude, for example) to the tool. You pay your AI subscription separately, and the tool does not resell intelligence.
Neither is wrong. The question underneath is where your context lives. If you already work daily in an assistant, with your projects, your instructions and your workflows, a tool you plug into it inherits all of that. A tool with its own assistant starts from what you teach it. Ask yourself which of the two you want to be building up over the next three years.
A good answer: a clear statement of which model the vendor uses, and what it means for cost and for where your context sits.
Pause if: the AI cost is unclear, or appears later as a usage line you did not expect.
8. Where does the data go?
Ask it early, because your IT team will.
Three sub-questions cover it. Where is the data stored, and can it stay in your own warehouse? Which AI provider processes it, under which terms? Is any of it used to train anything? If you work under client contracts that restrict AI use, the answers decide whether you can use the tool at all. We covered what those clauses actually say and how to assess AI data security for confidential financial data.
A good answer: precise, written, and the same in the contract as in the sales call.
Pause if: "it's secure" is the whole answer.
9. What do I get out?
Most finance work ends in front of someone who did not use the tool: a board, a sponsor, an auditor, a buyer's advisor. What they receive decides whether the work holds.
There are two kinds of output. A report (a PDF, a deck, a dashboard, an Excel file with values) is a conclusion. A workbook with live formulas is an argument: the reader can open it, follow a number to its inputs, change an assumption and see what happens. Both have their place. Only the second survives someone who wants to challenge the numbers.
A good answer: the vendor exports a real file in the demo, and you open it. Click a cell. Is there a formula behind it?
Pause if: export is a paid add-on, or the file is a picture of the result.
10. How long until it works, and who does the work?
Some tools are a service with software inside: a dedicated team builds your workflows over a few weeks, and you review. Others are self-serve: you connect them yourself in an afternoon and build as you go. Again, both are legitimate. A finance team of twenty with a complex ERP may want the first. A fractional CFO with ten clients, or an analyst who lives in Claude, usually wants the second.
Then ask the follow-up that matters more: can I leave? If the vendor disappears or the contract ends, what do you keep? Your workflows, your rules, your models, in a format something else can read?
A good answer: a realistic time to first value, a named owner on each side, and an export of everything you built.
Pause if: leaving means starting again.
A one-page scorecard
| Question | What you want to see |
|---|---|
| 1. Who does the math? | The AI writes the logic, an engine computes, and you can see which |
| 2. Where does each number come from? | A declared source per metric, outside the agent |
| 3. What if an assumption changes? | Everything downstream moves, immediately, traceably |
| 4. How is accuracy measured? | A replay on your own closed periods |
| 5. What does it remember? | Rules and model, both versioned, both inspectable |
| 6. Does a correction stick? | It becomes a written rule you can edit |
| 7. Which AI, whose account? | A clear model, a clear cost, a clear home for your context |
| 8. Where does the data go? | Precise and written, identical in the contract |
| 9. What do I get out? | A file with live formulas you can open and challenge |
| 10. Can I leave? | An export of everything, in an open format |
Run it in the demo, not after. Most of these take less than five minutes to test live.
Where Layerz sits in this
Here are our own answers, so you can hold us to them.
Layerz is a structured spreadsheet built for AI: it holds a financial model (assumptions, formulas, the links between lines, timelines, scenarios) as a structure separate from its data, and Claude drives it from the outside over MCP.
- Who does the math: Claude writes the structure. The Layerz engine calculates, deterministically, and every number traces back to its formula. Change an assumption and the model recalculates as you type, without a new call to the AI.
- Where numbers come from: mapping rules declare which account or bank line feeds which item, and each model carries a
FINANCE.mdwith its conventions (currency, sign, units, structure). - Memory: the model itself is the memory, versioned, with its history and branches.
FINANCE.mdholds the rules. - Which AI: you bring your own. Layerz is one of the rare finance tools you use with the Claude you already work with, with your projects and your context. Layerz does no AI processing on its side and trains nothing on your data.
- What you get out: a clean Excel workbook with live formulas, never paywalled, plus JSON and an open model format.
- Time to first value: self-serve, a few minutes to connect.
And what Layerz is not: it does not run your month-end close, reconcile your bank accounts on autopilot or post entries to your ERP. If that is the job, you want a different kind of tool, and the ten questions above still apply to it.
Further reading: Claude for Financial Modeling: Where the LLM Ends and the Model Layer Starts · How to Stop AI From Hardcoding Values in a Financial Model · The Verification Tax · Persistent Memory for a Financial Model: The Six Things That Must Survive a New Session · Finance Software Went Closed. Developer Tools Went Open. AI Is Ending the Divide.
Layerz keeps a financial model as structure separate from data. Claude writes the formulas, the engine calculates, and every number traces back to its formula. Excel export is clean, standard, and never paywalled. Explore Layerz →