Your AI made a decision, and Canadian regulators want to know how
How does an AI system arrive at the decisions it makes, and can you prove it in front of a regulator? For Canadian banks and insurers, that question now comes with a deadline.
The question has been around for years, but Canadian regulators are now asking for it in writing. Enterprise AI has moved from pilot to production quickly, and the rules are moving with it Microsoft says the number of active agents running in Microsoft 365 alone has grown fifteenfold in the past year.
The companies getting real value from AI are the ones with governance in place, according to Deloitte’s 2026 State of AI in the Enterprise report. Only one in five companies has a mature model for governing autonomous agents, and the ones that do are also the most likely to roll agents back after deployment.
Joseph Geraci has spent nearly a decade building AI in an industry where regulators have been demanding proof of how models reached their conclusions for years.
He’s the founder and chief scientific and technical officer of NetraMark, a publicly listed Canadian company that analyzes clinical trial data to identify patient subgroups. The pharmaceutical industry has long had to defend how it reached trial conclusions, even when no AI was involved.
Speaking remotely to a room of Canadian technology leaders at CIOCAN’s Peer Forum earlier this year, he talked about accountability.
“We need to move away from logging what an agent did, to auditing exactly why it made certain decisions,” said Geraci.
Canadian regulators are moving in the same direction.
Meet OSFI, and your new homework
If you work outside of the financial services industry, this might be the first time you’re hearing about Guideline E-23.
Here’s the short version.
OSFI, the Office of the Superintendent of Financial Institutions, is the federal regulator that oversees Canadian banks, insurance companies, trust companies, and foreign bank branches.
Last September, it finalized an update to its Guideline E-23 on Model Risk Management, the rulebook for how institutions manage the risks of the models they use. The 2017 version applied only to deposit-taking institutions. The rewrite expands E-23 across federally regulated financial institutions and explicitly addresses AI and machine learning alongside other models.
The updated version takes effect May 1, 2027, giving federally regulated institutions about nine and a half months of usable runway, minus whatever quarter-end and year-end blackouts get in the way.
E-23 requires banks and insurers to know every model they’re using and where it came from, including anything from a vendor. Not every model gets the same treatment. The ones that could affect customer decisions, financial results, or regulatory compliance get the full inventory-and-audit process.
Each model gets a risk rating based on what it does, how autonomously it operates, and how much damage a bad output could do. The higher the rating, the more the institution has to prove it understands how the model works and can monitor it for problems.
Fail the audit and OSFI can require additional capital or a detailed fix-it plan on its timeline.
If you’re a bank or insurance CIO reading this, the deadline is already on your calendar. If you’re not, you’ve got a front-row seat.
Quebec’s Autorité des marchés financiers has finalized a parallel AI guideline for financial institutions under its supervision, which also takes effect May 1, 2027. A 2026 Chambers guide says public sector buyers are increasingly putting AI governance terms into contracts, including documentation standards, audit rights, and incident reporting.
Logs tell you an agent made a decision. Audits tell you why.
A regulator looking at a bank’s mortgage model in 2027 will want more than a timestamp and an output. The institution will need to show what the model was built to do, what data it used, how its output can be explained, and who is accountable for using it.
Pharma has been living with that expectation for years.
Start with what pharma had to prove
Geraci wasn’t talking about E-23 at CIOCAN. Instead, he was talking about what autonomous AI systems need in order to work. They need local context, traceable reasoning, and an architecture that can be audited.
Pharma is where Geraci has been working on the problem. Drug developers have long had to defend how they reached trial conclusions.
A drug may work for one group of patients and do very little for everyone else. Blend those results together and the responders can disappear into the average. The trial can read as a failure, the drug gets shelved, and patients who might have benefited never get the option.
When you have a high-dimensional data set, Geraci said, everything starts to look like the average.
According to him, LLMs have the same weakness. They’re good at the average and can miss the weird, specific details that really matter to a company.
Geraci’s December 2025 paper in npj Digital Medicine works through a case study using data from a National Institute of Mental Health Phase II ketamine trial in treatment-resistant depression. The authors report that NetraAI found groups of patients who responded to the drug that standard analysis had missed, and could show its work using the underlying clinical and imaging data.
The study was retrospective and small, and the authors say independent, prospective validation is still needed. NetraMark funded much of the work, and Geraci and several co-authors are employees or shareholders.
Enterprise AI systems generally aren’t built to produce that kind of traceable reasoning. They’re trained on generic patterns, then pointed at a business whose specific structure they don’t understand.
They work well enough in a demo and start hallucinating in production.
“A hallucination, as of now, is embarrassing,” said Geraci. “But what these autonomous agents are producing, these hallucinations can propagate into catastrophic errors.”
A hallucination in a chatbot draft is a bad email, while the same failure in a system that approves refunds, adjusts pricing, or flags fraud can become a model-risk incident the institution has to detect, document, and address.
Knowing how to get a good answer from a chatbot is not the same as building a system you can defend when a regulator asks how it works.
“We don’t need prompt engineers. We need intelligence architects,” he said. “People who can design the integrated machine, not just those who can talk to it.”
The questions worth asking your team
Three questions are worth putting on the table. All three apply whether the regulator is already writing your rules or hasn’t gotten to you yet.
First, what does your model inventory capture?
Any technology leader has a reason to know what AI is running inside their environment. E-23 is the reason that question now has a deadline for banks and insurers. The regulation asks institutions to keep a record of each model they run, including AI.
But an inventory that just lists models isn’t going to survive an audit, or a serious conversation with the CFO. The practical version has to show what each agent is connected to, what data it can reach, and what it’s authorized to do without checking in with a human.
If your inventory lists the model and stops there, the audit is going to find what’s missing for you, your CFO, or your board.
AI security firms doing enterprise assessments have been finding between 350 and 430 AI services running inside a single organization, most of them never formally authorized.
One CIO with a clipboard isn’t going to cover it.
The second question is can they produce audit trails of the decisions their model makes, or only logs of what it does?
Vendor documentation on how models arrive at their decisions is thin. The vendor probably hasn’t offered to explain because they haven’t been asked, and E-23 is about to change that for anyone buying vendor AI in Canadian financial services. But the question matters wherever a vendor’s system is making decisions on your behalf.
Snowflake’s chief data and analytics officer Anahita Tafvizi told Digital Journal earlier this year that the honest version of the problem starts inside the organization, before the vendor even shows up. Many organizations haven’t documented their own data well enough to train a new hire on it.
“If you don’t have it documented to train a human, then how do you train an AI agent?” she asked.
If procurement teams aren’t asking a vendor version of that question, about the data behind their models, they should start.
Finally, how independently is each system operating, and can you rate that in a way an auditor can follow?
An agent that drafts emails for a human to approve is not the same risk as an agent that approves refunds on its own. E-23 requires banks and insurers to rate that difference. Anyone else running autonomous AI should be rating it too, whether a regulator has asked yet or not.
Geraci made the case with a hiring analogy. When you bring a new employee onto your team, he said, they have to get accustomed to the culture, learn new skills, and pick up a lot of information about how the business works.
Agents are no different. A model trained on the internet doesn’t know how your company approves an expense, what your customers care about, or which decisions your legal team wants to see before they go out the door.
Giving an agent autonomy without any of that is guesswork with a spreadsheet on top.
Underneath all three is the same question. What is your AI doing inside your organization, and can you show your work?
Turns out Grade 8 math class was training for something after all.
Final shots
- The distance between logging what an agent did and auditing why it decided is technical. It won’t close because someone wrote a better AI policy.
- E-23 leaves institutions accountable for vendor models too, which makes the vendor contract part of the model-risk program.
- Explainable AI has been a research topic for years. On May 1, 2027, much of Canada’s financial sector has to be ready to explain the models it put into production.
Digital Journal is the national media partner for the CIO Association of Canada.
Your AI made a decision, and Canadian regulators want to know how
#decision #Canadian #regulators