AI Governance: Who's Accountable When the Machine Decides?
E109

AI Governance: Who's Accountable When the Machine Decides?

Ryan Koch (00:22)
Hey folks, today I want to talk about governance as it relates to artificial intelligence. We've had a couple episodes recently that have touched on topics around this, including conversations about spec ops and AI readiness. Covering a bit about governance seems appropriate after all of that, as it's something folks should be thinking about should they try to make use of the practices we had talked about previously. And worry not, those episode links will be in the show notes for you to take a look at.

Folks in a lot of different roles and in a lot of different places are trying to figure out how AI tools are useful, where they fit, where they don't fit, and more. These aren't just tools in isolated development environments. These are things that are becoming integrated into the way work happens at various parts of what you'd call an enterprise stack. As you might imagine, this means a lot of different brains trying to do a lot of different things as we try to reach beyond our grasp.

As we think about the role technology should play in supporting a mission, an organization, a project, I'm of the mind that the goal should be to lower barriers to entry, to give folks close to the work access to tools, and then step back and see what creative stuff happens. Yeah, there are risks we'll have to control for, which is something we'll cover here together today, but building out bureaucratic walls that are too rigid can often lead to a cat and mouse game rather than compliance.

This leads to a thesis of a sort. If we provide a clear framework, avoid dictates forcing adoption, set up organizationally managed tools, and give easy access, and we focus on delivering real value to real people with our projects as quickly as we can that will land on solid footing. With this conversation, I'm hoping you'll get to see some key items that are useful for your personal or organizational playbook as you seek to navigate.

All this fun stuff.

The fundamental basis of the technologies we're talking about works on probabilities. It's all non deterministic, as you'll hear folks say. For many, this consideration is a change from what they're used to. Traditional IT governance relies heavily on things like static code reviews, structured change checklists, but those tools are built for predictable software in its cycles. With AI on models accuracy, for example, can silently degrade over time through something called data drift.

The input that's coming in from production use itself can bias the model's output. And since these tools involve semantic meaning, we also have to consider concept drift, where the meaning of what's being talked about can be the thing that's changing around your system. For example, the idea of what's good, whether it's a device, show, food, investment, can all change over time, which then impacts the relevance and quality of the output that came out.

Conversational AI interactions can also make security work a bit harder. Using data loss protection tools can lose its efficacy as the information we're trying to detect is buried in long interactions in which data structure doesn't easily match a signature of some sort. Or the interaction may just be happening somewhere the tools can't see. If an employee copy and pastes proprietary source code, customer PII, or trade secrets into a public model, that might go undetected.

And so we're put in a place where the challenges are a bit different than they've been in the past.

When you start thinking about structuring all this, it helps to separate your governance model from your operating model. Think of a governance model as deciding who has the decision rights and accountability, basically who decides. Your operating model is how things get resourced, organized, and executed, or how it runs. Most organizations look at three main paths that we'll talk about here centralized, decentralized, and hybrid models.

First, let's talk about the centralized model, which I've heard folks call the fortress. This is when decision making is concentrated in one core body, like a center of excellence, maybe a chief data officer, or some other role like that. It gives you really clean policy consistency, and the audit trail is obvious, but it does introduce operational bottlenecks that can make local innovation more difficult. On the other end of the spectrum,

You can go full decentralized with your approach. Authority is pushed out to individual departments or project teams. Innovation will likely move incredibly fast, but you'll pay for it with fragmented policies, technical debt, and reduced visibility from the center around what's going on. Then there's a hybrid approach, which is trying to apply the strengths of both ends in some sort of balanced manner. You apply guardrails from a centralized place.

But your individual business units have the autonomy to build and run deployments within those boundaries. it aligns with a data mesh philosophy of federated computational governance. I have some personal bias towards what the hybrid approach is trying to do. We shouldn't be handing down top-down mandates forcing people to use these tools.

Instead, let's give folks easy access to something organizationally managed and provide clear education and guidance and step back and watch. Once they have access, you can use solid user experience or UX evaluation practices to learn more. You can watch their behavior and iterate based on that real utility you're seeing in your organization.

From my own experience, I've seen engineering in the technology department I support lead and use these tools to quickly build context as they navigate large and sometimes old software repositories. A day of tedious exploration can be knocked down to something like twenty minutes of review. It's something that's akin to the SpecOps conversation I mentioned at the top of the episode, but it's again a combination of centralized tooling and guidance partnered with decentralized execution.

I've seen our product and design folks get prototypes together more quickly, allowing for a more robust conversation during our discovery and design processes, as another example. Ultimately, your focus should be on trying to line up with what works for the risk tolerance and threat model for your organization. If you're in a low risk tolerance, high attack surface kind of space, you may find that something more like that fortress, that centralized model.

Is the one that's going to fit best. But if you're on the opposite end, where with a high risk tolerance and maybe your need is more focused on innovation, it may well be decentralization that checks all your boxes. For a lot of organizations, I'd suspect it fits somewhere in between, with different parts of the organization demanding varied approaches depending on their portion of the organization's needs.

As you might imagine, a lot of smart folks have been thinking about these topics, so there's no need to start from scratch with most of it. There are some frameworks available to aid this work. There's ISO four two zero zero one, which seeks to define an AI management system, a key idea being that much of the risk of these tools comes not from the technology itself, but the structures or lack thereof for governing and using it.

In addition to that one, there's also the AI risk management framework from National Institute of Standards and Technology NIST. Frameworks like these can be a foundation for thinking through and building out your own organization's approach. Before we start down that path though, there's a foundation foundational principle that I'd suggest anchoring around, and that's that an AI tool cannot be held accountable.

A machine is not a legal or a moral agent. In any racy matrix, which stands for responsible, accountable, consulted, informed, or workflow you build, responsibility and accountability really need to land on humans. IS ISO 42001 really leans into this at a formal level. It explicitly states that generic team or system ownership shouldn't be part of how you design things.

To build a certifiable artificial intelligence management management system, you have to assign named primary and secondary human owners who have the documented authority to step in, pause, or hit a kill switch on that live system in real time. It requires writing over 20 mandatory documents covering everything from your top level AI policy to your risk treatment plan to ensure that that accountability we're talking about.

Is there at the at the institutional level. The NIST framework is also valuable, giving a lot of practical guidance if you're not pushing for that ISO certification. It covers governance topics like risk management, acceptable use policies, the work to align security with organizational objectives. It seeks to help with mapping activities like sorting out system boundaries, identifying dependencies, and noting constraints early on.

It nudges you to consider management around things like testing models for bias, monitoring privacy and data exposure, and taking the time to do red teaming. And it covers management areas like prioritization, human in the loop triggers, and runtime anomaly detection. Like all framework conversations in this space, there isn't really one right way to go or one right answer. Using one or some combination of elements.

with enough thought put into it can put you into a very solid place.

Let's also talk a bit about legal and regulatory conditions because there's quite a lot of them and they do change quickly. I'm not gonna go through every statute on the books in every place because we'd be here forever.

But instead, I want to focus on an idea that keeps showing up across different jurisdictions. one shift that should change how you think about AI risk, and then we can take a quick tour of stuff that's more about the tools you'll buy than the rules you'll have to follow.

Starting with that recurring idea I mentioned, it's something that lines up pretty well with something we talked about not too long ago the concept that an AI tool cannot be held accountable. Two of the largest regulatory regimes in the world are now busy with turning that principle into law, though they're maybe coming at it from different directions. Over in the European Union, there's GDPR Article 22. It's been around for a while, but

Worth understanding because it said in early baseline that a lot of other rules may now echo or use as a foundation. The core of it is a establishing a right. If you're a person being served inside the EU's borders, you have the right not to be subjected to a decision based solely on automated processing when that decision carries legal or similarly significant weight. Think like automated loan rejections.

or resume screening that filters you out before a human ever sees your application. The law says in some effect that a machine doesn't get the last word on something that important. There has to be a person in a process that you can appeal to.

We can also look at California, which is doing something somewhat similar with the C C PA's new automated decision making technology rules. These are enforced by the California Privacy Protection Agency, and they start landing in twenty twenty seven.

If you're using an automated system to screen people for employment, for housing, for credit, or health care, you're affected by this set of rules. You have to give people a pre-use notice that you're using the tool, and you have to offer them a way to opt out.

And there's also another big point, which is that you have to give them access to the actual decision logic not just, hey, AI said no, but some sort of account of what went into the no.

So you can picture a nonprofit's running a housing assistant program that uses some automated triage to decide who gets seen first. Under those rules, that's not just back office implementation details anymore. It's a decision that a real person has a legal right to understand and maybe even contest. And really notice what both of these laws are what they're trying to do. They're putting a human

in the loop through regulation and really having that need focused on the decisions that matter the most in people's lives. Which is again, you know, that human accountability thing that we keep circling around. Except now it's not just, hey, like you should have this as part of good system design, but it's a requirement with the force of law.

An old instinct was to treat an AI system as a piece of software. And when it misbehaves, you would think, that means there's a bug. With this newer kind of regulatory thinking, it treats your AI pipeline, your model endpoints, all of it as part of your attack surface. Europe's NIS2 directive is

does exactly this for critical services. The practical implication is maybe bigger than it just sounds coming from me. a prompt injection attack or a compromise of your training data. It's not just a a buggy file, but it's a security incident, something you have to log and report and respond to, much like any other security breach.

There's some signal that the United States is moving in a similar direction. There's a White House executive order from earlier this June that comes at AI through a national security lens. It directs CISA to issue binding directives hardening civilian government networks against AI orchestrated threats, and it sets up a voluntary process where developers of frontier models can

Give the government an early look before release.

What's relevant for the civic tech route specifically is that the order reaches past big tech. It talks explicitly about helping state and local government and critical infrastructure operators like rural hospitals, community banks, for them to get access to better defensive tooling.

And at the federal agency level, the trend has been to name actual humans to own this. Chief AI officer roles.

Those roles give formal accountability, a person whose job it is to answer for the risk. So we're on the same theme here again. Someone has to be accountable for these tools and their security.

And there's a another category of regulation or law that is like good to know, but might not really affect you that much, a lot of headline AI legislation right now is really aimed at those giant frontier model developers and not just organizations that are using them. California's Transparency and Frontier Artificial Intelligence Act.

which took effect in January, only applies to developers pulling in five hundred million dollars or more in revenue. It makes them publish how they guard against catastrophic risks. And that's defined somewhat narrowly as a model materially contributing to the death or serious injury of more than 50 people.

or more than a billion dollars in damage through something like

Helping create a weapon or running an autonomous

Cyberattack.

It also bakes in real whistleblower protections as an additional thing, which I imagine matters for a lot of folks in the public interest tech side of this. separately there's California's AI Transparency Act, which is more about content and kicks in this August and pushes large generative AI providers to watermark and disclose AI generated content. While companion law.

leans on those same developers to be transparent about the data those models are trained on.

You're, you know, out there listening, you're probably not likely going to be regulated by those things directly. but they may change or shape the tools that you're subscribing to or or using and the norms in which they operate on.

So with that in mind, it's probably worth at least having some awareness that this is going on around them.

I mean even now we have this situation with anthropic where there's a a frontier model that's gotten all tangled up in export control questions, which honestly wouldn't have been a thing I would have considered, you know, something like a month or even a year ago.

The through line across the whole messy landscape is actually pretty simple. It's the same thing we've been talking about all along. Make sure there's a human accountable for the decisions that matter, know what your tools are doing, and treat the AI in your stack as something to be governed, not just installed.

When you look at why these kinds of tech initiatives stall out or fail, it's usually because leadership treats governance like a static compliance checklist. Something you complete once, maybe throw it in a drawer, rather than something that's an active ongoing practice. And when IT approval cycles are too slow or bureaucratic.

Folks are naturally just gonna find ways around them. You know, they just wanna be able to complete their work for the day. It's rarely a malicious thing. It's usually just someone who's smart, who's trying to get their work done without it being too tedious or too much like drudgery.

But this does create some real vulnerabilities when it happens. As I was digging through some data, I ran across a stat from the cloud security lines showing that 82% of enterprises have discovered unauthorized AI agents or automated workflows running inside their network. In that same group, though, this is what makes that 82% interesting, is that 68% reported a high confidence in their visibility. So that's both an interesting and

Disconcerting gap there.

If an employee builds a custom workflow using a personal account, then leaves the organization, that history stays with them. And you'd end up with zero way to wipe that data or to revoke access. So how do we build a bridge over gaps like this using good modernization practice?

First, let's look at delivering value safely by providing sanctioned alternatives. Flat out banning popular AA pla AI platforms probably won't work, because it just ri drives your folks further into shadow IT. Instead, address the the root cause by putting a secure, corporate approved alternative in their hands, like it could be a private LLM instance, or you could get on some sanctioned enterprise plan with the one of your choice.

If you get that baseline tool into production quickly and you let people real people use it, get value out of it, you use good UX practices to learn from their use. you can iterate off of what they do and learn and get to a place where you know they're able to get something valuable and you're able to be assured of somewhat of the security. A second thing you can do is consider building an AI inventory that's living.

Doesn't have to be, you know, something super over engineered, but having some kind of repository to track where, you know, what tools are being used, maybe who the business owner is for the tool, where the data is coming from, where it fits in your organization's risk and threat models. you could even use a bit of like network scanning or DNS monitoring to discover those shadow tools we just talked about and try to pull them out of hiding and into an official path.

Also do discovery. You know, we've talked about UX practices, but this is one that really fits here. You know, gather a set of cross-functional folks and leaders from your different teams. if you just focus on what your tech staff are doing, you know, they're gonna naturalize for their kinds of workflows. Maybe you know, speed's important to them, latency, measuring model accuracy. But they're not naturally gonna look at data privacy laws or contract liabilities or structural bias necessarily.

But if you get a s lightweight committee up with that has, you know, your tech folks, but also domain experts, maybe even legal counsel or your compliance lead, your if you have privacy folks even better, you just want to make sure that those different perspectives have a voice in what you end up doing.

And lastly, think about the metrics you're gonna use, the things you're gonna measure. When you're looking at something like 90 days, you know, don't go don't reach for vanity stuff. You know, as I mentioned, like ad adoption's usually not your goal, so that wouldn't be the one I would reach for. but you could try to measure things like something that tries to represent administrative burden, like the number of touch points you've reduced in a bureaucratic process with your effort.

Or maybe it could be internal processing time. You know, an application takes this many less hours or days to complete. Or maybe on your tech team you're able to save compute on other operations because you made some process more automatic but have fewer steps. really the limit with this kind of thing is your imagination.

But what's important is that the thing is anchored in something that's actually valuable to people at the org.

At the end of this, what I really want to leave you with is that lowering the barrier to secure tools, keeping clear human accountability at the center of things like your racey charts, and focusing heavily on getting real value to real people quickly, if you do all that, you can build a governance framework that actually empowers your team instead of setting up something that ends up holding them back.

Episode Video