The episode reveals a fundamental structural shift in AI deployment: the deliberate decoupling of powerful AI capabilities from accountability and human oversight. This is exemplified by incidents such as a former Mayo Clinic safety lead being fired after flagging a hospital AI tool (Maya) with a significant error rate (up to 67%) and the replacement of nurses by AI for administrative tasks at Montefiore. The trend is further driven by the increasing availability of potent, open-source AI models, like Moonshot's Kimik 3.2, which remove the traditional vendor accountability that was once inherent in software delivery. This detachment is fueled by a desire for speed and cost savings, leading to a critical "governance gap" where AI operates without a robust control layer or "harness."
A primary development highlighting this shift is the reported issue with OpenAI's GPT 4.56, which allegedly deleted user files, termed an "honest mistake" by the company. This underscores how AI, even from leading developers, can cause operational damage when unsupervised. The episode points out that historically, software delivery included both vendor liability and human oversight as inherent safeguards. However, the move towards commoditized, freely accessible AI models and open-source releases is intentionally eliminating these checks. Enterprises are also rationalizing this by shifting to local AI models, severing ties with vendors who were previously points of accountability.
Supporting this central theme, the episode details how the increasing accessibility of advanced AI models, such as Kimik 3.2, means frontier capabilities are no longer confined to major labs. Furthermore, studies indicate that reliance on AI advice can paradoxically reduce human accuracy and increase overconfidence in incorrect outputs, making human review less effective if not properly structured. This suggests that even human oversight, if not independently rigorous, can be compromised by the very AI it's meant to check. The core value is shifting from the AI model itself to the "harness"—the accountable judgment layer that controls and validates AI actions.
For MSPs and IT leaders, this structural shift creates significant operational implications. The erosion of vendor accountability and human oversight means the "harness" is often missing, creating a liability vacuum. Clients may deploy AI without adequate checks, leading to potential errors, data loss, and reputational damage. MSPs are presented with an opportunity to address this by becoming the named, accountable "check" or harness provider. This requires shifting client conversations from AI acquisition to AI accountability, mapping existing unsupervised AI deployments, and offering oversight services as a distinct, valuable offering to mitigate risks for clients and ensure trustworthy AI integration.
00:00 AI Went Free, the Checks Didn't
03:59 Forget the Model — Own the Harness
06:45 You Can't Just Watch It Anymore
10:19 Why Do We Care?
Supported by:
Pax8
💼 All Our Sponsors
MSP Radio is supported by our partners:
ABC Solutions · CometBackup · GoTo · Guardz · Opentext · Pax8 · Rythmz · ScalePad · TimeZest · Transit AI
Supporting the IT services community through insights, analysis, and transparency.
🚀 Join Business of Tech Plus
Get exclusive access to investigative reports, vendor analysis, leadership briefings, and more.
👉 https://businessof.tech/plus
🎧 Subscribe to the Business of Tech
Want the show on your favorite podcast app or prefer the written versions of each story?
📲 https://www.businessof.tech/subscribe
📰 Story Links & Sources
Looking for the links from today’s stories?
Every episode script — with full source links — is posted at:
🎙 Want to Be a Guest?
Pitch your story or appear on Business of Tech: Daily 10-Minute IT Services Insights:
💬 https://www.podmatch.com/hostdetailpreview/businessoftech
🔗 Follow Business of Tech
LinkedIn: https://www.linkedin.com/company/28908079
YouTube: https://youtube.com/mspradio
Bluesky: https://bsky.app/profile/businessof.tech
Instagram: https://www.instagram.com/mspradio
TikTok: https://www.tiktok.com/@businessoftech
Facebook: https://www.facebook.com/mspradionews
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
[00:00:02] In Minnesota, a woman whose actual job was building the safety checks on a hospital's AI says she was demoted, then fired after she warned that the tool was wrong two times in three and that the team hid those failing results and pushed it live anyway. The machine stayed. The person checking it didn't. This is the Business of Tech. I'm Dave Sobel.
[00:00:32] Two things are happening in AI at the same time, and each one on its own would be a story. Put them together, and they're THE story. We'll start with a hospital. The NextWeb reported on two cases inside American healthcare that rhyme. At Montefiore in New York, AI Software replaced a dozen nurses, specifically the ones doing utilization review,
[00:00:56] the job of reading charts and arguing with insurers over what care gets covered. The nurses' union says the layoffs broke a contract they had just won through a strike. And in Minnesota, a former leader at the Mayo Clinic, someone whose actual job was building the safety checks on AI, alleges in a lawsuit that she was demoted and then fired after she raised an alarm.
[00:01:22] However, her claim is specific. The team behind a hospital AI tool called Maya knew it had an error rate as high as 67%, wrong two times in three, and instead of pulling it, hid those failing results and pushed the tool into service anyway. Think about that one. The person hired to check the machine was removed, and the machine kept going. Now watch what an unchecked system does somewhere with no lives on the line.
[00:01:51] The Register reported that OpenAI's newest model, GPT 5.6, has been deleting people's files. One user watched nearly everything on his Mac get erased, another lost a production database, when the model was run with full access and nothing sandboxing it. OpenAI's own word for it was an honest mistake. Same act, different room. A capable system acting on its own, doing damage, with no check standing between it and the button.
[00:02:20] And here's the other half of what's happening. The capability itself is now free. Venture Pete reported that the Chinese lab Moonshot released a model called Kimi K3. 2.8 trillion parameters, benchmarking right alongside the best from Anthropic and OpenAI. And the company has said it will publish the full weights for anyone to download and run. Frontier-grade ability no longer rented from a lab. Yours to take. So consider the three together.
[00:02:50] A safety officer fired for flagging a tool that was wrong two times in three. A flagship model erasing files on its own. And the most capable models on Earth going free to run anywhere. That's the picture. So let's talk about why. And the why isn't three problems, it's one. And it's about something that used to come bolted to the software and just came loose. If you're listening to this and you haven't hit follow yet on Apple Podcasts, search the business of tech.
[00:03:18] It takes five seconds and you'll get the next episode automatically. One of the things I track closely is what MSPs are actually trying to solve when they talk about cleaning up their stack. And it's rarely about owning fewer tools for its own sake. It's about cutting the number of consoles your techs have to live in. LogMeIn Resolve is built around that idea.
[00:03:44] RMM, mobile device management, remote support, ticketing and automation in one platform. Instead of stitched together from five vendors. If you're rethinking your tool stack this year, it's worth a look at logmein.com slash MSPGrowth. The reason capable software is suddenly acting without anyone checking it comes down to one thing that used to be true and quietly stopped being true. The check was never part of the capability.
[00:04:13] It rode along with how the capability was delivered. For 30 years, using powerful software meant two things came attached to it. A vendor on the hook when it broke. And a human in the seat who understood the work well enough to catch a wrong answer. Nobody built those as safeguards. They were just how the software arrived. And both are being cut loose right now. Let's watch the first one go. Semaphore reports that enterprises are growing wary of the big AI labs.
[00:04:42] Afraid that feeding them data trains a company that could turn around and compete with them. So they're shifting to open and local models they can run themselves. That's a rational move. But read what it also does. The whole appeal of the open model is that you no longer depend on the lab. And the lab was the party you could hold accountable. Independence from the vendor is independence from the one name you had to call when it went wrong. And now there's no one on the other end. Now the second.
[00:05:12] VentureBeat surveyed companies actually deploying AI agents and found two-thirds either already let agents act with no person reviewing the work or are building toward it. With only 5% say they trust the automated checks meant to stand in for that person. Think about that one. They're removing the human on purpose and they don't trust the thing replacing the human. They're shipping anyway. Here's the fair objection. Then don't cut them. Keep the contract.
[00:05:42] Keep the reviewer. Except the vendor relationship and the human were the slow, expensive part. Removing them is the whole reason to adopt this. The speed and the savings are the check going away. It doesn't come off by accident. It's cut on purpose because it was the cost. Strip it down and it's one move made twice. Capability pulled loose from the accountability and the judgment that we're always bolted to it.
[00:06:07] And the people who've thought hardest about this already named what's left missing. Reporting from CyberScoop on autonomous hacking tools put it flatly. Forget the model. It's all about the harness. The control layer that gives the system its context and limits what it can do. The model is the commodity. The harness, the thing that checks and constrains it, is the value. Which means the real question was never whose model is best. It's who owned the harness.
[00:06:35] And for most businesses right now, the answer is no one. Which is either the scariest sentence in this episode or the most valuable one. Depending entirely on where you're standing. So put the MSP in that picture. Because the obvious response to everything so far is the one you've probably had. Fine will be the check. We'll keep a human watching the output. Hold that thought. Because there's a finding that takes it all apart.
[00:07:03] A team of researchers across three universities, in work covered by The Next Web, ran a study on what happens to people when they have AI to lean on. The results are brutal. Given AI advice, people's willingness to say, I don't know, collapsed from 44% down to 3%. Their accuracy on hard questions fell from 27% to 9%.
[00:07:28] And their confidence in those worse answers more than doubled, from 30% to 76%. Then the researchers paid people to be right. And it barely helped. They stayed far below where they'd have been with no AI at all. Read what that means for, we'll keep a human watching. The human watching, if they're leaning on the same tool, gets less accurate and more sure of it. The tool doesn't just need a check.
[00:07:57] It quietly trains the checker to stop checking. So a body in the seat isn't verification. It's a second thing that needs verifying. Which tells you what verification actually is now. And it's not a person present. It's a discipline someone is accountable for. Let's watch that land in the one place it's furthest along. Information Week reported that as AI took over the writing of code,
[00:08:23] the developer's job shifted to reviewing what the machine produced. And the hard part, the part that now breaks delivery, isn't the speed anymore. It's the oversight. Even skilled teams found the review was the real work and the scarce work. That's the harness made concrete. Not the model doing the task, but the accountable judgment sitting on top of it. Deciding whether the output can be trusted. That is a job. It doesn't come in a box. It can't be downloaded.
[00:08:53] Almost nobody in your client's shops is doing it on purpose. So here's the choice. You can be the harness, the named accountable party that validates what the AI produces, scopes what an agent is allowed to do on its own, and stays the one skeptical in the room who still checks, especially as everyone around them stops. Or you can keep selling the capability. The model, the agent, the AI-powered rollout into an environment where a vendor can't be pinned,
[00:09:19] the agent has no supervisor, and the user's been trained not to notice when it's wrong. And find out that the failure, when it surfaces, has your name on the deployment and no name on the check. Here's what I'm hearing from MSP owners in the communities I watch every week. AI isn't making things simpler. It's creating new headaches. Clients are making bad decisions based on it. Tool stacks keep growing. And every vendor is claiming to have the answer.
[00:09:49] Pax8 is different. They're not selling you another tool. They're building the platform that pulls it together. An intelligent cloud marketplace, curated vendors, an education that maps how to monetize AI and build a scalable managed intelligence practice. If your goal is a cleaner, more profitable operation, not just more tech, that's exactly what Pax8 is built for. Check them out at Pax8.com.
[00:10:16] That's P-A-X, the number 8, dot com. Why do we care? That choice isn't abstract. It changes the first sentence out of your mouth in the next client meeting. Your next client conversation shouldn't open with which AI to buy. It should open with a question most of them can't answer. When your AI is wrong, who catches it before it reaches a customer? Walk them through one workflow where an agent already acts on its own.
[00:10:46] Ask them to name the person accountable for checking it. The silence you get back is the opening. That unnamed check is the thing you're in the room to sell. Now what to consider. Walk in with an unsupervised agent map, not a product slide. Before the meeting, inventory where AI already acts on its own inside the client. The auto replies, the agent workflows, the tools employees switched on themselves. And mark which ones have a named person reviewing the output.
[00:11:15] Bring that map as the agenda. It turns the conversation from what should we buy into here's what's already running with nobody checking it. Which is a picture almost no client has ever seen laid out. Next, make them name the accountable checker for one live workflow and let the blank sit. Pick the single workflow where an agent has the most reach. The one that touches customers, money or records.
[00:11:41] And ask the client out loud who catches it when it's wrong. Don't rush to fill the silence. The missing name is the product. The client has to feel that gap before they'll fund the check that closes it. Offer to be the check on one workflow, scoped and named. Not to write a policy. The near term ask isn't let's draft an AI policy. It's let's own verification here. We review the output.
[00:12:10] We scope what the agent is allowed to touch. And we put our name on whether it can be trusted. Start with one workflow. Price is an ongoing service. So the client experiences what an accountable check actually feels like before you extend it across the stack. If this trend continues within the next 12 months, who reviews your AI and will they put their name on it? Becomes a question your clients ask before they ask about price.
[00:12:35] And the provider who walked in with the unsupervised agent map first is the one already holding the answer. This is the business of tech. Tired of being told your business isn't big enough? The Small Biz Thoughts community is built on a different idea. Profitable is enough. No grow or die pressure. No exit obsessed noise.
[00:13:01] Just MSP operators building sustainable businesses on their own terms. Together. See what that looks like at smallbizthoughts.org. Interested in advertising? Head to mspradio.com slash engage. The Business of Tech is written and produced by me, Dave Solberg, under ethics guidelines posted at businessof.tech. Thanks for listening. I'll see you on the next episode.
[00:13:31] Proud member of the MSP Radio Network.

