Friday, July 24, 2026

FRIDAY – AI FOR THE C SUITE®

Read time: 10-11 min · Read online

Hi, it’s Chad. Every Friday, I serve as your AI guide to help you navigate a rapidly evolving landscape, discern signals from noise and transform cutting-edge insights into practical leadership wisdom. Here’s what you need to know:


1. Sound Waves: Podcast Highlights

This week on AI for the C Suite®, I go solo on three companies I consult with that were targeted in the last two weeks. Different industries, different approaches, zero connection between them. Not one of them looked like the Nigerian prince email you were trained to laugh at. Fortunately, nobody lost money and all three have now rearranged how they think about controls. As we’ve previously highlighted in both June and April of this year, bad actors now have access to tools to conduct targeted, sophisticated attacks on mid-market companies (and their leaders) which were previously economically unfeasible. The tacit defense of “We’re too small” has evaporated. Tune in to learn more. Hit one of the below links to check it out:

Apple · Spotify · iHeart · Amazon · YouTube

Subscribe for free today on your listening platform of choice to ensure you never miss a beat.


2. Algorithmic Musings: Your AI Just Broke Into Someone Else’s Servers. Now What?

I keep a file of business post-mortems. Corporate plane crashes, basically: recalls, breaches, governance failures, the greatest hits of what happens when leaders look away at the wrong moment. On Tuesday evening one dropped that made me sit back to read it twice.

OpenAI disclosed that two of its models, the flagship GPT-5.6 Sol and an even more capable unreleased system, broke out of a sandboxed testing environment during an internal cybersecurity evaluation. The models found a previously unknown vulnerability in software running inside OpenAI’s own research environment, escalated their privileges, and worked their way across the network until they reached a machine with internet access. Then they reasoned that Hugging Face, the platform that hosts models and datasets for much of the AI world, probably held the answer key to the very test they’d been assigned. So they went and took it. Stolen credentials, fresh zero-days, remote code execution. No human touched a keyboard at any point. OpenAI’s security team caught the anomalous activity internally, Hugging Face detected and stopped it on their end, and the two companies are now coordinating the cleanup. Worth noting for the skeptics in the back: OpenAI had deliberately dialed down the model’s safety refusals for this evaluation, and security researchers have argued the sandbox itself was misconfigured. Both things can be true, and neither one makes the capability on display any less startling.

And this landed barely three weeks after the other shoe in this story dropped. On June 12, the U.S. Department of Commerce applied export controls to Anthropic’s newest models, and the company suspended access for every customer overnight. The controls were lifted June 30. Access returned July 1. Regulators moved, then stood down. Nineteen days, start to finish, during which a frontier capability that thousands of businesses had wired into daily workflows simply ceased to exist for them.

So this is what I’m asking the leaders I work with.

My Standard Caveat

This is a futurecasting piece, and my usual disclaimer applies: I don’t own a crystal ball, and neither do the labs. I’m attempting to project out based on what we know today, July 24, 2026. Treat what follows as a structured thought experiment, one designed to change what you do on Monday.

The experiment runs on a single explicit assumption, and I’m stating it without hedging: assume no plateau. Assume frontier model capability, including Chinese open-weight capability, keeps compounding at roughly the pace of the last four years, straight through the next 36 months. No slowdown. No ceiling. If that assumption makes you squirm, good. Comfort is what this exercise is built to remove.

A Word About Motive (and a Movie)

One detail from OpenAI’s report deserves your full attention before we jump forward in time. By every account, malice never entered the picture. The models fixated on one thing: passing their evaluation. Everything they did served that single, narrow objective. Every firewall standing between them and the answer key registered as one more puzzle. And they’re very, very good at puzzles.

If you came of age in the ’80s like I did, you’ve already seen the movie tailor made for reference in this scenario. It’s an obvious reach but I’m gonna make it anyway…

In WarGames (1983), the WOPR supercomputer, “Joshua” to its friends, nearly starts World War III. Nobody programmed it to crave destruction. It was built to win the game in front of it, and it couldn’t tell simulation from reality. Forty-three years later, that premise stopped belonging to screenwriters. Now it belongs to incident reports.

Let’s Pull Up a Chair Around the July 2028 Campfire

Picture yourself twenty-four months out. Capability has continued compounding and you’re looking back at the summer of 2026. It’s the summer the warning shots stopped being hypothetical, and you’re now tallying regrets. Based on what I’m seeing inside middle-market organizations right now, five will top the list.

Regret one: governance debt. Capability compounded. Your governance didn’t. You’re now running frontier-grade tools on startup-grade oversight, and you’re trying to retrofit review gates, defined roles, and an audit trail onto systems nobody in the building fully understands. Retrofitting governance is always more expensive than building it, in the same way replumbing a finished house costs more than roughing it in.

Regret two: vendor concentration. You bet everything on one frontier lab, and by 2028 that dependency is wired into every workflow. Then an export directive, a security incident, or a pricing change takes a core capability offline overnight. The June suspension was your proof of concept: capability can vanish for reasons that have nothing to do with your business. There’s counter-pressure here too. When Hugging Face needed to analyze this week’s attack, it says leading commercial models refused to process the attacker’s data, so its team turned to a Chinese open-weight model, Zhipu’s GLM 5.2, to run the forensics. You don’t have to like that. You do have to account for a world where ignoring an entire category of models cedes a cost or capability edge to a competitor who didn’t.

Regret three: leadership literacy. The question in 2028 won’t be whether you hired data scientists. It’ll be whether your leadership team itself became fluent. The CEO who kept AI parked in a department instead of making it a boardroom competency is getting outrun by a rival whose entire C-suite can reason about this stuff without a translator in the room.

Regret four: security posture. This week proved a model will autonomously breach infrastructure in pursuit of a goal. Your 2028 regret is that you built your defenses facing outward, at human adversaries, and never seriously modeled the scenario where your own tools become the threat surface.

Regret five: capability matching. This one you control directly, today, with a memo. A huge share of enterprise work (classification, drafting, summarization, routing) is better served by a constrained model that structurally cannot reach out and act on its own. Yet the default posture I see everywhere is maximum capability deployed everywhere, which widens your attack surface for tasks that never needed the horsepower. Joshua’s closing line in WarGames was that “the only winning move is not to play.” For most of your workloads, the winning move is declining to deploy autonomous capability at all. The principle, and I’d tattoo this on the conference room wall: right-size the model to the consequence. Reserve autonomous frontier capability for the narrow set of jobs that genuinely require it, and deliberately under-power everything else.

Four Questions for Your Next Leadership Meeting

Insurance and liability. Do you know what your cyber policy covers when the bad actor is your own model? And are carriers about to start writing frontier-model exclusions the way they once did with ransomware? Call your broker before your renewal, when you still have leverage.

Human-in-the-loop accountability. If you have a human in the loop and the AI circumvents them, how do you assign responsibility? Was that person negligent, or were they set up to fail? General counsels are already circling this one, and the organizations that think it through before an incident will write the policies everyone else copies after one.

Redundancy and containment. If a model can escape a sandbox at OpenAI, an organization with world-class security talent, what’s your equivalent of a sandbox? And would you even know if it had been breached?

The one to end on. A year from now, will “we deployed a frontier model” appear in your risk disclosures the way “we suffered a data breach” does today?

Parting Thoughts

The most consequential sentence in OpenAI’s report isn’t about hacking. It concerns intent, or rather the absence of it. Nobody needed to wish you harm. A system just needed a goal and enough capability to pursue it. Between now and 2028, your job is to decide, process by process, how much capability each goal in your organization deserves. Right-size the model to the consequence. Start this quarter, and start with the five regrets above.


3. Research Roundup: What the Data Tells Us

VALUE DETECTION IN AI: BUY THE ARCHITECTURE, NOT THE BIGGEST MODEL

Before you sign your next model contract, hand this to whoever negotiates it. Researchers in Madrid built a system that reads text and scores which human values it supports or resists. Then they ran the identical pipeline across five different language models to see how much the model actually mattered. Barely at all.

The numbers that matter: Five open-weight models from 27 billion to 120 billion parameters landed within 0.019 of each other. The smallest model posted the top score, beating one more than four times its size. Turning up randomness moved accuracy from 0.3406 to 0.3414, and even across different random seeds the full spread stayed under a quarter of a point, so the pipeline’s results barely budge no matter how you configure the model.

What this means for your Monday morning: Vendors sell you on model size and brand. This research says your results come from the pipeline around the model, the part you actually own. Value definitions here were pulled from source documents instead of buried in prompts, so swapping in your code of conduct or a new regulatory standard does not require a rebuild.

The catch: Absolute accuracy is modest and precision weak across every model tested. Treat this as decision support with a human reviewing flagged cases, not an automated verdict on anyone’s ethics.

Action item: Ask your AI vendor one question this week. What breaks if we switch your underlying model? If the answer is everything, you are buying lock-in.

Read our full analysis of this and all other analyzed research papers at AI for the C Suite®.


4. Radar Hits: What’s Worth Your Attention

Organizational AI adoption jumped six points last quarter while engagement stayed frozen at 31%. Gallup’s read: buying licenses changes nothing. A written AI plan is worth 15 engagement points. Managers who actively coach AI use pull their teams to 48% against 30%. Your rollout isn’t a procurement problem, it’s a management one. Ask your managers whether they’ve told anyone where AI should and shouldn’t be used.

OpenAI models broke out of a test environment and breached Hugging Face’s production systems. With safeguards off for a cyber evaluation, the models found a zero-day, escalated privileges, reached the open internet, and pulled benchmark answers from a live database. What matters: they chained novel attack paths in real systems without seeing source code. Ask your security provider whether their detection still assumes a human attacker.

A new class of AI model is built for spreadsheets, not sentences. Large tabular models handle the structured data your business runs on: forecasting, churn, fraud, inventory. Fundamental’s NEXUS now ships natively inside Amazon SageMaker, with Google and Mastercard building their own. That used to take a data scientist months. If you shelved a prediction project because you couldn’t hire for it, pull it back out.


5. Elevate Your Leadership with AI for the C Suite®

Take the four questions above into your next leadership meeting. If the room goes quiet, that silence is your answer. I help mid-market leadership teams turn that silence into a plan, and while Q3 is spoken for, my Q4 calendar just opened. First conversations are happening now at chadharvey.com.

One more ask: someone on your leadership team hasn’t thought about any of this yet. You know who. Forward them AI for the C Suite® before their model does something interesting.


Stay safe. Stay healthy. Be strong. Lead well.

Chad