Category: AI Strategy

  • The Customization Tax Nobody Talks About

    The Customization Tax Nobody Talks About

    The Customization Tax Nobody Talks About

    A mid-sized logistics company spent $400,000 on enterprise AI software last year. Great product. Proven ROI. Used by dozens of their competitors.

    Six months later, they wondered why their AI investment hadn’t created competitive advantage.

    The answer was simple: they’d paid for efficiency, not differentiation. Every one of their competitors could write the same check.

    Most executives miss this: buying off-the-shelf AI tools gives you table stakes, not advantage. The customization you skip to save money today becomes the competitive gap that costs you market share tomorrow.

    But I’ve also seen startups burn through runway building custom AI systems for commodity capabilities where excellent solutions already existed. Eighteen months and three engineers building what they could have licensed in an afternoon.

    The question isn’t whether to build or buy. It’s knowing which capabilities deserve custom investment and which don’t.

    The Build Calculus Has Changed

    Three years ago, the rule was clear: if you’re not a software company, don’t build software. A hospital shouldn’t build its own electronic health records. A retailer shouldn’t build its own e-commerce platform. Software required specialized teams, long timelines, and maintenance that distracted from core business.

    That calculus shifted.

    I built my personal AI assistant system, Jarvis, over a weekend. It manages contacts, surfaces information before meetings, sends morning briefings, handles dozens of automated workflows. Three years ago, that would have required a development team and six months. Budget: $200,000 minimum.

    Current hard costs: about $50 per month in API fees and hosting.

    But here’s what matters for strategy: Jarvis is designed for exactly how I work. My specific meeting prep needs, my relationship tracking approach, my morning routine. No vendor would build that because the market for “Sean’s particular workflow preferences” is exactly one person.

    That customization is the point. Off-the-shelf tools are lowest common denominator by design. They serve the average user. When you build, you create tools that fit your specific workflows, your specific data, your specific competitive context.

    A caveat before anyone gets excited: I already understood cloud infrastructure, databases, and API integrations. AI accelerated my building, but it didn’t eliminate the need for foundational technical literacy. The barrier has dropped from “you need a development team” to “you need technical literacy plus AI assistance.” That’s massive. It’s not “anyone can build anything with zero knowledge.”

    Three Questions That Clarify the Decision

    When evaluating whether to build custom AI capabilities or buy existing solutions, three questions cut through the noise:

    Do you have proprietary data that makes your AI better than anything you could license? A payments company with transaction data across millions of businesses revealing fraud patterns no single merchant could detect? Building fraud detection makes strategic sense. Using the same public datasets everyone else accesses? Buy the mature solution.

    Is this capability core to how you differentiate? A pizza company built its own AI ordering system because they had data about ordering patterns nobody else had, and the capability was customer-facing and central to their experience. Their cloud infrastructure? Bought from a vendor. Hosting is commodity.

    Can you attract and maintain the talent to build and iterate? Building isn’t a one-time project. It’s an ongoing capability. If you can’t staff it sustainably, buy even when building seems strategically attractive.

    The honest assessment most leaders avoid: your organization is probably less ready to build than you think. Most companies think they’re more mature than they are. That gap matters because organizations that aren’t ready should be buying more than building, learning more than implementing.

    What You Can Do This Week

    List every AI capability your organization currently uses or is considering. For each one, answer the three questions above. Honestly.

    You’ll quickly see which capabilities justify custom development and which should be purchased.

    The pattern: build clusters around proprietary data and core differentiation. Buy clusters around speed to market and commodity capabilities. Alliance (the third option) emerges when you need complementary capabilities that would take too long to develop alone.

    Most organizations should be building fewer custom AI systems than their engineers want and more than their executives think possible. The mistake isn’t picking the wrong answer. It’s applying the same answer to every capability without thinking through what creates advantage.

    The full framework, including the maturity model, the strategic sourcing matrix, and worked examples across industries, is in AI Strategy for Business Leaders. I’ve also posted templates and assessment tools at my companion resources that walk through the build-buy-ally decision step by step. And if you want to understand how I approach the intersection of technology and strategy more broadly, my books cover the frameworks I return to again and again when the field shifts faster than the playbooks.

  • The Most Expensive Question You Never Ask

    The Most Expensive Question You Never Ask

    The Most Expensive Question You Never Ask

    A founder came to my office last month with a twelve-slide pitch deck. Beautiful design. Impressive revenue projections. Three different market opportunities, each requiring different capabilities, different partners, different brand positions.

    “Which one should I pursue?” he asked.

    Wrong question. The right question: “Which one am I actually built to win?”

    Most strategic failures don’t happen because people choose bad opportunities. They happen because people choose good opportunities that require someone else’s foundation. You can’t borrow another company’s mission. You can’t rent their accumulated capabilities. And when you try, every dollar you spend fighting uphill against your own structure is a dollar you’re not spending on the work you were designed to do.

    I learned this the expensive way. The story involves a company called Nouri, a product pivot, and the slow realization that chasing an attractive market while abandoning your competitive advantage is just expensive self-sabotage. You’ll find the specific numbers and the full postmortem in the book. What matters here is the framework that could have prevented it.

    The Pyramid You’re Standing On

    Every strategic decision rests on three layers, whether you’ve drawn them out or not.

    At the base: your values, mission, purpose. The stuff you’d defend even when it costs you money. Middle: your resources and capabilities. What you actually have and what you can actually do. Top: your activities. The work you’re doing right now.

    The layers either reinforce each other or they don’t. When they align, you get compounding advantage. When they conflict, you get expensive confusion.

    Most people skip straight to activities. “Should we add this feature? Enter that market? Hire for this role?” But activities divorced from the lower layers become random. You’re optimizing locally while drifting globally. You’re rearranging deck chairs while the hull fills with water.

    Start with the base. What are you here to do? Not the mission statement you workshopped for the website. The real answer. The one that determines which opportunities you’ll say no to even when they look profitable.

    Then the middle layer: what do you have to work with?

    Resources break into six categories I call BLINKA: Brand, Land, Information, Network, Knowledge, Assets. That’s the “what you possess” side. Capabilities are the “how you execute” side. A company can have a strong brand and terrible marketing capability. A person can have a powerful network and weak relationship management skills. Both matter. Confuse them and you’ll wonder why your advantages don’t translate.

    Only after you’ve mapped those two layers should you choose activities. Here’s the test: can you draw a straight line from each activity down through a capability or resource, all the way to a core value or mission priority?

    If you can’t, you’re spending energy on something that won’t compound.

    What Coherence Actually Looks Like

    Walt Disney built one of the most coherent pyramids in business history. Want to see what alignment looks like when it’s real? The book walks through the whole structure. Base layer: family-centered, optimism, quality, honoring the past. Middle layer: creativity, animation skills, business acumen, strategic location, marketing genius. Top layer: create animated, family-friendly stories.

    Every activity traced to a capability. Every capability served a value. The pyramid was load-bearing in both directions.

    Most pyramids aren’t coherent. They’re archaeological sites. Layers from different eras, serving different missions, built by different teams who never talked to each other. The activities at the top reflect last quarter’s priorities. The capabilities in the middle reflect whoever you managed to hire. The values at the base reflect a founding vision nobody’s looked at in years.

    The work isn’t building a pyramid from scratch. The work is excavating the one you’re already standing on, seeing where the layers conflict, and making the hard choices about which activities to stop.

    One Thing You Can Do Today

    Draw your pyramid. Personal or organizational, your choice. Three layers. Be honest about what’s there, not what the mission statement says should be there.

    Base: What do you actually value? What would you defend even if it cost you an opportunity?

    Middle: What resources do you really have (BLINKA: Brand, Land, Information, Network, Knowledge, Assets), and what can you do well?

    Top: What are you spending time on right now?

    Now draw lines. Which activities connect to which capabilities? Which capabilities serve which values? Where are the gaps?

    Where are you doing work that doesn’t rest on any capability you have? Those gaps are where you’re bleeding energy.

    The full framework, including how to use this with AI strategy and the complete Disney analysis, is in AI Strategy for Business Leaders. I’ve also put pyramid templates and worked examples in the companion resources so you can map your own organization without starting from a blank page.

    The pyramid doesn’t tell you what to build. It tells you what foundation you’re building on. Most people skip that question until it’s too expensive to ask.

  • Why Your AI-Calculated NPS Might Be Wrong (And Why That’s Not the Real Problem)

    Why Your AI-Calculated NPS Might Be Wrong (And Why That’s Not the Real Problem)

    I asked Claude to calculate the Net Promoter Score for a dataset last week. Forty-seven survey responses, mix of scores from zero to ten. The AI returned an NPS of 34, showed its work, and the math checked out.

    It had miscounted the categories. Treated some sevens and eights as promoters when they should have been passives. Miscategorized a few responses entirely. It arrived at a plausible answer through wrong math, which is worse than getting the wrong answer. A wrong answer you catch and fix. A right answer from flawed reasoning you trust and propagate.

    This happens more often than you’d expect with NPS. Something about the categorical boundaries (0-6 detractors, 7-8 passives, 9-10 promoters) confuses language models. Modern frontier models have gotten better by defaulting to code execution when they recognize math problems, but older models, smaller models, and chat-only contexts still slip.

    The lesson isn’t that AI is unreliable for feedback systems. The lesson is what that miscalculation reveals about verification, about what customer feedback actually does for strategy, and about the difference between measuring satisfaction and predicting what happens next.

    The Thermometer Problem

    Most companies treat NPS like a thermometer. You take a reading, write down the number, feel vaguely good or bad about it. Score above 30? Good. Above 50? Excellent. Above 70? World-class. The number goes in a dashboard somewhere, gets mentioned in a quarterly review, then everyone moves on.

    That’s operational hygiene, not strategy.

    The value of customer feedback shows up in three places that have nothing to do with whether your score is 34 or 42.

    First, feedback is a leading indicator. Declining NPS shows up in customer comments months before it shows up in revenue. By the time your finance team sees the churn line bend, your customer success team had the signal in detractor feedback ninety days earlier. By the time your sales team feels pricing pressure, NPS for price-sensitive segments had already drifted downward two quarters before.

    Second, feedback is balance-sheet infrastructure. When LexisNexis acquired my analytics company (full story in the book, Chapter 12), the feedback system and its high NPS were part of the valuation. Acquirers pay for customer satisfaction infrastructure because it predicts retention, expansion revenue, and referral economics better than trailing financials do.

    Third, feedback connects to decisions or it connects to nothing. Stripe’s CEO invites customers into bi-weekly leadership meetings for the first thirty minutes. Forty executives, one customer, direct voice to people who can act. Apple adjusted the iOS 15 Safari redesign after beta tester backlash, before public release. Both examples share the same pattern: the system connects customer voice to company response without bureaucratic delay.

    What AI Actually Changes

    AI changes the economics of every stage in a feedback system, but not the way most people assume. The value isn’t in calculating NPS (which, as we’ve established, it sometimes does incorrectly). The value is in making systematic collection and intelligent analysis cheap enough to run continuously.

    Survey design used to need expertise in question construction, bias avoidance, response design. Now you prompt an AI with constraints (keep it under three minutes, avoid leading questions, use consistent scales) and get a professionally designed survey in minutes. You still review and adjust for your context, but the first draft is done.

    Response processing used to need analyst hours. Now every survey response triggers an automated workflow: calculate the NPS category, feed the open-ended text to an AI node that extracts sentiment, themes, urgency, and a one-sentence summary, return structured JSON, store in your database. If a detractor response flags as urgent, alert your customer success team immediately via Slack. No manual compilation. No responses sitting unread in a spreadsheet.

    Weekly synthesis used to be someone’s job. Now a scheduled workflow pulls the week’s responses, calculates the trend, generates an executive summary with top themes, wins, concerns, and recommended actions. Different teams get different views: product sees feature requests, support sees common problems, sales sees competitive mentions. The intelligence flows to where it can drive action.

    But verification remains human responsibility. When AI calculates NPS, have your workflow also calculate it directly in code and flag discrepancies. When AI identifies urgent issues, have human review before action. When AI recommends changes, treat those as hypotheses to evaluate, not decisions to execute. The pilot doesn’t let the autopilot land without watching the instruments.

    One Thing You Can Do Today

    Audit what feedback you currently collect. Surveys, support tickets, sales notes, social media mentions, product reviews, usage analytics. List all sources. Then rate each one: systematic or ad hoc? Analyzed or accumulated? Actioned or ignored?

    Most companies collect far more feedback than they realize and use far less of it than they should. The gap between collection and action is where value gets lost. AI doesn’t fix that gap automatically. AI makes it cheaper to close the gap if you build the system intentionally.

    Start with one trigger point. Post-purchase survey sent seven days after delivery. Post-support survey sent twenty-four hours after ticket closure. Quarterly relationship health check for ongoing customers. Pick one, automate the collection, route the responses somewhere a human will actually read them, and connect what you learn to a decision that might change.

    The compound value of feedback grows as the system matures, but only if the system connects to decisions. Otherwise you’re just generating data that confirms what you already believe while the actual signals (the declining sentiment in a specific segment or the feature request that keeps appearing in different words) sit unnoticed in a database no one queries.

    The full feedback system architecture (the five-stage flywheel from collection through closing the loop, the n8n workflows, the AI analysis prompts, the competitive intelligence angle, the leading indicator math) is in Chapter 18 of AI Strategy for Business Leaders. The prompt templates, automation workflows, and database schema are free at the companion resources page. But the place to start is simpler: figure out what feedback you already have, pick one source, and build one path from collection to action. Everything else is scaling what works.

  • An AI Agent Hacked McKinsey in Two Hours for $20. Here’s What I Did Next.

    An AI Agent Hacked McKinsey in Two Hours for $20. Here’s What I Did Next.

    This week, a cybersecurity startup called CodeWall pointed an autonomous AI agent at McKinsey & Company’s internal AI platform. No credentials. No insider knowledge. No human in the loop. Just a domain name and what the researchers called “a dream.”

    Two hours later, the agent had full read and write access to the entire production database behind Lilli, the AI tool that 72% of McKinsey’s 43,000 consultants use every day.

    The total cost? Twenty dollars in API tokens.

    The basics matter more than the breakthroughs; they always have. And this McKinsey breach is one of the most important AI stories of the year. Not because the attack was clever. Because it wasn’t.

    The Attack Was Embarrassingly Simple

    CodeWall published the full technical writeup on their blog, and it’s worth 10 minutes of your time. But I’ll walk you through the core of it.

    McKinsey launched Lilli in July 2023. It handles chat, document analysis, RAG search over decades of proprietary research, and AI-powered queries across 100,000+ internal documents. Over 40,000 employees use it. 500,000+ prompts a month.

    CodeWall’s agent found publicly exposed API documentation. Over 200 endpoints, all neatly documented. Most required authentication.

    Twenty-two didn’t.

    One of those open endpoints wrote user search queries to the database. The values were safely parameterized (good), but the JSON keys, the field names in the request, were concatenated directly into SQL (very bad). That’s a SQL injection vulnerability. It’s been a known bug class since the 1990s. I was teaching people to watch for this when I was building ATAC Workstation in the late nineties.

    The agent ran 15 blind iterations. Each error message revealed a little more about the database structure. Then production data started pouring back.

    What it accessed: 46.5 million chat messages covering M&A strategy and client work. 728,000 files with confidential client data. 57,000 user accounts. 384,000 AI assistants. 94,000 workspaces. All in plaintext. Zero authentication required.

    Forget the Data. The Prompts Are the Real Story.

    The data leak is bad. Obviously. But the part of this story that keeps me up at night is something most people aren’t talking about yet.

    Lilli’s 95 system prompts were stored in the same database.

    Those prompts are the instructions that tell the AI how to behave. What questions to answer. What to refuse. How to cite sources. What guardrails to follow. And the agent had write access to all of them.

    Think about that for a second. An attacker could rewrite those instructions. Silently. No code deployment. No change management ticket. No alert in any monitoring system. Just one SQL UPDATE statement in one HTTP call.

    Now picture 43,000 McKinsey consultants trusting Lilli to help them build financial models and strategic recommendations for the world’s biggest companies. If someone quietly told the AI to skew its analysis, or to embed confidential data into responses that consultants then copy into client decks? Nobody would know. There’s no log trail for a modified prompt. The AI just starts behaving differently and everyone assumes it’s working fine.

    CodeWall nailed it in their writeup: organizations have spent decades securing their code, their servers, and their supply chains. But the prompt layer; the instructions that control how AI actually behaves; is the new high-value target. And almost nobody is treating it that way.

    The Irony Is Thick Enough to Cut

    McKinsey isn’t some scrappy startup that skipped security to ship faster. They have world-class engineering teams, real security budgets, and every resource you could ask for. Their CEO has said AI advisory work accounts for about 40% of revenue. They’ve built 25,000 AI agents for their own workforce. They point to Lilli as proof they practice what they sell.

    And a 30-year-old bug took them down. Their own internal scanners missed it for two years. An AI agent found it in minutes.

    I tell my strategy students at BYU: watch what companies do, not what they say. McKinsey sells AI strategy to Fortune 500 boards. They tell those boards to take AI security seriously. And their own AI platform was wide open to a decades-old attack.

    The moral? Speed kills when you’re not watching the basics. And the basics haven’t changed just because the technology got fancier.

    The Agent Chose Its Own Target

    The detail that really got my attention: CodeWall’s agent picked McKinsey on its own. The CEO, Paul Price, told The Register that the research agent cited McKinsey’s public responsible disclosure policy and recent updates to Lilli. No human selected the target.

    That’s new. I can’t think of a precedent for it.

    And CodeWall did it again days later. They pointed their agent at Jack & Jill, a well-funded AI recruitment platform whose clients include Anthropic, Stripe, and Monzo. The agent chained four individually harmless bugs into a complete organizational takeover in under an hour.

    Then it did something nobody expected. It gave itself a voice. It started a real-time conversation with the target’s AI agent. At one point, it impersonated Donald Trump and demanded access to all candidate data.

    I know that sounds ridiculous. It is. But it also proves a point: autonomous AI agents don’t follow human playbooks. They improvise. They chain things together in ways you wouldn’t predict. And they do it at machine speed, continuously, without getting tired or bored or distracted.

    A human pen tester might find one of those four bugs at Jack & Jill and think “interesting, but not exploitable.” The AI agent found all four and saw the connections between them.

    A Billion-Dollar Market Nobody Saw Coming

    If you’re reading this through a strategy lens (and you should be), the McKinsey breach is also a market signal.

    The global penetration testing market sits at roughly $3 billion in 2026, growing to over $7 billion by 2034, according to Fortune Business Insights. But those numbers reflect the old model: hire a firm, run a test, get a report, file it away, repeat next year.

    That model is dying.

    The new model is autonomous, continuous, and AI-driven. Deploy agents that attack your own systems around the clock. The same way CodeWall hit McKinsey, but with your permission and on your schedule.

    XBOW, which builds autonomous offensive security agents, has raised $117 million and is reportedly in talks for a valuation above a billion dollars. Aikido Security just became Europe’s fastest cybersecurity unicorn. Companies like Pentera, Novee, and RunSybil are all building platforms that do continuous AI-powered security testing.

    For anyone thinking about where value is going to be created in the next five years: the companies that help other companies secure their AI systems are going to be enormous. The attack surface grew overnight, and the old tools can’t keep up.

    What I Did After Reading the CodeWall Report

    I run SWORN.ai. We build AI-powered wellness monitoring for police, fire, and first responder agencies. We process sensitive health data, biometric data from Oura rings, and agency operational data. Our infrastructure runs on AWS GovCloud.

    After reading CodeWall’s writeup, I didn’t just think “that’s interesting.” I pulled up every API endpoint in my personal infrastructure and my company’s systems and ran them through the exact attack chain that hit McKinsey.

    Open webhooks with no authentication? Found some. A static API token that had been pasted into conversation logs? Yep. Database access policies that were too permissive? That too.

    I’m not sharing this to be self-deprecating (well, maybe a little). I’m sharing it because if someone who builds and ships AI products for a living had gaps in their own infrastructure? You probably do too.

    Here’s what I’d tell any founder, CTO, or board member who’s deploying AI right now:

    Count your endpoints. All of them. McKinsey had 200+ with 22 that needed no auth. Most organizations don’t even know their full API surface.

    Protect your prompts like you protect your source code. If your AI’s behavior instructions sit in the same database as user data, you have the McKinsey problem. Move them. Lock them down. Version them. Monitor them for changes.

    Test for AI-specific attacks. Traditional pen testing doesn’t cover prompt injection, RAG poisoning, or agent tool manipulation. You need testing that targets how the AI itself can be turned against you.

    Assume machine-speed attackers. If your incident response plan assumes a human working over days or weeks, it’s not built for this. An autonomous agent can go from discovery to full database access in two hours for $20.

    Rate-limit everything. McKinsey’s platform let the agent run 15 blind SQL injection attempts without throttling. Every public endpoint should have rate limiting and anomaly detection. No exceptions.

    The Gap That Matters

    McKinsey has more money, more engineers, and more security resources than 99% of companies deploying AI. They still got breached by a bug from the Clinton administration.

    The companies that will do well in the next decade aren’t going to be the ones that deploy AI the fastest. They’re going to be the ones that deploy it the most securely. Right now, the gap between those two groups is enormous. And the attackers just got a lot faster.


    Sean Bair is CEO of SWORN.ai and Professor of Strategy & Economics at BYU’s Marriott School of Business. He has built and sold AI companies for 30+ years, starting with BAIR Analytics (acquired by LexisNexis, 2015) and Nouri (acquired 2025). He’s the author of AI in Policing, Business Is Personal, and AI in Business Strategy.

    Read CodeWall’s full writeup. Edward Kiledjian’s independent analysis adds good critical context. The Register’s original report by Jessica Lyons broke the story.