Skip to main content

    Cheap to Run, Costly to Own

    by Juan HernandezRead on LinkedIn

    TL;DR — Key Takeaways

    • McKinsey: 32% of companies skipped a software purchase because AI agents could build it. Profit impact from AI stayed flat.
    • Meta priced real-time call transcription at $0.18 an hour, 80% below Google Cloud, with a 17.5% speaker-label error rate.
    • Anthropic's Fable 5.1 keeps its price but cuts cache reads 75%, making 38-hour agent runs up to 45% cheaper.

    McKinsey found that a third of companies now skip buying software because an AI agent can build it, yet the share seeing profit from AI didn't budge. Meta priced real-time transcription at 18 cents an hour. And Anthropic answered the "too expensive" complaint about Fable with the same sticker price and a 75% cut on the memory agents reuse. The work is getting cheap. Owning the result is the bill that's left.


    The Bottom Line

    This week the cost of doing the work fell again: a coding agent builds the tool, a phone call transcribes for pennies, and an agent can grind for 38 hours on a budget that finally adds up. What did not fall is the cost of owning what you started: the code someone has to maintain, the transcript someone has to check, the credentials an agent holds at 3 a.m. Buy the hard core, build the thin layer, and give every build a name and a run-cost line. Tell me in the comments: would you cancel a software subscription because an AI agent can build the replacement in a week? Or is maintenance the bill nobody counted?

    See what we build →


    Top 3 Stories This Week

    01 · McKinsey · State of AI 2026 · Build vs. Buy

    A Third of Companies Stopped Buying Software. Profits Didn't Move

    Every software renewal now arrives with a rival in the room: an AI coding agent that says it can build the same thing in a sprint. McKinsey's State of AI 2026 survey, published August 25, asked 1,719 business leaders whether they had acted on that. 32% said yes. Their organization declined at least one software product or feature because agentic coding tools could build it in-house. Tech firms led at 41%. Insurers sat at 19% and government at 17%, the sectors whose auditors make them count. Now the twist. The same survey found the share of companies attributing any earnings impact to AI stayed flat at 37%, and the "high performers" who get 5% or more of profit from AI stayed stuck at 6%. A third of the market walked away from a purchase, and the P&L did not notice. One reason: writing the first version is the cheapest part of software. Maintenance runs 60% to 90% of a system's lifetime cost, and that bill moved from the software line to payroll, where nobody tracks it against the license it replaced. The pressure to build is real. SaaS prices rose 16.4% in June, per procurement firm Vertice. So is the trap.

    Microservice Opportunity: Build the thin layer. Buy the hard core. The right build is narrow: one integration, no regulated data, a named owner for the next two years, replacing a tool you use three features of. A quote-follow-up bot, a job-costing rollup, a supplier-invoice matcher. Then treat it like a vendor: a five-year number, a run-cost line at 15% to 25% of the build per year, and a named person who gets the call when it breaks.


    02 · Meta · Muse Voice Transcribe · Speech AI

    Transcribing Every Phone Call Now Costs 18 Cents an Hour

    The phone is still where most small businesses win or lose the job, and almost none of those calls leave a record. On September 1, Meta Superintelligence Labs released Muse Voice Transcribe, its first real-time speech model, and priced it at $3 per 1,000 minutes. That is 18 cents an hour, about 80% below Google Cloud Speech-to-Text's standard $0.96. One model listens in 80-millisecond chunks, writes words as they land, tags who is speaking across 20-plus voices, and notices when a person has stopped talking. Meta trained it on more than 70 languages, verified 25 at launch, and built it to follow callers who switch languages mid-sentence. On Artificial Analysis's independent streaming benchmark it posted a 3.1% word error rate, ahead of Cartesia, ElevenLabs, OpenAI, and Google's Gemini 3.5 Transcribe, which shipped six days earlier. Read the fine print before you build on it. That benchmark is English only. And Meta's own number for labeling who said what is a 17.5% error rate. The words are near-perfect. The speaker tags are not.

    Microservice Opportunity: Give every call a paper trail. A microservice (a small AI tool that does one specific job) records the call, transcribes it for pennies, and writes a three-line summary plus a follow-up task into your CRM or job board. Add after-hours voicemail triage in English and Spanish. Keep the transcript, but treat speaker labels as a hint, not evidence, until a person checks them on any disputed call.


    03 · Anthropic · Claude Fable 5.1 · Long-Running Agents

    The AI Night Shift Just Got 45% Cheaper

    Three weeks ago, Ramp's card data showed businesses giving Anthropic's best model just 6% of their tokens because the price was too high. On September 1 Anthropic answered with Claude Fable 5.1. The sticker did not move: still $10 per million input tokens and $50 per million output. What moved was the memory. When an agent rereads context it has already processed (your instructions, your files, the conversation so far), that "cache read" now costs $0.25 per million tokens, down from $1.00. Anthropic says typical workloads get about 25% cheaper and heavily agentic ones up to 45%, because rereading is most of what a long-running agent does. And long-running is the point. Ramp told Anthropic that one 38-hour unattended run rechecked a prior result, found the flaw, launched six experiments overnight, and came back with next steps. Millennium says it traced a one-in-a-million crash nobody had explained in four to five years. On Anthropic's own business-workflow test, Fable 5.1 scored 31.4%, up from 17.1%. Nearly double, and still failing two tasks in three. Also new: Enterprise Frontier Safeguards let companies keep Claude's data in their own cloud at no extra charge, rolling out this fall.

    Microservice Opportunity: Stand up a night shift. One agent, one bounded job, with your business context cached once so the reread is nearly free: reconcile a month of invoices against bank lines, audit every product listing for wrong prices, chase every quote that went quiet. It checkpoints as it goes and writes a morning report a human reads before anything changes. Read-only credentials by default. The 38-hour run is impressive. The approval step is what keeps it safe.


    This week in numbers: 32%: companies that skipped a software purchase because an AI agent could build it. Profit impact from AI: flat at 37%. · $0.18: Meta's price for one hour of real-time transcription. Google Cloud's standard rate: $0.96. · 45%: cost drop on heavily agentic Fable 5.1 workloads after Anthropic cut cache reads by 75%.


    Sources

    Back to all posts
    G8 Engineering — Built, not briefed

    We architect conversational intelligence that makes every system in your organization accessible to anyone — with the discipline to do it securely, and the craft to make it last.

    © 2026 G8 Engineering. All rights reserved.