← BlogApril 24, 2026AI · Security · SOC · Incident Response · ISO 27001 · NIST CSF · MTTD · MTTR · Temporal Security

Slowness Is the Vulnerability Nobody Patches

Time is the real attack surface now. Why MTTD/MTTR are existential, how to move from spatial to temporal architecture, where to automate without blowing up like CrowdStrike, and what to actually measure when the adversary runs at machine speed.
The fastest recorded attack of 2025 moved from initial access to lateral movement in twenty-seven seconds. The record a year earlier was fifty-one seconds. A few years back, the typical eCrime breakout ran close to an hour. The average today is under half an hour, and the fast end of the distribution is well into single-digit seconds. Your ticketing system is slower than that. Your bridge call is slower than that. Your Tier 1 triage queue is slower than that. Most security programs haven't priced this in. We've spent decades buying perimeter solutions, hardening endpoints, encrypting everything at rest, and stacking identity providers in front of identity providers. All worthwhile. None of it tells you anything about the only variable that really determines blast radius now: how fast you can see, decide, and contain. Slowness in security operations isn't a process problem. It's a vulnerability. And unlike a missing patch or a misconfigured bucket, nobody puts it on the risk register.
Last September, an AI agent spent about a month systematically breaking into roughly thirty organisations — banks, government agencies, large tech firms, and a handful of chemical manufacturers. The humans driving it, later attributed to a Chinese state-sponsored group, acted mostly as supervisors. The agent ran reconnaissance, discovered vulnerabilities, exploited them, moved laterally and harvested credentials on its own. Depending on which analysis you read, the AI executed somewhere between eighty and ninety per cent of the tactical work. The humans stepped in only to authorise the big decisions: move from recon to exploitation, approve exfiltration, pull the trigger on the next target. That was the first publicly disclosed AI-orchestrated espionage campaign. It is not going to be the last. The story people missed isn't that the tradecraft was sophisticated. It wasn't, particularly. What should have chilled every SOC lead reading the writeup is that the entire campaign was run at a cadence no human operator can match, by an operator that doesn't sleep, doesn't get tired, doesn't miss a step because it's on a bridge call with its own internal stakeholders. It just executes. Run that alongside the other data points from the last year. An autonomous pentesting agent took the number-one spot on HackerOne's US leaderboard mid-way through 2025. Over 90 days, it submitted more than 1,000 validated vulnerabilities — 54 of them rated critical, 242 high. It found a zero-day in a major VPN product affecting more than 2,000 hosts. A single piece of software, running autonomously, outperformed the rest of the researcher pool. AI-generated phishing is now pulling click-through rates roughly four times higher than the old material, from somewhere around 12 per cent to 54 per cent. Deepfake videos capable of beating selfie liveness checks have grown by almost 200 per cent year-on-year. Identity attacks are up about a third compared to early 2024, and virtually all of them are automated. Meanwhile, on the defender side, the average SOC now handles close to three thousand alerts a day. More than sixty per cent of those alerts go untouched, which the industry politely calls "de-prioritised." False positives are the top complaint in almost every detection survey published last year. Over three-quarters of analysts report burnout. Most Tier 1 hires leave within three years. This is the asymmetry. It isn't ideological. It's physics. One side removed all the friction; the other side didn't.
Most programs report MTTD as if it means something. Almost all of them are measuring the wrong thing. What usually gets reported is alert-to-acknowledgment. The ticket pops in the queue, an analyst picks it up, the clock stops, and the number goes on the dashboard. That's useful for capacity planning. It's almost useless as a security metric because the actual dwell — the time from compromise to the first real signal — falls entirely outside that measurement. Ransomware intrusions get discovered in about six days on average, and the only reason the number looks that tame is that the attacker drops a ransom note. If the adversary isn't kind enough to notify you, internal discovery typically takes about 10 days. External notification — someone else calling to tell you your data is on a leak site — runs past three weeks. None of those numbers is survivable against an agent that finishes its work in a few hours. The second thing most teams get wrong is reporting MTTR. MTTD and MTTR are distributions, not single numbers. A five-minute median with a four-hour tail will absolutely destroy you on the day the tail fires. Report p3 (baseline) internally if needed. But your leadership needs to see P2 (major issues) and P1 (headline-grabbing worst cases), because that's where the breach lives. The worst day is the one that ends up in the news. And the chain needs to be separated out. Detect, investigate, contain, and recover are different problems with different fixes.
  • MTTD — compromise to the first real signal
  • MTTI — first signal to an analyst actually investigating, not just acknowledging
  • MTTC — decision to contain the blast radius
  • MTTR — start of the incident to restore normal operations
Most programs conflate three of these four and then wonder why their improvement work doesn't move the needle. It's hard to optimise a chain you can't see.
Most security architecture is spatial. Perimeters, zones, segments, DMZs, enclaves, trust boundaries. Even zero trust, for all its genuine value, is fundamentally a reorganisation of where you draw the lines. Defence-in-depth: more lines. Microsegmentation: smaller lines. The next leap is temporal. The question that actually matters now is: what can an attacker do in the first thirty seconds of access? The first five minutes? The first hour? Design the program around those windows, not around the diagram. A flat network with thirty-second containment beats a beautifully segmented one with a four-hour response loop. Not always, but often enough to change how you prioritise. When Scattered Spider made a ten-minute phone call to a Las Vegas casino's help desk and walked out with domain admin over Okta and Azure, no amount of segmentation was going to save that environment. The decision loop was simply too slow. MGM found that out the hard way. Caesars, who paid the ransom, found out in a different direction. Half the consumer sector spent the next eighteen months re-reading that playbook and realising they'd have lost the same way. In my own work, time-to-containment is the single best predictor of breach impact I've seen. Better than architectural complexity. Better than the control count. Better than headcount. Nothing else comes close. This isn't a new idea, by the way. Militaries have been thinking in OODA loops for sixty years. The security industry is finally catching up, largely because the adversary forced the issue.
The obvious response to a speed problem is to automate. It's also how you end up being CrowdStrike in July 2024. One bad content update, shipped as an automated security response, blue-screened roughly 8.5 million Windows machines in a few hours. Airports ground to a halt. Hospitals went back to paper charts. Delta alone put the damage at more than half a billion dollars. The full global direct loss probably cleared ten billion. It was the largest IT outage in history, and it came from the kind of tool you'd want on your side. The lesson isn't "don't automate." That argument is already over; the attackers made the decision for you. The lesson is that unchecked autonomy carries its own blast radius, and the real engineering problem is matching the degree of autonomy to the confidence of the signal and the consequence of the action. The model I keep coming back to when I'm advising teams looks like this:
  • High confidence, low blast radius. Full automation. Blocking known-bad indicators, quarantining confirmed phishing, killing a newly compromised service account, isolating an endpoint showing textbook beaconing. If you're routing these through a human, you're burning time you can't afford to lose.
  • High confidence, high blast radius. One-click approval. Evidence gets pre-assembled, the response is staged, and a human clicks go. Target: sixty seconds from alert to action. This is where most programs should be investing right now, and almost none are.
  • Lower confidence, any blast radius. Auto-enrich, human decides. The agent pulls context from fifteen sources, builds the timeline, and drafts the investigation note. The analyst makes the call.
The failure mode you are trying to avoid is automation that acts on ambiguous signals with a high blast radius. That's the CrowdStrike pattern, and it's also why most SOAR deployments quietly get rolled back after the first major false-positive outage. The other thing worth saying about automation, since nobody in the AI-security content mill seems willing to: you cannot automate your way around a compliance story. Every automated response needs to map cleanly to a documented control, a recoverable rollback, and an evidence trail an auditor can follow. If your SOC is running playbooks GRC has never seen, you are exactly one bad quarter from a very uncomfortable conversation.
Speed is a property. Measure it like one. Detection latency by attack pattern. Credential stuffing shows up in minutes. DNS-tunnelling exfil takes days. OAuth token abuse — how most of the big SaaS-to-SaaS breaches of the last year started, including the one that chained through a chatbot integration and hit seven hundred downstream companies — routinely goes undetected for a week or more because the traffic looks like normal API activity. Break your detection-latency numbers down by technique and go find your three-day gaps. Containment at time thresholds. Not "how fast on average." How many incidents can you actually contain at T+30s, T+5m, T+30m, T+2h? If the number at T+5m is zero, put that on the slide. Honesty is worth more than a pretty graph. Decision latency. The gap between an analyst having enough evidence to act and an action actually happening. This is almost always the longest link in the chain, and it's the one nobody instruments. It involves bridges, approvals, change management, and sometimes legal. Put timestamps on evidence-sufficiency and action-taken. You'll be surprised how much of your MTTR is just waiting for a human to be available. Recovery by the critical system. Not the number in your BCP document. The number from your last real test. If you haven't tested in the past 12 months, the number in the document is fiction. Tail numbers, not means. P1 (headline-grabbing worst cases), P2 (major issues), and P3 (baseline performance) on every latency metric. The tail is where you die. Review all of this quarterly. Put it on a dashboard next to your vulnerability backlog. Speed should be treated with the same seriousness as a CVE with a CVSS score of 9 or higher. In most places, it still isn't.
If I'm stepping into a new program tomorrow, with one quarter to move the real needle, this is roughly where I'd start. Pick the five highest-consequence incident scenarios. Ransomware with domain-wide encryption. OAuth or SaaS token compromise. Insider exfil of customer data. Domain admin takeover via help-desk vishing. Supply-chain compromise via a trusted integration. For each one, write down the honest answer to: What do we actually do in the first sixty seconds? For the scenarios where the answer is "nothing" — and there will be at least two — wire up a graduated-autonomy response. Even a scripted one-click isolation beats a forty-minute bridge call. Instrument decision latency immediately. Most teams don't. Once you can see it, you can shrink it, and the ROI on that single change tends to eclipse any shiny-new-tool purchase on the roadmap. Re-baseline MTTD and MTTR as distributions. Report p1 (critical outliers), p2 (major issues), and p3 (baseline performance) to prioritise fixes from headline-grabbing worst cases to everyday norms. Retire alert-to-acknowledgment as your headline detection metric. Replace it with compromise-to-first-real-signal. This is harder — you have to run actual purple-team scenarios and measure how long it took the first true detection to fire. It's uncomfortable. Do it anyway. The number you're afraid to see is the one you need most. Integrate the GRC side on day one. Every automation step gets a control mapping. Every playbook gets a compliance artifact. Do this while you're building it, not afterwards. Retrofitting compliance onto a live automation stack six months in is how you end up ripping the whole thing out and starting over. None of this is novel. Bits of it are in every mature threat report of the last two years. What's novel is that most programs still haven't done the work, and the adversary stopped waiting.
Most security teams I've worked with optimise for audits. The controls get built for spreadsheets. The playbooks get built for PDFs. The architectures are built for assessors. It's rational — your budget comes from the people who read those documents, not from the breach that hasn't happened yet. The attackers don't care about your spreadsheet. They care about seconds. The programs that survive the next few years won't be the ones with the most controls. They will be the ones with the fastest loops. Detect, decide, contain — measured in seconds, practised often, tested against the kind of agent that burned through thirty organisations in a month last year. If you stress-tested your SOC this week against a machine-speed adversary, where does the chain break first? Not the board-pack answer. The real one.