When AI runs the grid: Anthropic's push into critical infrastructure
For most of us, large language models still sit in a browser tab — a place to draft an email, summarise a PDF, or argue with a chatbot about whether a tomato is a fruit. But quietly, and quickly, that's changing. Anthropic's expansion of its Claude Mythos tier under Project Glasswing to critical infrastructure operators in more than 15 countries marks a shift in what AI is for. The same family of models that writes wedding speeches is being wired into the systems that keep the lights on, the water running, and the hospitals open.
That's a much bigger deal than another product launch. It's a quiet reordering of how essential services are run — and it deserves a clear-eyed look before it becomes invisible infrastructure that no one questions.
What Anthropic is actually doing
According to reporting from TechCrunch and the Financial Times, Anthropic is scaling its high-assurance Claude variant — branded Mythos — to government agencies and operators of critical infrastructure across more than 15 countries. The rollout sits under Project Glasswing, Anthropic's program for getting Claude into sensitive, regulated, and security-conscious environments.
In practice, that means the model is being positioned for use in sectors such as energy, healthcare, water, telecommunications and transport — the kind of systems that, when they fail, end up on the front page. As Let's Data Science notes, this is a deliberate broadening from corporate productivity use cases into the regulated stack — places where downtime is measured in lives, not in lost ad impressions.
It's worth being precise about what an LLM does in these settings. Mythos won't be physically opening circuit breakers or dosing chlorine. Instead, it's likely to be embedded as a layer above existing SCADA, ticketing, and operational systems — summarising alarms, drafting incident reports, helping engineers query thousands of pages of maintenance manuals in plain English, and accelerating the triage of cybersecurity alerts. The model becomes the interface between humans and the complexity they already can't fully hold in their heads.
Why operators want this — and why now
Critical infrastructure has a workforce problem. Grid operators, hospital IT teams and water utilities are losing experienced staff faster than they can train new ones, and the systems they run keep getting more complex. A model that can read a 600-page substation manual in a second, cross-reference it with last night's alarm logs and tell a tired night-shift engineer where to look first is, frankly, a godsend.
There's also a security pitch. Anthropic has marketed Claude as the “safety-first” frontier lab since its founding, and Glasswing leans hard into that brand. For a CIO at a transmission company or a public hospital network, “the responsible-AI vendor” is an easier story to take to a board than “we plugged ChatGPT into our control room.”
And the geopolitics matter. The expansion to 15+ countries is, in part, a statement that Western-aligned frontier AI is going to be the default substrate of critical services in allied economies. That has implications for Australia, where energy market operators, the major banks, and the public health system are all simultaneously trying to figure out how — and how fast — to bring generative AI into operations without ending up in front of a Senate committee.
The risks no glossy press release will mention
Here's where the hard questions start. LLMs, including Claude, have well-documented failure modes: they hallucinate, they're sensitive to how you phrase a prompt, and their behaviour can shift between model versions in ways that are hard to predict. Those are tolerable quirks when you're drafting a blog post. They're a different problem when the output is feeding into a decision about whether to shed load on a hot January afternoon in New South Wales.
A few specific risks deserve attention:
- Automation bias. Operators under pressure tend to trust the confident-sounding answer on the screen. If Mythos summarises an incident incorrectly at 3am, the tired engineer is the last line of defence — and a thin one.
- Model monoculture. If a single family of models ends up embedded across power, water and health in many countries, a subtle regression in a new Claude release could propagate failures across sectors simultaneously. That's a systemic risk regulators are not yet equipped to measure.
- Supply-chain dependency. Critical infrastructure is meant to be sovereign and resilient. Hard-wiring it to a US-based frontier lab — however well-intentioned — creates a dependency on that lab's pricing, policies, uptime, and, ultimately, the geopolitics of its home country.
- Opaque reasoning. Post-incident reviews depend on being able to reconstruct why a decision was made. “The model suggested it” is not a satisfying answer for a coroner or a regulator.
What this means for Australia
Australia is a useful test case for how this should be governed. The Security of Critical Infrastructure (SOCI) Act already imposes real obligations on operators in energy, water, healthcare, communications and finance. Layering an LLM into those environments is not a procurement decision in the ordinary sense — it touches risk management programs, incident reporting and, increasingly, cyber resilience uplift requirements.
The pragmatic question for an Australian utility or hospital network isn't “should we use Claude Mythos?” It's “under what conditions, with what guardrails, and with what fallback when it's wrong?” A few principles travel well:
- Advisory, not authoritative. Use the model to surface options and summarise context, not to make or execute control decisions.
- Bounded scope. Constrain what the model can read, what tools it can call, and what systems it can touch. The more glamorous “agentic” the deployment, the larger the blast radius if it misbehaves.
- Logging and replay. Every prompt, every retrieval, every output kept in a form auditors can reconstruct. If you can't replay it, you can't defend it.
- Vendor concentration limits. Treat reliance on a single frontier lab the way you'd treat reliance on a single cloud region — something to be measured, disclosed, and mitigated.
The bigger story
Strip away the branding and Project Glasswing is a milestone in a transition that's been quietly happening for two years: large language models are graduating from novelty tools to load-bearing components of essential services. Anthropic's pitch is that Mythos is the safe way to do this, and there's a credible argument that a deliberate, security-conscious rollout is better than the alternative — which is operators quietly pasting sensitive logs into consumer chatbots, as plenty already do.
But “safer than the shadow IT version” is a low bar. The harder, more interesting work — agreeing what an LLM should and shouldn't be allowed to do inside a hospital or a control room, and how we verify that in public — is only just beginning. The lights staying on in 15 countries may soon depend on getting that conversation right.