Recent reports about Europeans rushing to buy Chinese air conditioners have been everywhere. But here is another possibility:
The thing that needs air conditioning most may be AI supercomputing. (doge)
Over in the UK, something like that happened over the past few days:
Dawn, one of the UK's most powerful AI supercomputers, went down for a full week in temperatures above 30°C.
The supercomputer, housed at the University of Cambridge, comes with serious credentials:
It is a core part of the UK government's £300 million national AI computing power program, with 1,024 Intel GPUs and 256 liquid-cooled servers, and has already supported more than 350 research projects.
In January this year, it received £36 million for an expansion and upgrade that is expected to boost performance sixfold.
Then, in late June, a heat wave arrived, and Dawn went offline.
The more surreal part: some of the research running on the system included climate change simulations.
Yes, really. A machine built to help predict global warming was beaten by global warming.
The Darkest Hour for a National Supercomputer: 37.7°C
Here is what happened.
In June this year, the UK experienced its most intense June heat wave on record.
On June 26, the town of Lynford in Norfolk reached 37.7°C, breaking the June record of 35.6°C set in 1957 and 1976.
The UK Met Office took the rare step of issuing red extreme heat warnings for three consecutive days.
More than 1,000 schools closed, railway signals failed in the heat, and road surfaces began to melt.
Then, on June 27, as the heat wave peaked, the cooling system at the Cambridge West data center housing Dawn could no longer cope.
P.S. Lynford and Cambridge are both in eastern England, about 103 kilometers apart.
Dawn went down.
After the incident, a University of Cambridge spokesperson said:
Dawn experienced technical issues during the hot weather; cooling capacity has been fully restored, and access is expected to reopen on July 6.
The university did not specify the exact cause, but the situation was clear:
From June 27 to July 6, Dawn spent more than a week “cooling down.”
For a supercomputer that burns money by the hour and pushes scientific work forward by the second, a week of downtime is genuinely alarming.
The worst-hit case has already emerged.
A team led by Professor Vendruscolo at the University of Cambridge was using Dawn to screen molecules for new Parkinson's disease drugs.
Dawn's machine-learning capabilities can screen billions of molecules in a matter of days, searching for compounds that bind to protein aggregates associated with Parkinson's.
Using traditional methods? That would start at six months, cost millions of pounds, and cover only a small fraction of what Dawn can scan in a few hours.
A one-week outage meant this life-saving pipeline simply stopped.
Lennard Lee of the University of Oxford, who leads the UK's AI and supercomputing project for cancer vaccines, had secured 10,000 GPU hours on Dawn to use AI to accelerate target discovery for personalized cancer vaccines.
Lee had previously said:
Discoveries that used to take years can now be completed in weeks.
Lee later said no data was lost and no work needed to be redone, but the relief in his comments itself showed how serious the situation was.
The British Antarctic Survey's IceNet sea-ice forecasting model trained on Dawn was paused, and Cambridge PhD student Bill McGough's AI kidney cancer screening project trained on Dawn also stopped. Of the more than 350 projects running on Dawn, almost none escaped the disruption.
And all of this was caused by just 37.7°C.
So, the “culprit” has been found. The next question is: who, exactly, should take responsibility?
After going around in circles, it seems no one wants to own the problem.
Dawn's cooling system was supplied by USystems, a unit of France's Legrand Group. After the incident, USystems issued a statement:
Our equipment operated fully within its design specifications throughout the incident and performed normally.
Translation: the cooling failed, but do not blame us. Our equipment was never designed for this temperature.
So were the design standards too conservative, or is climate change moving too fast?
The answer may be: both.
The UK's historic June extreme temperature was only 35.6°C, and Dawn's cooling system was probably designed around that range.
37.7°C exceeded the limit.
And that “exceedance” came with almost no warning, because the last time the record was set was nearly 50 years ago.
Dawn was not the only victim.
That same week, the cooling units at Queen Alexandra Hospital in Portsmouth failed, prompting the hospital to declare a critical incident.
Operating rooms stopped, cardiac catheterization labs stopped, and imaging services stopped. The hospital told patients:
Please bring plenty of drinking water, as the hospital is very hot.
Norfolk and Norwich University Hospital, or NNUH, had it even worse:
The cooling systems for all MRI scanners failed because of high heat and humidity, forcing at least 254 outpatient appointments to be canceled.
So, in a sense:
It is not that supercomputers are fragile. The UK's entire temperature-control infrastructure was not ready for weather like this.
How Can Temperatures Just Above 30°C Knock Out a Supercomputer?
Seen over a longer timeline, Dawn's heat-related failure is not surprising at all.
In July 2022, the UK experienced what was then its hottest day on record, at 40.3°C.
Google's London data center suffered “simultaneous failures of multiple redundant systems” in its cooling system and had to shut down to protect hardware. Google Cloud's London region was disrupted for more than 18 hours before it fully recovered.
Oracle's London South data center went down the same day. Oracle used an interesting phrase in its statement: “unseasonably high temperatures.”
From 2022 to 2026, four years have passed, and a similar incident has happened again.
Which raises the question: is this problem really so difficult that it cannot be prevented in advance?
In practice, there is a real explanation for how temperatures in the 30s can bring down a supercomputer. The hardest bottleneck is cooling.
In Europe especially, facilities commonly use free cooling, a method that is naturally constrained by outdoor temperatures.
How should we understand that?
No matter how advanced a cooling system is, it ultimately has to dump heat into the outside air. Outdoor air temperature is the final bottleneck in the entire chain.
The chain looks like this:
The chip transfers heat to the heat sink; the heat sink transfers it to coolant or air; the coolant transfers it to the cooling tower; and the cooling tower transfers it to the atmosphere.
The atmosphere is the last party left holding the heat.
So when the atmosphere itself is already 37°C, it starts to run out of capacity.
More specifically, when outdoor temperature jumps from 20°C to 37°C, the cooling efficiency of cooling towers and dry coolers can drop by 40% to 50%.
Why not just turn on the air conditioning? Because compressors become less efficient in high temperatures, current rises, and they can easily overheat and trip.
Oracle's 2022 incident report put it plainly: two cooling units failed when they were required to operate beyond their design limits.
Dawn's situation this time may reasonably have been similar.
Its Dell PowerEdge XE9640 servers are equipped with direct liquid cooling, a far more advanced cooling approach than traditional air cooling.
The coolant flows directly along the chip surface, removing heat far more efficiently than blowing air over it.
But the same point still applies: liquid cooling improves efficiency inside the rack. After heat is carried away by the coolant, it still has to pass through the coolant distribution unit, the facility chilled-water loop, and the cooling tower before finally reaching the outside atmosphere. The last link is still constrained by outdoor temperature.
Once the cooling system stops, it also triggers a chain reaction.
Research shows that if cooling goes offline, server inlet temperatures can jump from 22°C to above 35°C within five minutes.
When that happens, chips start protecting themselves:
First comes thermal throttling, where they actively reduce operating speed to cut heat, causing performance to collapse; if temperatures keep rising past safe thresholds, they force a shutdown.
Operators then have only two choices:
Let the equipment power itself off, potentially damaging data;
Or carry out an orderly shutdown, protecting hardware while bringing operations to a halt.
Google, Oracle, and Cambridge's Dawn all chose the latter.
The Stronger AI Gets, the More It Fears Heat
There is something even more worrying.
As AI data centers keep expanding, temperature may have an increasingly significant impact on AI.
A few days ago, I watched Xiao Lin's on-site visit to a Huawei data center on Bilibili, and one comparison stood out:
A traditional data center rack has a power density of about 5 to 10 kilowatts, but AI training racks have already reached 30 to 50 kilowatts. NVIDIA's latest GB200 NVL72 rack reaches 120 to 132 kilowatts, while the next-generation Rubin may reach 600 kilowatts.
What does that mean? A 100-kilowatt AI rack produces as much heat as running 50 electric heaters at the same time in a space the size of a phone booth.
Imagine those small radiant heaters used in winter, then cram all of them into one cabinet. That is the cooling pressure today's AI computing power infrastructure faces.
The bigger problem is that GPUs themselves are getting “hotter.”
NVIDIA's V100 in 2017 was about 300 watts. The H100 in 2023 jumped to 700 watts. The B200 in 2024 reached 1,000 watts. By 2025 to 2026, the B300 and AMD MI355X are set to reach 1,400 watts.
In seven years, heat output from a single chip has increased three- to fivefold.
So whether measured by the number of chips or by each individual chip, the stronger AI becomes, the more it fears heat and the more cooling it needs.
At this point, two curves are clearly colliding:
Chips are getting exponentially hotter, and the planet is also heating faster.
The problem is becoming more difficult.
Google went to Finland as early as 2011 to build a data center, and Meta went to northern Sweden, both to use cold climates as natural cooling.
Elon Musk has even floated the idea of building AI data centers in space.
But the UK government only just put £36 million into expanding Dawn in January, and it is also planning a new national supercomputer in Edinburgh.
Are these facilities' cooling designs based on the British summers of the previous era, or on the new normal that is arriving? No one can say for sure.
But one thing is certain:
A supercomputer used to predict climate change was shut down by climate change-driven heat.
That is no longer a punchline. It is a real infrastructure challenge in the AI era.
Comments
00No comments yet. Be the first to weigh in.