The Conversation: "High-Performance Computing and Supercomputers: The Race to 'Exascale'"

Search
January 3, 2023
Frontier, the first "exascale" computer in the United States. Oak Ridge Leadership Computing Facility, Oak Ridge National Laboratory, Flickr, CC BY
Frontier, the first "exascale" computer in the United States. Oak Ridge Leadership Computing Facility, Oak Ridge National Laboratory, Flickr, CC BY
This year, supercomputers reached a milestone—performing one billion billion operations per second. Why and how did they get there?

On May 27, 2022, thehigh-performance computing(HPC) community announced with great fanfare the arrival of the first “exascale” supercomputer—that is, one capable of performing10¹⁸ “FLOPS,” or one billion billion operations per second (on real numbers in floating-point notation, to be precise).

The new supercomputer, Frontier—operated by the U.S. Department of Energy at Oak Ridge National Laboratory in Tennessee and featuring several million cores—has surpassed the Japanese supercomputer Fugaku, which has dropped to second place in the TOP500 ranking of the world’s most powerful supercomputers.

Frontier, not content with being (for now) the world’s most powerful supercomputer, also ranks highly in terms of energy efficiency… at least relative to its computing power, since it consumes enormous amounts of energy—the equivalent of a city with tens of thousands of residents. And the problem doesn’t stop with Frontier, since it is merely the flagship of the thriving global fleet of several thousand supercomputers.

A long-distance confrontation

This return of the Americans to the lead highlights a new battleground between the U.S. and Chinese superpowers, with the Europeans watching from the sidelines. In fact, China had caused a surprise in 2017 by snatching the top spot from the United States: at that time, we witnessed a massive influx of more than 200 Chinese supercomputers into the TOP500. Today, the top Chinese machine has been relegated to sixth place, and the Chinese have chosen to remove their machines from this ranking.

In 2008, the Roadrunner supercomputer at the U.S. Los Alamos National Laboratory became the first to reach the “petaflop” mark—one million billion FLOPS (10¹⁵). The exascale became a strategic goal for the United States, even though this goal seemed technically unattainable.

To achieve exascale performance, it was necessary to rethink the architecture of the previous PetaFlops generation. For example, at these extreme scales, the reliability of millions of components becomes crucial. Just as a grain of sand can jam a gear, the failure of a single component prevents the entire machine from functioning.

The “Energy Wall”

However, the U.S. Department of Energy (U.S. DoE) has imposed a constraint on this technological development by setting a maximum power limit of 20 megawatts for exascale deployments—a constraint known as the “power wall.” The U.S. Exascale Computing Initiative received more than $1 billion in funding in 2016.

[Nearly 80,000 readers rely on The Conversation’s newsletter to better understand the world’s major issues. Subscribe today]

To overcome this “performance barrier,” it was necessary to rethink all software layers (from the operating system to applications) and design new algorithms to manage heterogeneous computing resources—namely, standard processors and accelerators, memory hierarchies, and interconnects.

Ultimately, Frontier’s power consumption is measured at 21.1 megawatts, or 52.23 gigaflops per watt, which roughly corresponds to 150 metric tons ofCO2 emissions per day, taking into account the energy mix in Tennessee, where the platform is located. This is just below the 20-megawatt threshold set in the DoE’s goal (if we divide Frontier’s 1.102 exaflops by 21.1 megawatts, we get 19.15 megawatts).

This places Frontier second on the Top Green500 list of supercomputers that consume the fewest operations (FLOPS) per watt—a ranking that was launched in 2013 and reflects the community’s growing concern about energy issues. This ranking on the Top Green500 is good news: Frontier’s performance gains are accompanied by improvements in energy efficiency.

Overly optimistic estimates

But these estimates of digital energy consumption are too low, as is often the case in this area: they take into account only usage and overlook the significant portion of energy consumed during the manufacturing of the supercomputer and associated infrastructure—such as buildings—as well as its future decommissioning. Based on my research experience and that of my colleagues in academia and industry, we estimate that usage accounts for only about half of the total energy cost, calculated over an average lifespan of 5 years. There are few studies on this topic, due to the systemic complexity of the issue and the limited availability of data, but one notable example is a recent study measuring the energy consumption of a single core of the Dahu computing platform over one hour, which concluded that operating costs account for barely 30% of the total energy cost.

Furthermore, technological improvements that enable energy savings lead to an overall increase in consumption: this is known asthe “rebound effect.” New features and increased usage ultimately result in higher energy consumption. A recent example in computing is natural language processing(NLP), which gains new capabilities as computational performance increases [link].

The Tree That Hides the Forest

The technological advances needed to reach exascale are undeniable, but the direct and indirect impact on global warming remains significant, regardless of what the optimists say—who consider it a drop in the bucket compared to the 40 billion metric tons ofCO2 emitted each year by all human activities.

Furthermore, this is not just about a single supercomputer: Frontier is the tree that hides the forest. Indeed, the community has long observed that the advances achieved by building a new generation of high-performance computing systems spread rapidly: new platforms quickly replace those already deployed in university computing centers or in companies. If replacement occurs prematurely, the effective lifespan of the replaced machines is reduced, and their environmental impact increases.

The TOP500 represents only a fraction of the vast array of HPC platforms deployed worldwide. It is very difficult to estimate their total number because many platforms fly under the radar: a large number of large-scale platforms are located in private companies, and many smaller-scale platforms are deployed locally.

A small study based directly on TOP500 data shows that the actual performance of the most powerful platform has increased 33-fold over the past ten years (while the average performance of the 500 machines has increased by only a factor of 20). Over the same period, the energy efficiency of the TOP Green500 has improved by a factor of barely 15 (and 18 on average). The overall balance in terms of energy consumption is therefore negative—it has, in the end, increased.

What should we do with this progress?

A counterargument can be made: advances toward increasingly powerful platforms could lead to technical solutions for combating climate change. This way of thinking reflects the mindset of our technology-centered society, but unfortunately, it is virtually impossible to measure the impact of these new technologies on reducing the carbon footprint. In fact, most of the time, these measures focus on the usage phases and ignore the “side effects,” such as the manufacturing of new equipment, for example.

It is reasonable to ask what mechanism drives this race for performance. One reason cited by the designers of Frontier "is scientific progress: the more complex the phenomena we seek to model and understand become, the more we need simulations—and the only way to conduct simulations is to build ever-more-powerful HPC platforms…"The Conversation

This article is republished from The Conversation under a Creative Commons license. Readthe original article.
Published on January 3, 2023
Updated on January 26, 2023