Data: an activity with a high environmental impact
Many companies focus their CSR strategies around actions such as limiting aerial transport or reducing waste. While important, these strategies are insufficient compared to the growing impact of other activities, such as data management. It is therefore important to reassess the priorities of companies and the pros and cons that each of these actions can represent.
In 2021, data represented between 2-4% of worldwide greenhouse gas emissions; by way of comparison, civilian air transport represented between 2-3%. The transfer and storage of data, as well as the manufacture of electronic components required for the servers, are energy-consuming processes by nature: it has been estimated that the energetic consumption of data centers in 2024 was 415 TWh, or 1,5% of total worldwide energy consumption. The rapid rise of large language models (LLMs) and GPU-intensive computing has exacerbated this significantly, transforming data infrastructure into a major and increasingly visible driver of energy consumption and carbon emissions. Additionally, one study suggests that the amount of data created or replicated per year in the world may double by the end of 2026 relative to the end of 2022. If preventative measures are not taken, the environmental impact of data could therefore explode in the coming years: in France alone, it is estimated that the digital sector will be responsible for 7% of greenhouse gas emissions by 2040.
In addition to greenhouse gas emissions, data centers have a massive impact on water resources. A U.S. federal report estimates U.S. data centers indirectly used ~211 billion gallons of water in 2023 via electricity consumption, which equates to roughly 1.2 gallons per kWh. To put this into perspective, this volume is comparable to the annual household water use of approximately 1.9 million U.S. families. Moreover, with demand projected to reach up to 1,050 TWh by 2030, water use is expected to rise accordingly.
Beyond the direct environmental consequences such as greenhouse gas emissions and water usage, another critical factor exacerbating the environmental impact of data is the growing economic competition between countries. Driven by concerns over technological sovereignty and autonomy, this competition has led to significant investments in large-scale data center projects, such as Project Stargate in the United States and recent initiatives by Mistral AI in Europe.
Despite an environmental footprint that can no longer be overlooked and is increasingly acknowledged by the public, data has remained “under the radar” for a long time and its impact underestimated. There are three possible reasons for this:
- The illusion of the immateriality of data and its relatively low financial cost, which leads to exaggerated data generation;
- Storage practices that consume too much energy and are ill-suited to current uses;
- A lack of education surrounding recent, increasingly energy-consuming, technological advancements.
To fully grasp the scale of the challenge, it’s essential to examine one of the primary culprits behind this environmental impact, hot storage.
Hot data storage hell
Today, often due to lacking optimization and organization, we have the tendency to store all of our data using “hot storage”.
Contrary to industries based on data transfer such as videoconferencing (mentioned in the first article of this series on sustainable data management), the environmental impact of data in most industries is largely due to storage. Indeed, the difference between data storage and transfer is that the first consumes electricity every day, while the second only consumes at the moment of transfer. Additionally, storage consumes more electricity per byte than data transfer. This is due to our method of storing data.
Often due to lack of optimization and organization, we have the tendency to store most data in “hot storage”. This means that our data are stored on servers which are accessible at any moment, and therefore permanently turned on. This hot storage method is the most common, but also the most consuming. Indeed, ingesting, storing, and accessing 1 GB of data on a cloud server, for example, has a carbon footprint of around 105g of CO2eq per year, equivalent to a 40km drive in a car. While this may seem anecdotal for a year’s time, we are talking about a single gigabyte; companies generate and store thousands of gigabytes of data every year. The impact of storage is even higher if these collaborative documents are stored on local servers, as they are powered by communal electricity and rarely offer algorithmic optimization of data management, unlike the Cloud. Their usage can still be necessary for cybersecurity reasons, however.
Read more: Green cloud computing: a sustainable solution for reducing the environmental impact of data storage?
Cold data storage: a more ecological data storage solution
What is cold data storage?
When data is no longer used on a daily basis, “cold storage” is sufficient.
Immediate accessibility at any time is not practical for a large part of our data, notably for those that we keep for archiving or compliance reasons, and therefore consult rarely (if ever). Cold storage technology, which is on the rise for all data storage solution providers (Cloud or not), places cold data in servers which are only turned on upon request; the stored data is thus “frozen” and the servers do not consume energy unless an access request is made. This can slow down the access to information, as the data that is cold-stored can usually only be consulted several hours after the access request. However, this approach considerably reduces electricity consumption linked to data storage, and thus their financial costs and environmental impact, by a factor of up to 3.5 in 2022. Although this multiple may have changed since 2022 as storage technologies and data-center efficiency have evolved, the potential savings remain substantial given the rapid growth in data volumes and storage demand.
A leading example of a company embracing cold storage is Amazon, which has made large-scale investments in cold data storage through its AWS S3 Glacier storage classes. This service is specifically designed for long-term storage of rarely accessed data, with retrieval times ranging from just minutes to hours, enabling massive cost and energy optimization at scale.
4 steps for adopting cold data storage starting today
Cold storage of data is an effective and easy solution to implement, given that simply migrating data to cold servers is sufficient. Here are the steps to follow for implementing this migration today:
- Step 1: identify the data currently in hot storage that should be archived
- Step 2: establish a policy of data management for the future, with periods of hot and cold storage that are well-defined for the different types of data collected (for certain types of data, hot storage could even be completely avoided)
- Step 3: organize regular and systematic archiving campaigns to ensure that this policy is respected by migrating archived files to cold servers whenever possible
- Step 4: delete the cold data at the end of the defined storage period
Digital technology has truly become a high-priority environmental challenge whose impact is still exponentially growing, notably due to data storage. Establishing an internal policy that aims to better manage stored data, as well as a larger adoption of cold data storage, can reduce (and even reverse) this trend. In the third and final category of our series, we will explore the advantages of the Cloud in terms of environmental impact. Alcimed is at your side to help you meet your challenges of data storage through a digital sobriety strategy. Contact our dedicated data and AI team, Nautilus.ai, to learn more!
About the authors,
Matthieu, Project Manager, and Maxwell, Consultant in Alcimed’s Data team in USA.