Can AI Chips Actually Survive in Space?



Uploaded image Image credit: Google (https://blog.google/innovation-and-ai/technology/research/google-project-suncatcher/)

Google is preparing to test its AI hardware in orbit as part of Project Suncatcher. But moving high-performance computing from a data centre to space introduces some very different engineering problems.

Modern AI accelerators normally operate in carefully controlled data centres, surrounded by power supplies, cooling equipment and high-speed networking. Google is now preparing to find out what happens when that hardware is sent into low Earth orbit.

Project Suncatcher is Google's research programme exploring whether space could eventually host scalable machine learning infrastructure. The idea is ambitious: future constellations of satellites could carry Tensor Processing Units (TPUs), generate electricity from solar arrays and communicate through high-bandwidth optical links, effectively distributing AI computing across multiple spacecraft.

The immediate experiment is much smaller. Google is preparing a prototype satellite for SpaceX's Transporter-18 rideshare mission, developed in partnership with Planet, to gather real-world data on how its TPUs cope with launch and the space environment. It is not an orbital data centre, but it represents an important first test of whether hardware developed for terrestrial AI computing can operate reliably in space, where launch forces, radiation and thermal management present very different engineering challenges.

Getting An AI Accelerator Into Orbit

The first challenge begins before the satellite reaches space. Google says the journey into low Earth orbit lasts around ten minutes, during which the spacecraft can experience sustained acceleration loads of up to 10 g alongside intense vibration. Individual components can experience considerably higher forces, with Google citing loads of 50 to 100 g for components such as its TPU chips.

Project Suncatcher hardware has therefore undergone vibration testing on three axes to reproduce frequencies associated with launch. According to Google, the hardware survived those ground tests, but the orbital mission will provide the opportunity to see how the complete system performs through an actual launch.

Once in orbit, the electronics face a different problem. Energetic particles associated with cosmic radiation and solar activity can interact with semiconductor devices. One possible result is a change in the state of stored or processed data, such as a bit flip, even where the device itself has not permanently failed.

Google has already subjected its Trillium TPUs, also known as its v6e Cloud TPUs, to proton-beam testing. The devices were exposed to a 67 MeV proton beam while running AI workloads, allowing researchers to examine total ionizing dose and single-event effects.

The results provide an interesting indication of how hardware designed for terrestrial data centres might behave in orbit. Google's research found that the high-bandwidth memory (HBM) subsystems were the most sensitive part of the tested hardware, with irregularities beginning after a cumulative dose of 2 krad(Si). Google estimates that a shielded five-year mission would expose the hardware to around 750 rad(Si), while no permanent failures attributable to total ionizing dose were observed on a tested chip up to 15 krad(Si).

That does not make a Trillium TPU a radiation-hardened space processor, nor does a ground-based proton test reproduce every condition encountered in orbit. It does, however, provide evidence that commercial AI hardware could potentially operate within the radiation environment expected for the proposed low Earth orbit application.

How Do You Cool An AI Chip In A Vacuum?

Radiation is only one problem. High-performance computing hardware also generates substantial heat, and the cooling infrastructure surrounding terrestrial AI accelerators cannot simply be transferred into space.

The key difference is the absence of convection. NASA notes that heat transfer in a vacuum is limited to conduction and radiation, with no surrounding air available to carry heat away. Heat generated by processors can still be conducted through the spacecraft and transported away from the electronics, but ultimately excess heat has to be rejected to space through thermal radiation. This makes radiator area an important consideration as the amount of computing hardware, and therefore the heat load, increases.

For Project Suncatcher, Google is investigating a combination of heat pipes and radiators. Heat pipes provide a way of moving thermal energy away from heat-generating electronics towards a cooler surface, typically one coupled to a radiator. NASA describes conventional heat pipes as passive devices in which a working fluid evaporates at the hot end, travels through the pipe and condenses at the cooler end before returning through a wick structure.

Google has already tested its TPU cooling approach in a thermal vacuum chamber. This type of testing is established practice in spacecraft development, with thermal vacuum facilities used to expose hardware to vacuum and temperature conditions before flight. The orbital experiment provides the next opportunity to see how the system performs in the environment for which it is being developed.

This becomes more than a packaging problem if orbital AI is ever expected to scale. Adding more compute increases electrical power demand and the amount of heat that must ultimately be rejected. Future orbital computing systems would therefore have to treat power generation, processor density, heat transport and available radiator area as closely connected parts of the spacecraft design.

One Satellite Does Not Make A Data Centre

Even if AI accelerators can survive launch, radiation and thermal conditions, a single satellite carrying compute hardware does not provide the equivalent of a terrestrial AI cluster.

Modern machine learning workloads depend on moving large quantities of data between processors. Google's longer-term Project Suncatcher concept therefore proposes connecting multiple compute satellites using free-space optical links. Instead of copper traces or fibre running between racks, data would have to travel between spacecraft using light.

This creates a very different networking problem. Google's research explores compact satellite formations specifically because high-bandwidth communication becomes increasingly difficult as the physical distance between nodes grows. One example model considers an 81-satellite cluster within a radius of approximately one kilometre, requiring the spacecraft to maintain a controlled formation while communicating through optical inter-satellite links.

The satellites would therefore have to behave as more than independent computers that happen to share an orbit. A scalable system would require compute, communications and orbital control to work together closely enough for distributed machine learning workloads to operate across the constellation. Project Suncatcher's initial orbital experiment will not demonstrate such a large-scale AI constellation. Its purpose is much more fundamental: determine how the hardware behaves in the real environment before attempting to scale the architecture.

Why Put AI Computing In Space At All?

The reason Google is investigating such a difficult engineering problem comes down largely to energy. AI infrastructure requires large amounts of electrical power, and scaling terrestrial data centres means providing that electricity along with land, grid connections and cooling infrastructure. Space presents an unusual alternative because solar generation behaves differently outside Earth's atmosphere.

Google estimates that, in a suitable orbit, a solar panel could generate up to eight times more energy than an equivalent panel on Earth and could produce power almost continuously. A sufficiently large solar-powered satellite constellation could therefore access substantial energy without drawing that power from terrestrial electrical grids.

That advantage comes with considerable engineering and economic costs. Hardware has to be launched into orbit, failed systems cannot simply be reached by a technician, heat rejection remains a major constraint and an orbital computing cluster requires high-bandwidth communications between moving spacecraft. Google's research also identifies launch cost as a critical factor and models a scenario in which launch prices to low Earth orbit fall below $200 per kilogram by the mid-2030s. That is a projection rather than a guarantee, but it illustrates how dependent the concept is on developments outside semiconductor technology itself.

From Four Walls To Low Earth Orbit

Project Suncatcher should not be mistaken for Google moving its AI data centres into space today. The current programme is experimental, and the upcoming satellite is intended to answer much narrower questions about how TPU hardware, cooling and other systems perform in a real orbital environment.

Those questions are interesting precisely because modern AI hardware was not originally developed around the requirements of spaceflight. Google's proton testing has already identified HBM as particularly sensitive to radiation effects, while its thermal experiments show why cooling has to be reconsidered when convection is removed. Scaling beyond an individual spacecraft would introduce another set of problems around optical communications, formation flying, power generation and distributed computing.

If those problems can be solved, the attraction is clear: access to substantial solar energy, increasingly capable AI hardware and potentially lower launch costs could make orbital computing more plausible than it once appeared.

For now, however, Project Suncatcher is testing a much simpler proposition. Before anyone can build an AI computing cluster in space, the chips first have to survive getting there.


You may also like

The Component Club

About The Author

The Component Club Editorial Team covers new electronic components, emerging technologies and the engineering developments shaping the electronics industry.

Avnet Silica IoT Podcast
Avnet Silica At The Edge
DigiKey
Avnet Silica At The Pulse