Compute refers to the computing power used to run software, and in AI more specifically the resources (GPUs, accelerators, servers) needed to train and serve models.
The word moved from engineering jargon to executive vocabulary with the rise of large language models. A model's capabilities improve largely with the amount of compute invested in training, which makes it both a major expense and a strategic issue. People talk about "access to compute" the way they would talk about access to capital.
In practice, compute is bought (in-house servers and datacenters) or rented (cloud, specialized GPU providers). AI companies constantly trade off training compute, to build or adapt a model, against inference compute, to serve it to users.
Why it matters when hiring
In an AI start-up, the compute constraint shapes the profiles to hire: you need engineers who can get the most out of a limited compute budget. In interviews, ask how the candidate reduced a training or inference cost, and with which trade-offs. An answer quantified in GPU hours or latency is worth more than a general speech about "scaling".
