
TileLang
TileLang is a programming language specifically designed for writing computations on graphics processing units (GPUs) in a way that executes as fast as possible. It is primarily used when operating large AI models, where computational speed directly translates into cost savings.
TileLang is a programming language for graphics processing units, or GPUs for short — the chips on which AI models run. It was developed to solve a specific bottleneck: normal languages describe what should be computed. TileLang additionally describes precisely how the data is split up and processed in parallel on the chip. This splitting into small blocks, called “tiles” in English, is baked directly into the name. The goal is to utilize the hardware as fully as possible — because a GPU that is waiting still costs electricity and money.
Why speed is a cost problem in AI
Large language models such as GPT or Gemini are used by millions of people simultaneously. Every answer such a model computes costs the operator real money — for electricity, for rented server time. Whoever computes the same answer in half the time pays half as much. That is why it pays off to invest enormous effort into optimizing the computations.
The problem: modern GPUs have a complex internal structure. They consist of thousands of small computing units organized into groups, and each group has its own very fast local memory. If this memory is used poorly, the computing units are constantly waiting for data. It’s as if a hundred cooks were standing in a kitchen, but only a single one fetches the ingredients from storage — most of them are just standing around. TileLang gives developers a precise language to control exactly that: who fetches which data when, and who computes with whom when?
How TileLang splits up computations
The core idea is tiling: a large computation — such as multiplying two huge numerical matrices — is cut into many small blocks, the tiles. Each block fits into the fast local memory of a computing unit group. This way, the slow main memory of the GPU doesn’t need to be accessed all the time. The result: the computing units are kept busy almost continuously, instead of waiting for new supplies.
TileLang makes this pattern explicitly controllable. Developers write directly in the language how big a tile should be, which computing units are responsible for it, and in which order the blocks are processed. That sounds like a detail, but it is the decisive difference between code that uses 20% of the possible computing performance and code that achieves 80%. Other approaches — such as Triton, an older language with a similar goal — abstract this layer away more strongly. TileLang gives it back to the developer, but in return demands more knowledge about the target hardware.
It is also important to note: TileLang is not a general-purpose language. You don’t write web applications or databases with it. It is what’s called a domain-specific language — a language built precisely for one area of application. In this case: highly optimized computational code for AI hardware.
TileLang in practice and in the news
TileLang mainly appears in the context of open-source AI projects. The Chinese AI lab DeepSeek made the language known because it used it to optimize parts of its model computations so that the cost per answer dropped drastically — an important factor in the competition with expensive American AI hardware. When tech media report that a model is “more efficient” or “cheaper to operate,” it’s often precisely this kind of low-level optimization behind it.
For ordinary users, TileLang remains invisible. It is infrastructure, like the transmission in a car: you don’t notice it directly, but without good tuning the car drives slower and consumes more. Developers working on AI frameworks — for example on open-source projects like FlashAttention or vLLM, which accelerate the inference of large models — are the actual target audience. They use TileLang to hand-write individual, computationally intensive operations that automatically generated code cannot make more efficient.
In the long run, the question is how much control developers will hand over to automatic tools. Compilers — programs that automatically optimize code — are getting better and better. But for the most performance-hungry AI applications, automatic optimization is not yet sufficient today. As long as that remains the case, languages like TileLang have their place.