Artificial Intelligence is Now Inside Chips

Artificial Intelligence is Now Inside Chips


While everyone's attention in the world of artificial intelligence is focused on the new models being trained in massive data centers, the real battle is being waged much more quietly and subtly. This battle isn't about training, but about "inference." Chatbots, AI agents, and integrated systems, run billions of times a day, rely on this inference process. Companies, seeing the enormous profit margins in the inference market, are fiercely competing to surpass the limitations of traditional hardware. AMD's quiet acquisition of Taalas, a Canadian AI chip startup, sends a crucial message about the future of this market.

Looking at the current state of the industry, we see that giant companies are pursuing different strategies. Recently, Nvidia made a 20 billion dollar acquisition of Groq, which produces SRAM-based LPUs offering low latency and model-independent operation. Nvidia's goal was to increase speed while maintaining flexibility. AMD, on the other hand, adopted a completely opposite and quite bold philosophy. Instead of producing a generic chip independent of any model, they acquired Taalas, a company distinguished by its ASIC architecture that directly tailors the hardware to a single AI model.

At this point, it's important to fully understand what Taalas is doing. Traditional AI chips constantly have to move data between memory and the processor during processing. This leads to both time loss and significant energy waste, creating a bottleneck known in the industry as the "memory wall." Taalas solves this VRAM-related problem at its root by permanently engraving the AI ​​model's weights directly into the silicon, i.e., the transistors. There's no model to load or data to read from external memory; the chip itself becomes the embodiment of the model.

The results of this revolutionary approach are truly incredible. Taalas' HC1 chip, running the Llama 3.1 8B model on a simple PCIe card, achieves a production speed of 17,000 tokens per second. Moreover, it achieves this performance with only one-tenth the energy consumed by an Nvidia H200 chip. The company is already known to be planning larger 27B class models that will run on desktop systems and can produce over 10,000 tokens per second with near-zero power consumption. Thus, the idea of ​​AI agents running at lightning speed on our local hardware, without pouring huge sums of money into cloud-based servers, is getting very close to reality.

Of course, a custom model etched into the hardware has a very obvious disadvantage. If the model changes or is updated, the chip becomes completely unusable because it cannot physically run anything else. Initially, many experts found this approach contrary to the nature of the hardware. However, AMD is presenting a grand vision that some models on the market have matured and become stable enough to be worth etching into the sand. Furthermore, the entire chip is not fixed; the unchanging weights of the model are written into a persistent mask ROM, while the KV cache and LoRA adapters are kept in SRAM, maintaining their flexibility. When the model is completely changed, re-manufacturing the metal layers requires approximately two months of hardware cycle time.

Strategically, this acquisition demonstrates AMD's acceptance of the inference efficiency limitations of its general-purpose Instinct GPUs and its choice to overcome this limitation with a proven hardware architecture rather than solving it from scratch. Taalas, which has received a total of $219 million in investment to date, will continue its work within AMD's AI organization under the leadership of Vamsi Boppana. The company plans to integrate this new technology not as a complete replacement, but as a complement to its existing Instinct GPUs, EPYC CPUs, and Helios systems. The initial stages of complex workloads will be computed on flexible GPUs, while the final stage, text generation, will be handled by these dedicated Taalas chips.

There is also a regional and sociological aspect to this. While Canada is a world leader in cultivating skilled AI and chip hardware teams, it faces significant challenges in independently scaling these initiatives. As we've seen before with examples like Untether AI, CentML, and Tenstorrent, the most talented hardware engineers and startups in Toronto are either acquired by US-based tech giants or relocated south once they reach a certain size. The fact that Taalas founder Ljubisa Bajic is also a former Tenstorrent and AMD employee demonstrates how intertwined this talent migration is. The saying "Education happens in Toronto, graduations happen elsewhere," often heard in the industry, remains a structural problem in the Canadian technology ecosystem.

In conclusion, the biggest debate in the AI ​​ecosystem in the coming period will no longer be about how intelligent the models are, but about how cheaply and quickly this intelligence can be delivered to the end user. On one hand, there are flexible GPUs developed with billions of dollars of investment, capable of running everything, but expensive and high power-consuming. On the other hand, there are custom silicon chips that perform a single task flawlessly, provide huge cost advantages, and are directly etched into transistors. A similar era to the specialized ASICs we saw in the crypto world is now beginning in the AI ​​market. As we move towards 2028, we will all see whether flexibility or absolute efficiency will dominate the hardware market.

How do you rate this article?

6



Cryptocurrencies and Stocks Articles
Cryptocurrencies and Stocks Articles

In this section, I will have articles about the stock market and cryptocurrencies.

Publish0x

Send a $0.01 microtip in crypto to the author, and earn yourself as you read!

20% to author / 80% to me.
We pay the tips from our rewards pool.

Page not displaying correctly?