AMD2 mins read

AMD acquires Taalas to push AI models directly into silicon

AMD is buying Canadian startup Taalas, whose inference-chip approach hard-codes AI model architecture and trained parameters into silicon for very fast performance, with the tradeoff of locking each chip to one model.

AMD logo wall
Image credits:The Decoder

The deal: AMD is buying Taalas

AMD is acquiring Canadian AI startup Taalas, which builds specialized chips for AI inference. Taalas was founded in Toronto in 2023 and came out of stealth in February with a design that embeds a model’s architecture and trained parameters directly into the chip.

The move gives AMD another path for serving the rapidly growing AI inference market. AMD plans to fold the technology into its accelerator roadmap and offer it alongside Instinct GPUs as a system-level solution.

Why the chip approach is fast — and restrictive

Taalas demo chip benchmark for Llama hard-coded inference performance
Image credits:Taalas

Taalas’ core idea is to hard-code model weights directly into silicon. That can make inference extremely fast, but it also locks each chip to a single model.

The reported tradeoff is clear: speed and efficiency on one side, flexibility on the other. Buyers and developers would need to weigh whether model-specific hardware fits their deployment plans.

The benchmark that makes Taalas notable

A Taalas demo chip running Llama 3.1-8B reportedly reached more than 16,000 tokens per second per user. The article describes that result as many times faster than competing hardware.

The same constraint still applies: the demo chip was locked to that one model. That makes the performance figure striking, but it should be read together with the limits of model-specific silicon.

What to watch next

The acquisition is subject to standard regulatory approvals. AMD says the deal strengthens its AI portfolio, while Taalas says AMD provides the scale and reach the startup needs.

The broader signal is that major AI hardware players are exploring more specialized inference designs. Google is reportedly working on a similar approach for Gemini, suggesting model-in-silicon hardware could become a more visible part of the AI chip race.

Discover More