The Architecture of Groq's LPU

Abhinav Upadhyay

Mar 1, 2024

What powers the ground breaking performance of Groq's Langauge Processing Unit?

Read →

12 Comments

Logan Thorneloe

Oct 30, 2024

Revisiting this because it's incredible. Good work!

Thanks, Logan!

Great post, thank you very much!

Viswa Kumar

Mar 2, 2024

Great post Abhinav. Learnt a lot. I wonder if you also cover or point me in the direction to understand how a tensor operation would become different at the library level from application point of view . For eg what changes (if any) needs to be done in either pytorch or the likes to better conquer this massive parallelisms offered by TSUs or is this completely taken by the Groq’s compiler behind the scenes . I understand Groqs hasn’t published anything yet but if you came across any nuggets on your research pls do share!

Reply (1)

Abhinav Upadhyay

Mar 3, 2024

Not a lot of details on it. But looks like their compiler can take a pytorch Or tensorflow model and compiler for their hardware. But the groq twitter account also hints that sometimes they have to rewrite the code. So it's not quite clear in what situations the compiler works without any manual intervention.

I'm just guessing here.