Unfortunately the issue (alluded to in the blog post you linked) is that transpo...

musebox35 · 2025-06-07T12:39:50 1749299990

I totally agree that the resulting kernel will be rarely useful. I just wanted to highlight that it is a commonly used educational exercise to showcase how to optimize for memory throughput. If the post showed how to fuse a transpose + rmsnorm epilogue to a gemm then the kernel would be more functional but the blog post would be much harder to follow for newcomers.

Jay Shah’s later articles contain examples that involve epilogue fusion. IMHO, understanding how to write an efficient transpose helps with following the more involved ones.

saagarjha · 2025-06-16T11:41:52 1750074112

It's less that the result is kind of useless and more that hitting memory throughput on a simple algorithm like this is not very difficult. It takes a complex example to actually have trouble doing this.

simon_vtr · 2025-06-07T13:51:47 1749304307

That was exactly my reason to write this blogpost and optimise transpose. It is a simple educational yet not trivial example to learn the basics.